Dome Systems

NVIDIA NIM and Dome

NVIDIA NIM on your GPUs and NVIDIA's, behind one Gateway.

Connect hosted NIM and your own NIM containers to Dome. Pool them under one name, with your GPUs first and NVIDIA's API behind them. Rules decide which callers may leave your infrastructure.

PeopleAgentsDomePool: llama-3.3-70bCallersPeopleAgentsclinical-notesAgentsscheduling-agentAgentsresearch-agentGatewaysprod-gatewaySelf-hosted NIMPriority 0, your GPUsNemotron 3 Supernvidia/nemotron-3-super-120b-a12bHosted NIMPriority 1, integrate.api.nvidia.comRulesGuardsAuditsaudit-trail
clinical-notes→llama-3.3-70b· as dr.haleAllowed

How Dome helps

Dome provides model brokering and routing for NVIDIA NIM

Your GPUs first

Pool a self-hosted NIM with hosted NIM under one name. Calls go to your container, then to NVIDIA's API.

One model name in both places

NIM containers serve the same model ids as NVIDIA's API. Agents send the pool's name and never change.

Sensitive callers stay home

A rule reads the verified user's groups. Clinical users never reach the hosted API.

Get started

Connect NVIDIA NIM in three steps

Add hosted NIM and your own container as connections, pool them, then point agents at the Gateway.

  1. 01

    Add hosted NIM

    The provider id is nvidia. Dome uses https://integrate.api.nvidia.com/v1 and NVIDIA's model ids.

    $ dome models add llama-nim \
    --provider nvidia \
    --model meta/llama-3.3-70b-instruct \
    --api-key "$NVIDIA_API_KEY" \
    --attributes '{"hosting":"nvidia-cloud"}' \
    --gateway prod-gateway
  2. 02

    Add your NIM container and pool both

    A self-hosted NIM speaks the same API. Set its endpoint, and make sure the Gateway can reach it: a Dome-hosted gateway can't reach your private network.

    $ dome models add llama-nim-local \
    --provider nvidia \
    --endpoint http://nim.internal.example.com:8000/v1 \
    --model meta/llama-3.3-70b-instruct \
    --auth-method none --credential-type none \
    --attributes '{"hosting":"self"}' \
    --gateway prod-gateway
     
    $ dome models pool create llama-3.3-70b \
    --failover-max all --gateway prod-gateway
    $ dome models pool member add llama-3.3-70b llama-nim-local --priority 0
    $ dome models pool member add llama-3.3-70b llama-nim --priority 1
  3. 03

    Point your agent at the Gateway

    Any OpenAI-compatible client works. The agent sends its Dome key and the pool's name.

    from openai import OpenAI
     
    client = OpenAI(base_url=f"{GATEWAY_URL}/v1", api_key=DOME_AGENT_KEY)
    client.chat.completions.create(
    model="llama-3.3-70b",
    messages=[{"role": "user", "content": "Draft a visit summary from these notes."}],
    )

Commands and rules tested against a Dome workspace on October 1, 2026. For anything about NVIDIA NIM itself, see NVIDIA's documentation.

Rules

Clinical users stay on your GPUs

Applied to an agent, this refuses hosted NIM whenever the verified user is in the clinical group. Everyone else can still fail over to NVIDIA's API.

forbid (principal, action == Dome::Action::"llm:invoke", resource is Dome::LLMModel)
when {
resource has hosting && resource.hosting == "nvidia-cloud" &&
principal has act_as && principal.act_as.groups.contains("clinical")
};

Agent workflow

Bringing it together

Connecting NVIDIA NIM to registered agents, tools, and identity in Dome completes a governed agent application.

Dome

Control point

Gateway

  • Rules
  • Guards
  • Quotas

Every call decided and audited

FAQ

Common questions

How does Dome connect to NVIDIA NIM?

Through NIM's OpenAI-compatible API at https://integrate.api.nvidia.com/v1, with your NVIDIA key sent as a bearer token. The key stays in Dome's vault.

Does a self-hosted NIM work with Dome?

Yes. Add it with provider nvidia and your container's /v1 URL as the endpoint, on a Gateway that can reach that URL.

Which NIM models can I use?

Any model id NIM serves, such as meta/llama-3.3-70b-instruct or nvidia/nemotron-3-super-120b-a12b. Dome passes the id through unchanged.

Next steps

Talk with our FDE team

Our forward deployed engineers work with your platform team to get your agents into production and under control: the first one governed on your own systems, and a pattern your teams can repeat for every agent after it.