Dome Systems

vLLM and Dome

Your vLLM servers behind one Gateway, with every call governed.

vLLM puts open models on your own GPUs behind an OpenAI-compatible server. Connect it to Dome, let Together AI take the overflow, and a rule decides which agents may ever leave your hardware. The Gateway needs a route to the server.

PeopleAgentsDomePool: llama-3.3-70bCallersPeopleAgentscontracts-agentAgentssupport-agentAgentssummarizerGatewaysprod-gatewayvLLMmeta-llama/Llama-3.3-70B-Instruct, priority 0vLLMQwen/Qwen3.6-35B-A3B, no poolTogether AILlama overflow, priority 1RulesGuardsAuditsaudit-trail
contracts-agent→llama-3.3-70b· as h.bergAllowed

How Dome helps

Dome provides model brokering and routing for vLLM

Any model vLLM serves

Point Dome's openai_compatible provider at the server's /v1 address. The served model name is the model id.

Overflow to a hosted provider

Pool vLLM with Together AI serving the same Llama. Your GPUs answer first and Together takes what they can't.

Some agents never leave your GPUs

Tag connections by where they run. A rule keeps an agent on self-hosted models only.

Get started

Connect vLLM in three steps

Connect your server, give the Gateway a route to it, and add a hosted fallback.

  1. 01

    Add the connection

    Use the provider id openai_compatible and the server's /v1 URL. If you start vLLM with --api-key, give Dome the same key.

    $ dome models add llama-vllm \
    --provider openai_compatible \
    --endpoint https://vllm.internal.example.com/v1 \
    --model meta-llama/Llama-3.3-70B-Instruct \
    --api-key "$VLLM_API_KEY" \
    --attributes '{"hosting":"self"}' \
    --gateway prod-gateway
  2. 02

    Give the Gateway a route

    The hosted gateway can't reach inside your network. Publish vLLM on an address it can call.

  3. 03

    Pool it with a hosted provider

    Add Llama 3.3 70B on Together AI as a second connection. Your GPUs answer first, and agents ask for llama-3.3-70b.

    $ dome models pool create llama-3.3-70b \
    --failover-max all --gateway prod-gateway
     
    $ dome models pool member add llama-3.3-70b llama-vllm --priority 0
    $ dome models pool member add llama-3.3-70b llama-together --priority 1

Commands and rules tested against a Dome workspace on October 1, 2026. For anything about vLLM itself, see vLLM's documentation.

Rules

This agent stays on your GPUs

Applied to one agent with --agent, this refuses any connection not tagged hosting: self. The contracts agent never reaches the hosted fallback.

forbid (principal, action == Dome::Action::"llm:invoke", resource is Dome::LLMModel)
unless { resource has hosting && resource.hosting == "self" };

Agent workflow

Bringing it together

Connecting vLLM to registered agents, tools, and identity in Dome completes a governed agent application.

Dome

Control point

Gateway

  • Rules
  • Guards
  • Quotas

Every call decided and audited

FAQ

Common questions

How does Dome connect to vLLM?

Through vLLM's OpenAI-compatible server, connected as an openai_compatible provider. Give Dome a key only if vLLM was started with --api-key.

Can a Dome-hosted gateway reach my vLLM server?

Only on a URL it can reach. A server on a private network needs a reachable endpoint first.

Which model name do I use?

The name vLLM serves, which defaults to the Hugging Face repo id, such as meta-llama/Llama-3.3-70B-Instruct.

Next steps

Talk with our FDE team

Our forward deployed engineers work with your platform team to get your agents into production and under control: the first one governed on your own systems, and a pattern your teams can repeat for every agent after it.