vLLM and Dome
Your vLLM servers behind one Gateway, with every call governed.
vLLM puts open models on your own GPUs behind an OpenAI-compatible server. Connect it to Dome, let Together AI take the overflow, and a rule decides which agents may ever leave your hardware. The Gateway needs a route to the server.
How Dome helps
Dome provides model brokering and routing for vLLM
Any model vLLM serves
Point Dome's openai_compatible provider at the server's /v1 address. The served model name is the model id.
Overflow to a hosted provider
Pool vLLM with Together AI serving the same Llama. Your GPUs answer first and Together takes what they can't.
Some agents never leave your GPUs
Tag connections by where they run. A rule keeps an agent on self-hosted models only.
Get started
Connect vLLM in three steps
Connect your server, give the Gateway a route to it, and add a hosted fallback.
01
Add the connection
Use the provider id openai_compatible and the server's /v1 URL. If you start vLLM with --api-key, give Dome the same key.
$ dome models add llama-vllm \--provider openai_compatible \--endpoint https://vllm.internal.example.com/v1 \--model meta-llama/Llama-3.3-70B-Instruct \--api-key "$VLLM_API_KEY" \--attributes '{"hosting":"self"}' \--gateway prod-gateway02
Give the Gateway a route
The hosted gateway can't reach inside your network. Publish vLLM on an address it can call.
03
Pool it with a hosted provider
Add Llama 3.3 70B on Together AI as a second connection. Your GPUs answer first, and agents ask for llama-3.3-70b.
$ dome models pool create llama-3.3-70b \--failover-max all --gateway prod-gateway$ dome models pool member add llama-3.3-70b llama-vllm --priority 0$ dome models pool member add llama-3.3-70b llama-together --priority 1
Commands and rules tested against a Dome workspace on October 1, 2026. For anything about vLLM itself, see vLLM's documentation.
Rules
This agent stays on your GPUs
Applied to one agent with --agent, this refuses any connection not tagged hosting: self. The contracts agent never reaches the hosted fallback.
forbid (principal, action == Dome::Action::"llm:invoke", resource is Dome::LLMModel)unless { resource has hosting && resource.hosting == "self" };Agent workflow
Bringing it together
Connecting vLLM to registered agents, tools, and identity in Dome completes a governed agent application.
Acting for
Control point
Gateway
- Rules
- Guards
- Quotas
Every call decided and audited
Models
FAQ
Common questions
How does Dome connect to vLLM?
Through vLLM's OpenAI-compatible server, connected as an openai_compatible provider. Give Dome a key only if vLLM was started with --api-key.
Can a Dome-hosted gateway reach my vLLM server?
Only on a URL it can reach. A server on a private network needs a reachable endpoint first.
Which model name do I use?
The name vLLM serves, which defaults to the Hugging Face repo id, such as meta-llama/Llama-3.3-70B-Instruct.
Explore
More of what Dome works with
Model
Claude Fable
Fable 5.1 from Anthropic and Amazon Bedrock in one failover pool, open to one group and capped by quota.
Read moreMCP server
GitHub
The GitHub MCP server behind the Tool Gateway, with rules that decide each call on its owner and repository.
Read moreIdentity
Okta
Okta tokens verified on every agent call, so rules and audit name the person each agent acted for.
Read moreRuntime
OpenAI Agents SDK
OpenAI Agents SDK agents with their models on the Model Broker and their MCP tools on the Tool Gateway.
Read moreClient
Claude Code
Claude Code on a Dome Gateway with per-developer sign-in, rules on every tool call, and audit by name.
Read moreAgent service
TinyFish
TinyFish's web agents behind the Tool Gateway, with rules that decide each run on the site it targets.
Read moreNext steps
Talk with our FDE team
Our forward deployed engineers work with your platform team to get your agents into production and under control: the first one governed on your own systems, and a pattern your teams can repeat for every agent after it.
No card required to start. Register your first agent in minutes.
LLM providers for AI agents: one governed path to every model
Connect a provider once. Its key stays in Dome, and every agent call to it is authorized, metered and audited.
See them all