Llama and Dome
Run Llama from three providers under one name.
Six of Dome's providers serve Llama, so no one vendor has to hold your support queue. Dome pools Groq, Together AI and SambaNova under llama-3.3-70b, keeps the model to the support team, and puts a daily ceiling on tokens.
How Dome helps
Dome provides model brokering and routing for Llama
No single vendor
Groq, Together AI and SambaNova each serve Llama 3.3 70B under their own id. The pool gives it one name, and a failed call moves to the next provider.
Llama for the support team
Tag Llama connections with a family attribute. A rule refuses them unless the person the agent acts for is in support.
A daily token ceiling
20M tokens a day on the pool, summed across Groq, Together and SambaNova. The count resets daily.
Get started
Llama with failover in five steps
Three providers, three model ids, one pool name, then the support rule and a token cap.
01
Connect Groq
Groq names it llama-3.3-70b-versatile. Its key stays in Dome; support agents hold Dome keys only.
$ dome models add llama-groq \--provider groq \--model llama-3.3-70b-versatile \--api-key "$GROQ_API_KEY" \--attributes '{"family":"llama"}' \--gateway prod-gateway02
Connect Together AI and SambaNova
Same model, each provider's own id. Tag them the same way.
$ dome models add llama-together \--provider together \--model meta-llama/Llama-3.3-70B-Instruct-Turbo \--api-key "$TOGETHER_API_KEY" \--attributes '{"family":"llama"}' \--gateway prod-gateway$ dome models add llama-sambanova \--provider sambanova \--model Meta-Llama-3.3-70B-Instruct \--api-key "$SAMBANOVA_API_KEY" \--attributes '{"family":"llama"}' \--gateway prod-gateway03
Pool them for failover
Groq leads, Together and SambaNova follow. The provider ids stay inside the pool.
$ dome models pool create llama-3.3-70b \--failover-max all --gateway prod-gateway$ dome models pool member add llama-3.3-70b llama-groq --priority 0$ dome models pool member add llama-3.3-70b llama-together --priority 1$ dome models pool member add llama-3.3-70b llama-sambanova --priority 204
Point your agent at the Gateway
All three speak the OpenAI API, so the support agent's client just changes its base URL and key.
from openai import OpenAIclient = OpenAI(base_url=f"{GATEWAY_URL}/v1", api_key=DOME_AGENT_KEY)client.chat.completions.create(model="llama-3.3-70b",messages=[{"role": "user", "content": "Classify this ticket: billing, bug or how-to."}],)05
Cap the tokens
Tokens are counted at the pool, not per provider.
$ dome quotas set --subject pool --pool llama-3.3-70b \--unit tokens --limit 20000000 --window daily
Commands and rules tested against a Dome workspace on October 1, 2026. For anything about Llama itself, see Meta's documentation.
Rules
Llama for the support team
Put it on an agent with --agent and Llama answers only when the person behind the call is in support. Everything else that agent uses is left alone.
forbid (principal, action == Dome::Action::"llm:invoke", resource is Dome::LLMModel)when { resource has family && resource.family == "llama" }unless { principal has act_as && principal.act_as.groups.contains("support") };Try it
One call, two outcomes
Switch the caller or the argument and watch the same call decide differently. Every decision lands in audit.
Acting for
- Agenttriage-agent is registered and active
- Callerr.alvarez verified, groups: support
- RuleLlama is open to the support group
- Quota6.2M of 20M tokens today
Agent workflow
Bringing it together
Connecting Llama to registered agents, tools, and identity in Dome completes a governed agent application.
Acting for
Control point
Gateway
- Rules
- Guards
- Quotas
Every call decided and audited
Models
FAQ
Common questions
Which providers serve Llama through Dome?
Groq, Together AI, SambaNova, DeepInfra, NVIDIA NIM and Amazon Bedrock all serve Llama 3.3 70B. Dome passes each provider's model id through unchanged.
Why does each provider use a different model id?
Providers name open models their own way, such as llama-3.3-70b-versatile on Groq. Your agents send the pool's name instead.
What happens when a new Llama ships?
Connect the new id on each provider and create a pool named for it. Agents move when they send the new name.
Explore
More of what Dome works with
Provider
Anthropic
The Claude API behind the Model Broker: the key held in Dome, every call authorized, metered and audited.
Read moreMCP server
GitHub
The GitHub MCP server behind the Tool Gateway, with rules that decide each call on its owner and repository.
Read moreIdentity
Microsoft Entra ID
Entra access tokens verified on every agent call, so rules read app roles and audit names the person.
Read moreRuntime
LangGraph
LangGraph agents with their model calls on the Model Broker and their MCP tools on the Tool Gateway.
Read moreClient
Claude Code
Claude Code on a Dome Gateway with per-developer sign-in, rules on every tool call, and audit by name.
Read moreAgent service
TinyFish
TinyFish's web agents behind the Tool Gateway, with rules that decide each run on the site it targets.
Read moreNext steps
Talk with our FDE team
Our forward deployed engineers work with your platform team to get your agents into production and under control: the first one governed on your own systems, and a pattern your teams can repeat for every agent after it.
No card required to start. Register your first agent in minutes.
AI models for enterprise agents: failover, rules and quotas
Every model your agents call goes through the Model Broker. Pool providers for failover, decide who may call each model, and cap what it costs.
See them all