Dome Systems

Llama and Dome

Run Llama from three providers under one name.

Six of Dome's providers serve Llama, so no one vendor has to hold your support queue. Dome pools Groq, Together AI and SambaNova under llama-3.3-70b, keeps the model to the support team, and puts a daily ceiling on tokens.

PeopleAgentsDomePool: llama-3.3-70bCallersPeopleAgentstriage-agentAgentsintake-routerAgentsdocs-writerGatewaysprod-gatewayGroqPriority 0Together AIPriority 1SambaNovaPriority 2RulesGuardsAuditsaudit-trail
triage-agent→llama-3.3-70b· as r.alvarezAllowed

How Dome helps

Dome provides model brokering and routing for Llama

No single vendor

Groq, Together AI and SambaNova each serve Llama 3.3 70B under their own id. The pool gives it one name, and a failed call moves to the next provider.

Llama for the support team

Tag Llama connections with a family attribute. A rule refuses them unless the person the agent acts for is in support.

A daily token ceiling

20M tokens a day on the pool, summed across Groq, Together and SambaNova. The count resets daily.

Get started

Llama with failover in five steps

Three providers, three model ids, one pool name, then the support rule and a token cap.

  1. 01

    Connect Groq

    Groq names it llama-3.3-70b-versatile. Its key stays in Dome; support agents hold Dome keys only.

    $ dome models add llama-groq \
    --provider groq \
    --model llama-3.3-70b-versatile \
    --api-key "$GROQ_API_KEY" \
    --attributes '{"family":"llama"}' \
    --gateway prod-gateway
  2. 02

    Connect Together AI and SambaNova

    Same model, each provider's own id. Tag them the same way.

    $ dome models add llama-together \
    --provider together \
    --model meta-llama/Llama-3.3-70B-Instruct-Turbo \
    --api-key "$TOGETHER_API_KEY" \
    --attributes '{"family":"llama"}' \
    --gateway prod-gateway
     
    $ dome models add llama-sambanova \
    --provider sambanova \
    --model Meta-Llama-3.3-70B-Instruct \
    --api-key "$SAMBANOVA_API_KEY" \
    --attributes '{"family":"llama"}' \
    --gateway prod-gateway
  3. 03

    Pool them for failover

    Groq leads, Together and SambaNova follow. The provider ids stay inside the pool.

    $ dome models pool create llama-3.3-70b \
    --failover-max all --gateway prod-gateway
     
    $ dome models pool member add llama-3.3-70b llama-groq --priority 0
    $ dome models pool member add llama-3.3-70b llama-together --priority 1
    $ dome models pool member add llama-3.3-70b llama-sambanova --priority 2
  4. 04

    Point your agent at the Gateway

    All three speak the OpenAI API, so the support agent's client just changes its base URL and key.

    from openai import OpenAI
     
    client = OpenAI(base_url=f"{GATEWAY_URL}/v1", api_key=DOME_AGENT_KEY)
    client.chat.completions.create(
    model="llama-3.3-70b",
    messages=[{"role": "user", "content": "Classify this ticket: billing, bug or how-to."}],
    )
  5. 05

    Cap the tokens

    Tokens are counted at the pool, not per provider.

    $ dome quotas set --subject pool --pool llama-3.3-70b \
    --unit tokens --limit 20000000 --window daily

Commands and rules tested against a Dome workspace on October 1, 2026. For anything about Llama itself, see Meta's documentation.

Rules

Llama for the support team

Put it on an agent with --agent and Llama answers only when the person behind the call is in support. Everything else that agent uses is left alone.

forbid (principal, action == Dome::Action::"llm:invoke", resource is Dome::LLMModel)
when { resource has family && resource.family == "llama" }
unless { principal has act_as && principal.act_as.groups.contains("support") };

Try it

One call, two outcomes

Switch the caller or the argument and watch the same call decide differently. Every decision lands in audit.

Acting for

agent triage-agent · acting as r.alvarez
llm:invoke(model: "llama-3.3-70b")
  1. Agenttriage-agent is registered and active
  2. Callerr.alvarez verified, groups: support
  3. RuleLlama is open to the support group
  4. Quota6.2M of 20M tokens today
DecisionAllowed

Agent workflow

Bringing it together

Connecting Llama to registered agents, tools, and identity in Dome completes a governed agent application.

Dome

Control point

Gateway

  • Rules
  • Guards
  • Quotas

Every call decided and audited

FAQ

Common questions

Which providers serve Llama through Dome?

Groq, Together AI, SambaNova, DeepInfra, NVIDIA NIM and Amazon Bedrock all serve Llama 3.3 70B. Dome passes each provider's model id through unchanged.

Why does each provider use a different model id?

Providers name open models their own way, such as llama-3.3-70b-versatile on Groq. Your agents send the pool's name instead.

What happens when a new Llama ships?

Connect the new id on each provider and create a pool named for it. Agents move when they send the new name.

Next steps

Talk with our FDE team

Our forward deployed engineers work with your platform team to get your agents into production and under control: the first one governed on your own systems, and a pattern your teams can repeat for every agent after it.