Gemma and Dome
Run Gemma 4 from two providers under one name.
Gemma is Google's open-weight model, small enough to be the sensible default for people who don't need a frontier one. Dome serves it from DeepInfra with NVIDIA NIM behind, keeps contractors on it, and counts calls per day.
How Dome helps
Dome provides model brokering and routing for Gemma
DeepInfra first, NVIDIA next
DeepInfra serves google/gemma-4-31B-it and NVIDIA NIM serves google/gemma-4-31b-it. The pool gives them one name.
Contractors use Gemma
Tag Gemma connections with a family attribute. When the person behind a call is a contractor, a rule refuses every other model.
Calls per day
The pool allows 50,000 calls a day. The count is calls, not tokens: each request counts once, however long.
Get started
Gemma with failover in five steps
DeepInfra and NVIDIA in one pool, a contractor rule, then a daily call cap.
01
Connect DeepInfra
DeepInfra's id is google/gemma-4-31B-it. Dome keeps the DeepInfra key out of every agent's hands.
$ dome models add gemma-deepinfra \--provider deepinfra \--model google/gemma-4-31B-it \--api-key "$DEEPINFRA_API_KEY" \--attributes '{"family":"gemma"}' \--gateway prod-gateway02
Connect NVIDIA NIM
The provider id is nvidia. NVIDIA writes the model id in lower case.
$ dome models add gemma-nvidia \--provider nvidia \--model google/gemma-4-31b-it \--api-key "$NVIDIA_API_KEY" \--attributes '{"family":"gemma"}' \--gateway prod-gateway03
Pool them for failover
The two ids differ only in case. The pool hides both behind gemma-4-31b.
$ dome models pool create gemma-4-31b \--failover-max all --gateway prod-gateway$ dome models pool member add gemma-4-31b gemma-deepinfra --priority 0$ dome models pool member add gemma-4-31b gemma-nvidia --priority 104
Point your agent at the Gateway
Contractor-facing agents use the OpenAI SDK pointed at the Gateway, with their own Dome key.
from openai import OpenAIclient = OpenAI(base_url=f"{GATEWAY_URL}/v1", api_key=DOME_AGENT_KEY)client.chat.completions.create(model="gemma-4-31b",messages=[{"role": "user", "content": "Find the VPN setup steps in the knowledge base."}],)05
Cap daily calls
One call counts once, on whichever host served it.
$ dome quotas set --subject pool --pool gemma-4-31b \--unit calls --limit 50000 --window daily
Commands and rules tested against a Dome workspace on October 1, 2026. For anything about Gemma itself, see Google's documentation.
Rules
Contractors use Gemma
When the person behind a call is in contractors, any connection not tagged Gemma is refused. Employees using the same agent are untouched.
forbid (principal, action == Dome::Action::"llm:invoke", resource is Dome::LLMModel)when { principal has act_as && principal.act_as.groups.contains("contractors") }unless { resource has family && resource.family == "gemma" };Try it
One call, two outcomes
Switch the caller or the argument and watch the same call decide differently. Every decision lands in audit.
Model
- Agenthelpdesk-agent is registered and active
- Callerj.smith verified, groups: contractors
- RuleConnection is tagged family: gemma
- Quota18,400 of 50,000 calls today
Agent workflow
Bringing it together
Connecting Gemma to registered agents, tools, and identity in Dome completes a governed agent application.
Acting for
Control point
Gateway
- Rules
- Guards
- Quotas
Every call decided and audited
Models
FAQ
Common questions
Which providers serve Gemma through Dome?
DeepInfra and NVIDIA NIM serve Gemma 4 31B, and SambaNova lists it in preview. Dome passes each provider's model id through unchanged.
Can I use Gemma on Amazon Bedrock?
Not yet. Bedrock serves Gemma 4 only on its bedrock-mantle endpoint, and Dome's Bedrock connections use the Converse API.
What happens when a new Gemma ships?
Add the new Gemma on DeepInfra and NVIDIA, swap the members, and contractors move with it.
Explore
More of what Dome works with
Provider
Anthropic
The Claude API behind the Model Broker: the key held in Dome, every call authorized, metered and audited.
Read moreMCP server
GitHub
The GitHub MCP server behind the Tool Gateway, with rules that decide each call on its owner and repository.
Read moreIdentity
Okta
Okta tokens verified on every agent call, so rules and audit name the person each agent acted for.
Read moreRuntime
LangGraph
LangGraph agents with their model calls on the Model Broker and their MCP tools on the Tool Gateway.
Read moreClient
Claude Code
Claude Code on a Dome Gateway with per-developer sign-in, rules on every tool call, and audit by name.
Read moreAgent service
TinyFish
TinyFish's web agents behind the Tool Gateway, with rules that decide each run on the site it targets.
Read moreNext steps
Talk with our FDE team
Our forward deployed engineers work with your platform team to get your agents into production and under control: the first one governed on your own systems, and a pattern your teams can repeat for every agent after it.
No card required to start. Register your first agent in minutes.
AI models for enterprise agents: failover, rules and quotas
Every model your agents call goes through the Model Broker. Pool providers for failover, decide who may call each model, and cap what it costs.
See them all