Dome Systems

Gemini and Dome

Run Gemini on Vertex AI, with Pro kept for the work that needs it.

Dome calls Gemini through Vertex AI with your service account. Pool Gemini 3.8 Flash with an older Flash behind it, and agents keep working when one model is unavailable. Rules keep Gemini 3.1 Pro to the people who need it.

PeopleAgentsDomeVertex AI modelsCallersPeopleAgentsresearch-agentAgentssupport-agentAgentsdoc-extractorGatewaysprod-gatewayGemini 3.8 FlashPriority 0Gemini 2.5 FlashPriority 1Gemini 3.1 Progemini-3.1-proRulesGuardsAuditsaudit-trail
doc-extractor→gemini-flash· scheduledAllowed

How Dome helps

Dome provides model brokering and routing for Gemini

A fallback model behind Flash

Pool Gemini 3.8 Flash with Gemini 2.5 Flash behind it. When the first fails, the call goes to the second.

Pro for the teams that need it

Tag the Pro connection with its tier. A rule refuses it unless the person the agent acts for is in ml-research.

A ceiling on spend

Cap Pro at $400 of provider spend a month. Once it's spent, Pro calls are refused until the window resets.

Get started

Gemini on Vertex AI in five steps

Connect each model with your Google Cloud project and service account, pool the Flash models, and point your agents at the Gateway.

  1. 01

    Connect Gemini 3.8 Flash

    Pass the service account's JSON key as the credential. Dome stores it in its vault and exchanges it for Google access tokens at call time.

    $ dome models add gemini-38-flash \
    --provider google \
    --model gemini-3.8-flash \
    --provider-config '{"project":"acme-agents","location":"us-central1"}' \
    --api-key "$(cat vertex-sa.json)" \
    --gateway prod-gateway
  2. 02

    Connect the fallback and Pro

    One connection per model. Tag Pro so a rule can name it.

    $ dome models add gemini-25-flash --provider google --model gemini-2.5-flash \
    --provider-config '{"project":"acme-agents","location":"us-central1"}' \
    --api-key "$(cat vertex-sa.json)" --gateway prod-gateway
     
    $ dome models add gemini-pro --provider google --model gemini-3.1-pro \
    --provider-config '{"project":"acme-agents","location":"us-central1"}' \
    --api-key "$(cat vertex-sa.json)" --attributes '{"tier":"pro"}' \
    --gateway prod-gateway
  3. 03

    Pool the Flash models

    Members are tried in priority order. Agents send the pool's name.

    $ dome models pool create gemini-flash \
    --failover-max all --gateway prod-gateway
     
    $ dome models pool member add gemini-flash gemini-38-flash --priority 0
    $ dome models pool member add gemini-flash gemini-25-flash --priority 1
  4. 04

    Point your agent at the Gateway

    Any OpenAI-compatible client works. Dome translates the request to Gemini's format.

    from openai import OpenAI
     
    client = OpenAI(base_url=f"{GATEWAY_URL}/v1", api_key=DOME_AGENT_KEY)
    client.chat.completions.create(
    model="gemini-flash",
    messages=[{"role": "user", "content": "Extract the invoice fields."}],
    )
  5. 05

    Cap the spend

    A model quota counts every call to that connection, through pools or directly.

    $ dome quotas set --subject model --model gemini-pro \
    --unit provider_usd --limit 400 --window monthly

Commands and rules tested against a Dome workspace on October 1, 2026. For anything about Gemini itself, see Google's documentation.

Rules

Gemini Pro for research, refused for everyone else

Any connection tagged pro is refused unless the person the agent acts for is in ml-research. The group arrives in their identity token.

forbid (principal, action == Dome::Action::"llm:invoke", resource is Dome::LLMModel)
when { resource has tier && resource.tier == "pro" }
unless {
principal has act_as &&
principal.act_as.groups.contains("ml-research")
};

Try it

One call, two outcomes

Switch the caller or the argument and watch the same call decide differently. Every decision lands in audit.

Acting for

agent research-agent · acting as p.nakamura
llm:invoke(model: "gemini-pro")
  1. Agentresearch-agent is registered and active
  2. Callerp.nakamura verified, groups: ml-research
  3. RulePro is open to ml-research
  4. Quota$150 of $400 this month
DecisionAllowed

Agent workflow

Bringing it together

Connecting Gemini to registered agents, tools, and identity in Dome completes a governed agent application.

Dome

Control point

Gateway

  • Rules
  • Guards
  • Quotas

Every call decided and audited

FAQ

Common questions

Does Dome call Gemini through Vertex AI or the Gemini API?

Vertex AI, with a Google Cloud service account. Gemini API keys from Google AI Studio are not supported.

Which providers serve Gemini through Dome?

Google, through Vertex AI. A pool can fall back between Gemini models in the same project.

Next steps

Talk with our FDE team

Our forward deployed engineers work with your platform team to get your agents into production and under control: the first one governed on your own systems, and a pattern your teams can repeat for every agent after it.