Claude Haiku and Dome
Run Claude Haiku 4.5 as the fast default for high-volume agents.
Haiku is the volume model. Classifiers, routers and triage agents call it thousands of times a day, so an outage is a queue and a runaway loop is a bill. Dome spreads it across two providers, caps each agent's tokens, and makes it the only model contractors reach.
How Dome helps
Dome provides model brokering and routing for Claude Haiku
Volume without a single point of failure
Pool Anthropic and Bedrock under claude-haiku-4-5. A provider outage costs one retry, not a backlog.
Haiku as the floor for some callers
Tag Haiku connections as small. A rule keeps contractors on them and refuses everything larger.
A token cap per agent
Cap the batch classifier at 5M tokens a day. A runaway job stops at its own ceiling.
Get started
Haiku with failover in five steps
Two connections tagged small, a pool, then a token cap on the batch classifier.
01
Connect Anthropic
tier: small marks it as the model contractors may use. Triage agents call with their own keys; the Anthropic key stays put.
$ dome models add haiku-45-anthropic \--provider anthropic \--model claude-haiku-4-5 \--api-key "$ANTHROPIC_API_KEY" \--attributes '{"tier":"small"}' \--gateway prod-gateway02
Connect Amazon Bedrock
Bedrock's Haiku id is us.anthropic.claude-haiku-4-5-20251001-v1:0. AWS keys go in through the dashboard, with tier: small on this connection too.
03
Pool them for failover
A failed classification retries on Bedrock instead of joining a backlog.
$ dome models pool create claude-haiku-4-5 \--failover-max all --gateway prod-gateway$ dome models pool member add claude-haiku-4-5 haiku-45-anthropic --priority 0$ dome models pool member add claude-haiku-4-5 haiku-45-bedrock --priority 104
Point your agent at the Gateway
A triage agent written against the OpenAI SDK only needs the Gateway URL and its Dome key.
from openai import OpenAIclient = OpenAI(base_url=f"{GATEWAY_URL}/v1", api_key=DOME_AGENT_KEY)client.chat.completions.create(model="claude-haiku-4-5",messages=[{"role": "user", "content": "Classify this ticket: billing, bug or other."}],)05
Cap tokens per agent
An agent quota counts every model call that agent makes, through any pool.
$ dome quotas set --subject agent --agent batch-classifier \--dimension llm --unit tokens --limit 5000000 --window daily
Commands and rules tested against a Dome workspace on October 1, 2026. For anything about Claude Haiku itself, see Anthropic's documentation.
Rules
Contractors get Haiku and nothing larger
When the person an agent acts for is in the contractors group, only connections tagged small are allowed. The group arrives in their identity token.
forbid (principal, action == Dome::Action::"llm:invoke", resource is Dome::LLMModel)when { principal has act_as && principal.act_as.groups.contains("contractors")}unless { resource has tier && resource.tier == "small" };Try it
One call, two outcomes
Switch the caller or the argument and watch the same call decide differently. Every decision lands in audit.
Model
- Agentintake-agent is registered and active
- Callercontractor.lee verified, groups: contractors
- RuleConnection is tagged tier: small
- Quota1.2M of 5M tokens today
Agent workflow
Bringing it together
Connecting Claude Haiku to registered agents, tools, and identity in Dome completes a governed agent application.
Acting for
Control point
Gateway
- Rules
- Guards
- Quotas
Every call decided and audited
Models
FAQ
Common questions
Which providers serve Claude Haiku through Dome?
Anthropic and Amazon Bedrock. Running Claude through Vertex AI isn't supported by Dome.
Can an agent be limited to Haiku?
Yes. Grant the agent only the claude-haiku-4-5 pool, and it can't route anywhere else.
How do I stop a batch job from burning tokens?
Set a token quota on the agent. The job is refused at its ceiling while other agents keep working.
Explore
More of what Dome works with
Provider
OpenAI
The OpenAI API behind the Model Broker: the key held in Dome, every call authorized, metered and audited.
Read moreMCP server
GitHub
The GitHub MCP server behind the Tool Gateway, with rules that decide each call on its owner and repository.
Read moreIdentity
Okta
Okta tokens verified on every agent call, so rules and audit name the person each agent acted for.
Read moreRuntime
LangGraph
LangGraph agents with their model calls on the Model Broker and their MCP tools on the Tool Gateway.
Read moreClient
Claude Code
Claude Code on a Dome Gateway with per-developer sign-in, rules on every tool call, and audit by name.
Read moreAgent service
TinyFish
TinyFish's web agents behind the Tool Gateway, with rules that decide each run on the site it targets.
Read moreNext steps
Talk with our FDE team
Our forward deployed engineers work with your platform team to get your agents into production and under control: the first one governed on your own systems, and a pattern your teams can repeat for every agent after it.
No card required to start. Register your first agent in minutes.
AI models for enterprise agents: failover, rules and quotas
Every model your agents call goes through the Model Broker. Pool providers for failover, decide who may call each model, and cap what it costs.
See them all