Cerebras API and Dome
Fast inference on Cerebras, with every call governed.
Cerebras is the provider you reach for when a person is waiting on the answer. Connect it to Dome once and its capacity goes to those agents: a rule keeps batch work off it, and Fireworks AI takes gpt-oss-120b calls when Cerebras fails.
How Dome helps
Dome provides model brokering and routing for Cerebras
One Cerebras key, many agents
Dome holds the Cerebras key and sends it as a bearer token. Each voice and routing agent calls with its own Dome key, and audit names it.
Same model, second provider
Fireworks AI serves gpt-oss-120b too. Pool the two and a failed call on Cerebras goes to Fireworks.
Fast capacity where it counts
Tag Cerebras connections as the fast lane. A rule keeps batch agents off it.
Get started
Connect Cerebras in three steps
Add Cerebras, back it with Fireworks AI, then point agents at the Gateway.
01
Add the connection
The provider id is cerebras. Dome uses https://api.cerebras.ai/v1 and Cerebras' own model ids.
$ dome models add gpt-oss-cerebras \--provider cerebras \--model gpt-oss-120b \--api-key "$CEREBRAS_API_KEY" \--attributes '{"lane":"fast"}' \--gateway prod-gateway02
Pool it with Fireworks AI
Add accounts/fireworks/models/gpt-oss-120b as a second connection. Members are tried in priority order.
$ dome models pool create gpt-oss-120b \--failover-max all --gateway prod-gateway$ dome models pool member add gpt-oss-120b gpt-oss-cerebras --priority 0$ dome models pool member add gpt-oss-120b gpt-oss-fireworks --priority 103
Point your agent at the Gateway
Any OpenAI-compatible client works. The agent's Dome key replaces the Cerebras key.
from openai import OpenAIclient = OpenAI(base_url=f"{GATEWAY_URL}/v1", api_key=DOME_AGENT_KEY)client.chat.completions.create(model="gpt-oss-120b",messages=[{"role": "user", "content": "Classify this caller's intent."}],)
Commands and rules tested against a Dome workspace on October 1, 2026. For anything about Cerebras itself, see Cerebras's documentation.
Rules
Keep batch agents off the fast lane
Applied to a batch agent with --agent, this refuses any connection tagged lane: fast. Cerebras capacity goes to the agents a person is waiting on.
forbid (principal, action == Dome::Action::"llm:invoke", resource is Dome::LLMModel)when { resource has lane && resource.lane == "fast" };Agent workflow
Bringing it together
Connecting Cerebras to registered agents, tools, and identity in Dome completes a governed agent application.
Acting for
Control point
Gateway
- Rules
- Guards
- Quotas
Every call decided and audited
Models
FAQ
Common questions
How does Dome connect to Cerebras?
Through Cerebras' OpenAI-compatible API at https://api.cerebras.ai/v1, with your key sent as a bearer token. The key stays in Dome's vault.
Can I keep Cerebras for agents a person is waiting on?
Yes. Tag the Cerebras connections lane: fast and apply the rule above to each batch agent. Batch calls are refused before they leave Dome. Interactive agents are untouched.
What happens when Cerebras is unavailable?
The gpt-oss-120b pool sends the call to Fireworks AI, the next member by priority. The agent still asks for gpt-oss-120b, and audit records which provider answered.
Explore
More of what Dome works with
Model
Claude Fable
Fable 5.1 from Anthropic and Amazon Bedrock in one failover pool, open to one group and capped by quota.
Read moreMCP server
GitHub
The GitHub MCP server behind the Tool Gateway, with rules that decide each call on its owner and repository.
Read moreIdentity
Okta
Okta tokens verified on every agent call, so rules and audit name the person each agent acted for.
Read moreRuntime
LangGraph
LangGraph agents with their model calls on the Model Broker and their MCP tools on the Tool Gateway.
Read moreClient
Claude Code
Claude Code on a Dome Gateway with per-developer sign-in, rules on every tool call, and audit by name.
Read moreAgent service
TinyFish
TinyFish's web agents behind the Tool Gateway, with rules that decide each run on the site it targets.
Read moreNext steps
Talk with our FDE team
Our forward deployed engineers work with your platform team to get your agents into production and under control: the first one governed on your own systems, and a pattern your teams can repeat for every agent after it.
No card required to start. Register your first agent in minutes.
LLM providers for AI agents: one governed path to every model
Connect a provider once. Its key stays in Dome, and every agent call to it is authorized, metered and audited.
See them all