Dome Systems

Cerebras API and Dome

Fast inference on Cerebras, with every call governed.

Cerebras is the provider you reach for when a person is waiting on the answer. Connect it to Dome once and its capacity goes to those agents: a rule keeps batch work off it, and Fireworks AI takes gpt-oss-120b calls when Cerebras fails.

PeopleAgentsDomeCerebras modelsCallersPeopleAgentsvoice-agentAgentsintent-routerAgentsbatch-taggerGatewaysprod-gatewaygpt-oss-120bgpt-oss-120bQwen 3.8 27Bqwen-3.8-27bFireworks AIgpt-oss failover, priority 1RulesGuardsAuditsaudit-trail
voice-agent→gpt-oss-120b· as c.diazAllowed

How Dome helps

Dome provides model brokering and routing for Cerebras

One Cerebras key, many agents

Dome holds the Cerebras key and sends it as a bearer token. Each voice and routing agent calls with its own Dome key, and audit names it.

Same model, second provider

Fireworks AI serves gpt-oss-120b too. Pool the two and a failed call on Cerebras goes to Fireworks.

Fast capacity where it counts

Tag Cerebras connections as the fast lane. A rule keeps batch agents off it.

Get started

Connect Cerebras in three steps

Add Cerebras, back it with Fireworks AI, then point agents at the Gateway.

  1. 01

    Add the connection

    The provider id is cerebras. Dome uses https://api.cerebras.ai/v1 and Cerebras' own model ids.

    $ dome models add gpt-oss-cerebras \
    --provider cerebras \
    --model gpt-oss-120b \
    --api-key "$CEREBRAS_API_KEY" \
    --attributes '{"lane":"fast"}' \
    --gateway prod-gateway
  2. 02

    Pool it with Fireworks AI

    Add accounts/fireworks/models/gpt-oss-120b as a second connection. Members are tried in priority order.

    $ dome models pool create gpt-oss-120b \
    --failover-max all --gateway prod-gateway
     
    $ dome models pool member add gpt-oss-120b gpt-oss-cerebras --priority 0
    $ dome models pool member add gpt-oss-120b gpt-oss-fireworks --priority 1
  3. 03

    Point your agent at the Gateway

    Any OpenAI-compatible client works. The agent's Dome key replaces the Cerebras key.

    from openai import OpenAI
     
    client = OpenAI(base_url=f"{GATEWAY_URL}/v1", api_key=DOME_AGENT_KEY)
    client.chat.completions.create(
    model="gpt-oss-120b",
    messages=[{"role": "user", "content": "Classify this caller's intent."}],
    )

Commands and rules tested against a Dome workspace on October 1, 2026. For anything about Cerebras itself, see Cerebras's documentation.

Rules

Keep batch agents off the fast lane

Applied to a batch agent with --agent, this refuses any connection tagged lane: fast. Cerebras capacity goes to the agents a person is waiting on.

forbid (principal, action == Dome::Action::"llm:invoke", resource is Dome::LLMModel)
when { resource has lane && resource.lane == "fast" };

Agent workflow

Bringing it together

Connecting Cerebras to registered agents, tools, and identity in Dome completes a governed agent application.

Dome

Control point

Gateway

  • Rules
  • Guards
  • Quotas

Every call decided and audited

FAQ

Common questions

How does Dome connect to Cerebras?

Through Cerebras' OpenAI-compatible API at https://api.cerebras.ai/v1, with your key sent as a bearer token. The key stays in Dome's vault.

Can I keep Cerebras for agents a person is waiting on?

Yes. Tag the Cerebras connections lane: fast and apply the rule above to each batch agent. Batch calls are refused before they leave Dome. Interactive agents are untouched.

What happens when Cerebras is unavailable?

The gpt-oss-120b pool sends the call to Fireworks AI, the next member by priority. The agent still asks for gpt-oss-120b, and audit records which provider answered.

Next steps

Talk with our FDE team

Our forward deployed engineers work with your platform team to get your agents into production and under control: the first one governed on your own systems, and a pattern your teams can repeat for every agent after it.