Open Computer / Products / Inferno Router

Inferno Router

Inferno Router is the inference gateway in front of your models. Call compatible Claude, GPT, Gemini, Grok, and open models through one endpoint. The router handles provider choice, failover, usage logs, and Fusion routing so applications do not hard-wire a single vendor — and so routine work does not hit frontier prices.

One API. Every model.

Call Claude, GPT, Gemini, Grok, and open models through one OpenAI-compatible endpoint. Point existing SDKs at Inferno Router. Swap a model without rewriting the client.

Stop paying frontier prices for easy work

Fusion routing sends routine steps to self-hosted open models and reserves frontier models for the hard parts. Prompt caching cuts repeated spend.

If a provider dies, the call does not

When a provider or region is unhealthy, Inferno Router fails over to the next healthy target. Clients keep the same URL.

Run your own weights

Host open or fine-tuned models on your capacity. Frontier models stay available for the steps that need them. Data can stay in your environment.

See the bill before it surprises you

Usage logs, latency, and per-key spend in one place. Quotas and rate limits stay isolated per team so one noisy workload cannot eat another budget.

Repeated prompts should not cost full price

Prompt caching cuts latency and token spend on work you already ran. Purge the cache when you need a fresh inference.

Fusion routing

A task comes in. Routine work goes to self-hosted open models. Hard steps go to frontier models. Caching applies across the mix.

  • Self-hosted models for the routine majority of work
  • Frontier models only when the step needs them
  • Prompt caching applied automatically

Where it is used

One gateway for the product

Every application that calls a model goes through one URL. Routing, spend, and failover live in one place.

What Open Computer Agent uses

Each agent step uses a right-sized model instead of a single frontier default.

Private or hybrid

Serve your own models, still reach frontier models for hard steps, keep data in your environment.

Uptime, latency, and cost depend on the configured providers, regions, and traffic. Inferno Router supports routing and failover; it does not guarantee every model is available in every region. Cost-reduction figures are illustrative of right-sized routing versus all-frontier list prices.