anonrouterdocs

Provider routing and failover

Choose which provider serves a model, and how AnonRouter moves to another provider when one rate-limits or fails, without lowering privacy.

Model routing and provider routing are separate steps. /auto chooses a model. Provider routing chooses who serves the model you named. Several providers can serve the same model, and switching between them never changes the model id you send.

By default you do not need to set anything. AnonRouter picks the most private route for the model, prefers providers that are working right now, and, if the chosen provider rate-limits or errors before sending any output, tries the next one automatically.

Models and provider routes

A model has one canonical creator/model id, for example deepseek/deepseek-v3.2. Each provider that serves it is a provider route with its own privacy tier and price. GET /v1/models lists the routes under the canonical model in provider_routes, and auto_provider names the provider Auto would pick right now. See models and privacy labels.

There are no provider-prefixed public model ids. To choose a provider, use the provider field described below.

The provider request field

Send provider in the chat completion body. It is either a provider id (an exact pin) or a policy object:

FieldValuesPurpose
orderarray of provider idsTry these providers first, in this order. Remaining eligible routes fill in by sort.
onlyarray of provider idsUse only these providers.
ignorearray of provider idsNever use these providers.
allow_fallbacksboolean, default truefalse makes a single attempt on the first eligible route.
sortprivacy (default), price, latency, throughputHow eligible routes are ordered.
minimum_privacyanonymous, private, tee, e2eeThe lowest tier any attempt may use.
max_price{ "input": number, "output": number }Ceiling in USD per million tokens. Routes above it are skipped.
require_parametersbooleanOnly use routes that support every parameter in the request, such as tools or reasoning.
max_attempts1 to 3How many providers may be tried for this request. Defaults to 3, or your workspace's limit.

A bare string is an exact pin with no fallback. "provider": "venice" is the same as { "only": ["venice"], "allow_fallbacks": false }. Omitting provider means Auto.

Provider ids are the catalog's provider values, such as venice, chutes, or deepinfra. An id also matches that provider's variants and regions. Models served through partner providers that are not listed by name appear under the provider Other, with id other. other works in order, only, ignore, and as a pin, and matches every route listed under Other.

The policy is strict. A provider in both only and ignore, or an order entry that only excludes or ignore removes, is a provider_policy_conflict. An unknown key or an out-of-range value, such as max_attempts: 4, is a 400 invalid_request. The provider object is never forwarded to the upstream provider.

Ticket flow: bind the same policy

In the private ticket flow, send the same provider value in the ticket request (POST /v1/inference/tickets) and in the completion body. The ticket binds the policy, and a body whose policy differs is refused with 409 ticket_provider_policy_mismatch before any provider work. In compatibility mode you send it in the body only.

How Auto picks a provider

When you omit provider, or leave the choice to sort, AnonRouter:

  1. Drops routes that are unavailable, disabled, outside your only/ignore and workspace settings, missing a capability the request needs, or too expensive for max_price.
  2. Keeps the strongest privacy tier among the routes that remain. Routes below it are not used unless you allow a lower tier.
  3. Skips routes whose context window cannot hold the prompt plus the requested output. If no route at that tier can, the request fails with context_too_large rather than moving to a weaker tier.
  4. Prefers healthy routes over degraded ones, including routes that failed recently (see Recently failed providers go last).
  5. Picks the lowest price, then breaks ties by latency and a fixed provider order.

E2EE routes serve only E2EE requests, so Auto never picks one for a plaintext request. See end-to-end encrypted inference.

With sort, the order changes:

  • privacy (default): strongest tier first, then health, then price.
  • price: cheapest first, never below the privacy floor.
  • latency: fastest measured time to first token over the last 24 hours first.
  • throughput: most measured output tokens per second over the last 24 hours first.

When there is not enough latency or throughput data (fewer than 20 recent requests), routes are ordered by health, then price. A non-default sort with no minimum_privacy uses Private as the floor. Anonymous routes need an explicit "minimum_privacy": "anonymous".

Under every sort, a route that recently failed goes behind the healthy routes. Under privacy, privacy rank still comes first.

Automatic failover

If the provider serving your request fails before any output reaches you, AnonRouter tries the next route in the plan, up to 3 providers in total. You get one response, from whichever provider succeeded. Before any output means: for a non-streaming request, before the JSON body is sent; for a streaming request, before the response headers are sent.

What happens depends on what the provider returned:

Provider resultTries the next provider?Recently-failed penalty
Rate limited (429)YesThat route
Server error (5xx)YesThat route
TimeoutYesThat route
Network failure before any outputYesThat route
Malformed or incomplete response before any outputYesThat route
Model unavailable (404)YesThat route
Provider authentication failure (401)YesEvery route of that provider
Provider quota exhausted (402)YesEvery route of that provider
Invalid request (400)NoNone
Other 4xx (403, 409, 413, 422, ...)YesNone
Content-policy refusalNoNone

A 401 or 402 from a provider is about AnonRouter's own account with that provider, never about your request or your key, so the whole provider is penalized and the request moves on.

A 400 means the request itself is invalid for that model. Sending it to another provider would show your prompt to one more party, almost always for the same answer, so AnonRouter stops. You receive the provider's error as a sanitized 400 (raw upstream error bodies are never returned). If the 400 comes from a later attempt, after a fallback already happened, you receive 503 provider_fallback_exhausted.

A 400 or a refusal says nothing about the provider's health, so neither one penalizes the route.

The Default sort

New workspaces start with the Default provider sort, a blend that leans private. Workspaces created earlier keep the sort they had (Privacy unless you changed it), and you can switch on the Routing page.

  1. Your privacy boundary, disabled providers, fallback floor and price ceilings apply first, as always.
  2. It prefers the more private tier. A TEE provider is chosen over a Private one unless it costs more than 1.5 times as much, is more than twice as slow, or keeps failing. An Anonymous provider is used only when every more private provider costs more than 3 times as much, is more than 3 times as slow, or keeps failing, or when the model has no more private provider and your privacy boundary allows Anonymous.
  3. Within a tier it weighs the expected cost of a typical request (counting cached-input pricing), measured speed and success rate over the last 24 hours. Providers with fewer than 20 recent requests are treated as average, never guessed.
  4. Providers that score within 10% of the best share traffic, which also helps avoid rate limits on any one provider.
  5. Fallback follows the same order and never steps below the first provider's tier unless your fallback floor allows it.

The measurements aggregate traffic by provider route and model. They contain no prompt content and nothing about your account. AnonRouter does not send a conversation back to the same provider to improve cache hits, because that would mean linking your requests.

For now Default is a workspace setting only; a request cannot send "sort": "default".

Recently failed providers go last

When a route fails with one of the penalized results above, AnonRouter tries it last for a while:

  • 30 seconds after the first failure;
  • each further failure within 10 minutes doubles it: 60, 120, 240 seconds, up to 5 minutes;
  • the count resets 10 minutes after the last failure.

During that time the route is tried after every healthy route, under every sort. It is never excluded: if every route for the model is in this state, they are tried in their normal order, so a model with a single route keeps working.

The penalty never crosses a privacy boundary. It only reorders routes that already passed your privacy floor and every other filter. Under the default privacy sort, a recently failed route is tried after the healthy routes of its own tier and still ahead of any weaker tier you allowed.

An explicit order is honored as written. The penalty reorders only the routes that Auto or sort places. If you need a fixed provider sequence, set order or pin a provider.

The penalty describes the provider, not you. AnonRouter records only the route or provider, the kind of failure, a strike count, and an expiry. No account, key, request id, or content.

When there is no failover

  • After output has started. Once the JSON body or the stream headers are sent, a failure is reported to you rather than retried elsewhere.
  • E2EE requests. The ciphertext and attestation are bound to one enclave. If that provider fails before output you get 503 provider_fallback_not_supported_for_e2ee. Retry with fresh attestation and a fresh ticket.
  • Below your privacy floor. Failover never moves to a weaker tier than the request and your workspace allow, and never below the workspace's privacy boundary.
  • Disabled or ignored providers. They are never used, as first choice or as fallback.
  • An invalid request (400) or a content-policy refusal.
  • Exact pins, allow_fallbacks: false, or max_attempts: 1.
  • After you cancel the request.

Workspace defaults

Each workspace sets its defaults in the dashboard, split across two pages. Every section says what it applies to.

  • Privacy page:
    • Privacy boundary. The weakest privacy tier any request from the workspace may use: Standard (the default), Private, or Maximum. See privacy boundary.
    • Provider fallback:
      • Allow fallback to another provider. On by default. Off means one attempt, and the request fails if that provider fails.
      • Retries on another provider. Up to 1 or up to 2, default up to 2. Shown while fallback is on. Each retry goes to a different provider for the same model and follows the same privacy rules as the first attempt. This is the API's max_attempts minus the first attempt.
      • How far fallback may step down. The lowest tier a fallback may use. The default, No downgrade, keeps every request at its own tier. Choices below the privacy boundary are greyed out.
    • What we keep. A summary of what AnonRouter retains, with a link to the privacy policy.
  • Routing page:
    • Auto router. How /auto picks a model: the strategy, which privacy tiers it may use (Standard, Private by default, or Maximum, never below the privacy boundary), and which models it may use (all models, only these, or all except these). Only affects requests to /auto.
    • Default provider sort. Every request that does not send its own provider settings. Default, Privacy, Price, Latency, or Throughput. See the Default sort.
    • Providers. Every request. A disabled provider is never used for this workspace, and a request cannot turn it back on.

The same settings are available with a session through GET and PUT /v1/routing/preferences?workspace_id=<id> (fields provider_allow_fallbacks, provider_max_attempts, provider_minimum_privacy, provider_sort, and provider_ignore). Without workspace_id they apply to the Default Workspace. PUT replaces the whole preferences object, so read it first and send it back with your changes.

Workspace defaults and request settings merge restrictively: a request can make routing stricter, never looser.

  • Request only is intersected with the workspace's allowed providers.
  • Request ignore is added to the workspace's disabled providers.
  • A request cannot lower the workspace's privacy floor or raise its price ceiling.
  • A request can turn fallback off or lower max_attempts, but cannot allow more attempts than the workspace does.

See workspaces.

Billing across attempts

Before the first attempt, AnonRouter reserves enough of your balance for the most expensive attempt your policy permits, so a fallback can never land outside the amount you authorized. Failed attempts are not charged: AnonRouter absorbs their cost. You pay only for the attempt that succeeded, at that route's price, and the rest of the reservation is released. See metering and settlement.

Response headers

Every completion reports the route that actually served it:

x-anonrouter-provider: venice           # provider that served the request
x-anonrouter-routing: exact             # "auto" when /auto chose the model
x-anonrouter-provider-attempts: 2       # providers tried for this request
x-anonrouter-provider-fallback: true    # whether a fallback provider served it
x-anonrouter-privacy-class: private     # tier of the serving route

x-anonrouter-provider uses public provider ids, so routes listed under Other report other. x-anonrouter-privacy-class is the tier of the route that served the request: anonymous, private, tee, or e2ee. The ticket's privacy_class is the weakest tier your policy could reach, so it never overstates the guarantee.

Errors

StatustypeMeaning
400invalid_provider_policyThe provider value is malformed, for example an invalid provider id
400provider_policy_conflictContradictory fields, such as a provider in both only and ignore
400no_provider_routeNo eligible provider route is available for the model under your policy
403no_provider_route_meets_privacyNo route meets the privacy floor
403privacy_boundary_not_metNo route of the model meets the workspace's privacy boundary. Refused before any provider sees the request
400no_provider_route_meets_priceNo route is within max_price
409ticket_provider_policy_mismatchThe body's provider differs from the policy the ticket was issued for
503provider_fallback_exhaustedEvery permitted attempt failed before output after a fallback, or a later attempt got a 400
503provider_fallback_not_supported_for_e2eeAn E2EE provider failed before output; E2EE cannot switch providers

When only one provider was tried, you receive that provider's own sanitized error instead of provider_fallback_exhausted.

Examples

These use a compatibility-mode key. In the ticket flow, put the same provider value in the ticket request as well.

# Auto: strongest privacy, healthy providers first, then lowest price.
curl -i https://api.anonrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $ANONROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v3.2",
    "messages": [{ "role": "user", "content": "Hi" }]
  }'

# Exact pin: one provider, no fallback.
curl https://api.anonrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $ANONROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v3.2",
    "provider": "venice",
    "messages": [{ "role": "user", "content": "Hi" }]
  }'

# Cheapest route at or above Private, at most two providers.
curl https://api.anonrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $ANONROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v3.2",
    "provider": { "sort": "price", "minimum_privacy": "private", "max_attempts": 2 },
    "messages": [{ "role": "user", "content": "Hi" }]
  }'
const response = await fetch("https://api.anonrouter.ai/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.ANONROUTER_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "deepseek/deepseek-v3.2",
    // Try DeepInfra first, then Venice, then anything else Auto allows.
    provider: { order: ["deepinfra", "venice"], allow_fallbacks: true },
    messages: [{ role: "user", content: "Hello" }],
  }),
});

console.log(response.headers.get("x-anonrouter-provider"));
console.log(response.headers.get("x-anonrouter-provider-fallback"));

A provider policy can also constrain /auto. The model is chosen only from models with an eligible route under the policy, and provider selection then runs for the chosen model:

{
  "model": "/auto",
  "provider": { "only": ["venice"], "minimum_privacy": "private" },
  "messages": [{ "role": "user", "content": "Hello" }]
}

Or hand the whole thing to your agent

One prompt carries the entire setup. Give it to your agent, review what it configures, and you are done.

On this page