Provider routing and failover
Choose which provider serves a model, and how AnonRouter moves to another provider when one rate-limits or fails, without lowering privacy.
Model routing and provider routing are separate steps. /auto
chooses a model. Provider routing chooses who serves the model you named. Several
providers can serve the same model, and switching between them never changes the
model id you send.
By default you do not need to set anything. AnonRouter picks the most private route for the model, prefers providers that are working right now, and, if the chosen provider rate-limits or errors before sending any output, tries the next one automatically.
Models and provider routes
A model has one canonical creator/model id, for example
deepseek/deepseek-v3.2. Each provider that serves it is a provider route
with its own privacy tier and price. GET /v1/models lists the routes under the
canonical model in provider_routes, and auto_provider names the provider
Auto would pick right now. See models and privacy labels.
There are no provider-prefixed public model ids. To choose a provider, use the
provider field described below.
The provider request field
Send provider in the chat completion body. It is either a provider id (an
exact pin) or a policy object:
| Field | Values | Purpose |
|---|---|---|
order | array of provider ids | Try these providers first, in this order. Remaining eligible routes fill in by sort. |
only | array of provider ids | Use only these providers. |
ignore | array of provider ids | Never use these providers. |
allow_fallbacks | boolean, default true | false makes a single attempt on the first eligible route. |
sort | privacy (default), price, latency, throughput | How eligible routes are ordered. |
minimum_privacy | anonymous, private, tee, e2ee | The lowest tier any attempt may use. |
max_price | { "input": number, "output": number } | Ceiling in USD per million tokens. Routes above it are skipped. |
require_parameters | boolean | Only use routes that support every parameter in the request, such as tools or reasoning. |
max_attempts | 1 to 3 | How many providers may be tried for this request. Defaults to 3, or your workspace's limit. |
A bare string is an exact pin with no fallback. "provider": "venice" is
the same as { "only": ["venice"], "allow_fallbacks": false }. Omitting
provider means Auto.
Provider ids are the catalog's provider values, such as venice, chutes, or
deepinfra. An id also matches that provider's variants and regions. Models
served through partner providers that are not listed by name appear under the
provider Other, with id other. other works in order, only, ignore,
and as a pin, and matches every route listed under Other.
The policy is strict. A provider in both only and ignore, or an order
entry that only excludes or ignore removes, is a provider_policy_conflict.
An unknown key or an out-of-range value, such as max_attempts: 4, is a
400 invalid_request. The provider object is never forwarded to the upstream
provider.
Ticket flow: bind the same policy
In the private ticket flow, send the same provider value
in the ticket request (POST /v1/inference/tickets) and in the completion
body. The ticket binds the policy, and a body whose policy differs is refused
with 409 ticket_provider_policy_mismatch before any provider work. In
compatibility mode you send it in the body only.
How Auto picks a provider
When you omit provider, or leave the choice to sort, AnonRouter:
- Drops routes that are unavailable, disabled, outside your
only/ignoreand workspace settings, missing a capability the request needs, or too expensive formax_price. - Keeps the strongest privacy tier among the routes that remain. Routes below it are not used unless you allow a lower tier.
- Skips routes whose context window cannot hold the prompt plus the requested
output. If no route at that tier can, the request fails with
context_too_largerather than moving to a weaker tier. - Prefers healthy routes over degraded ones, including routes that failed recently (see Recently failed providers go last).
- Picks the lowest price, then breaks ties by latency and a fixed provider order.
E2EE routes serve only E2EE requests, so Auto never picks one for a plaintext request. See end-to-end encrypted inference.
With sort, the order changes:
privacy(default): strongest tier first, then health, then price.price: cheapest first, never below the privacy floor.latency: fastest measured time to first token over the last 24 hours first.throughput: most measured output tokens per second over the last 24 hours first.
When there is not enough latency or throughput data (fewer than 20 recent
requests), routes are ordered by health, then price. A non-default sort with no minimum_privacy uses Private
as the floor. Anonymous routes need an explicit "minimum_privacy": "anonymous".
Under every sort, a route that recently failed goes behind the healthy routes.
Under privacy, privacy rank still comes first.
Automatic failover
If the provider serving your request fails before any output reaches you, AnonRouter tries the next route in the plan, up to 3 providers in total. You get one response, from whichever provider succeeded. Before any output means: for a non-streaming request, before the JSON body is sent; for a streaming request, before the response headers are sent.
What happens depends on what the provider returned:
| Provider result | Tries the next provider? | Recently-failed penalty |
|---|---|---|
Rate limited (429) | Yes | That route |
Server error (5xx) | Yes | That route |
| Timeout | Yes | That route |
| Network failure before any output | Yes | That route |
| Malformed or incomplete response before any output | Yes | That route |
Model unavailable (404) | Yes | That route |
Provider authentication failure (401) | Yes | Every route of that provider |
Provider quota exhausted (402) | Yes | Every route of that provider |
Invalid request (400) | No | None |
Other 4xx (403, 409, 413, 422, ...) | Yes | None |
| Content-policy refusal | No | None |
A 401 or 402 from a provider is about AnonRouter's own account with that
provider, never about your request or your key, so the whole provider is
penalized and the request moves on.
A 400 means the request itself is invalid for that model. Sending it to
another provider would show your prompt to one more party, almost always for the
same answer, so AnonRouter stops. You receive the provider's error as a
sanitized 400 (raw upstream error bodies are never returned). If the 400
comes from a later attempt, after a fallback already happened, you receive
503 provider_fallback_exhausted.
A 400 or a refusal says nothing about the provider's health, so neither one
penalizes the route.
The Default sort
New workspaces start with the Default provider sort, a blend that leans private. Workspaces created earlier keep the sort they had (Privacy unless you changed it), and you can switch on the Routing page.
- Your privacy boundary, disabled providers, fallback floor and price ceilings apply first, as always.
- It prefers the more private tier. A TEE provider is chosen over a Private one unless it costs more than 1.5 times as much, is more than twice as slow, or keeps failing. An Anonymous provider is used only when every more private provider costs more than 3 times as much, is more than 3 times as slow, or keeps failing, or when the model has no more private provider and your privacy boundary allows Anonymous.
- Within a tier it weighs the expected cost of a typical request (counting cached-input pricing), measured speed and success rate over the last 24 hours. Providers with fewer than 20 recent requests are treated as average, never guessed.
- Providers that score within 10% of the best share traffic, which also helps avoid rate limits on any one provider.
- Fallback follows the same order and never steps below the first provider's tier unless your fallback floor allows it.
The measurements aggregate traffic by provider route and model. They contain no prompt content and nothing about your account. AnonRouter does not send a conversation back to the same provider to improve cache hits, because that would mean linking your requests.
For now Default is a workspace setting only; a request cannot send
"sort": "default".
Recently failed providers go last
When a route fails with one of the penalized results above, AnonRouter tries it last for a while:
- 30 seconds after the first failure;
- each further failure within 10 minutes doubles it: 60, 120, 240 seconds, up to 5 minutes;
- the count resets 10 minutes after the last failure.
During that time the route is tried after every healthy route, under every
sort. It is never excluded: if every route for the model is in this state,
they are tried in their normal order, so a model with a single route keeps
working.
The penalty never crosses a privacy boundary. It only reorders routes that
already passed your privacy floor and every other filter. Under the default
privacy sort, a recently failed route is tried after the healthy routes of its
own tier and still ahead of any weaker tier you allowed.
An explicit order is honored as written. The penalty reorders only the routes
that Auto or sort places. If you need a fixed provider sequence, set order
or pin a provider.
The penalty describes the provider, not you. AnonRouter records only the route or provider, the kind of failure, a strike count, and an expiry. No account, key, request id, or content.
When there is no failover
- After output has started. Once the JSON body or the stream headers are sent, a failure is reported to you rather than retried elsewhere.
- E2EE requests. The ciphertext and attestation are bound to one enclave. If
that provider fails before output you get
503 provider_fallback_not_supported_for_e2ee. Retry with fresh attestation and a fresh ticket. - Below your privacy floor. Failover never moves to a weaker tier than the request and your workspace allow, and never below the workspace's privacy boundary.
- Disabled or ignored providers. They are never used, as first choice or as fallback.
- An invalid request (
400) or a content-policy refusal. - Exact pins,
allow_fallbacks: false, ormax_attempts: 1. - After you cancel the request.
Workspace defaults
Each workspace sets its defaults in the dashboard, split across two pages. Every section says what it applies to.
- Privacy page:
- Privacy boundary. The weakest privacy tier any request from the workspace may use: Standard (the default), Private, or Maximum. See privacy boundary.
- Provider fallback:
- Allow fallback to another provider. On by default. Off means one attempt, and the request fails if that provider fails.
- Retries on another provider. Up to 1 or up to 2, default up to 2.
Shown while fallback is on. Each retry goes to a different provider for
the same model and follows the same privacy rules as the first attempt.
This is the API's
max_attemptsminus the first attempt. - How far fallback may step down. The lowest tier a fallback may use. The default, No downgrade, keeps every request at its own tier. Choices below the privacy boundary are greyed out.
- What we keep. A summary of what AnonRouter retains, with a link to the privacy policy.
- Routing page:
- Auto router. How
/autopicks a model: the strategy, which privacy tiers it may use (Standard, Private by default, or Maximum, never below the privacy boundary), and which models it may use (all models, only these, or all except these). Only affects requests to/auto. - Default provider sort. Every request that does not send its own
providersettings. Default, Privacy, Price, Latency, or Throughput. See the Default sort. - Providers. Every request. A disabled provider is never used for this workspace, and a request cannot turn it back on.
- Auto router. How
The same settings are available with a session through
GET and PUT /v1/routing/preferences?workspace_id=<id> (fields
provider_allow_fallbacks, provider_max_attempts,
provider_minimum_privacy, provider_sort, and provider_ignore). Without
workspace_id they apply to the Default Workspace. PUT replaces the whole
preferences object, so read it first and send it back with your changes.
Workspace defaults and request settings merge restrictively: a request can make routing stricter, never looser.
- Request
onlyis intersected with the workspace's allowed providers. - Request
ignoreis added to the workspace's disabled providers. - A request cannot lower the workspace's privacy floor or raise its price ceiling.
- A request can turn fallback off or lower
max_attempts, but cannot allow more attempts than the workspace does.
See workspaces.
Billing across attempts
Before the first attempt, AnonRouter reserves enough of your balance for the most expensive attempt your policy permits, so a fallback can never land outside the amount you authorized. Failed attempts are not charged: AnonRouter absorbs their cost. You pay only for the attempt that succeeded, at that route's price, and the rest of the reservation is released. See metering and settlement.
Response headers
Every completion reports the route that actually served it:
x-anonrouter-provider: venice # provider that served the request
x-anonrouter-routing: exact # "auto" when /auto chose the model
x-anonrouter-provider-attempts: 2 # providers tried for this request
x-anonrouter-provider-fallback: true # whether a fallback provider served it
x-anonrouter-privacy-class: private # tier of the serving routex-anonrouter-provider uses public provider ids, so routes listed under Other
report other. x-anonrouter-privacy-class is the tier of the route that
served the request: anonymous, private, tee, or e2ee. The ticket's
privacy_class is the weakest tier your policy could reach, so it never
overstates the guarantee.
Errors
| Status | type | Meaning |
|---|---|---|
400 | invalid_provider_policy | The provider value is malformed, for example an invalid provider id |
400 | provider_policy_conflict | Contradictory fields, such as a provider in both only and ignore |
400 | no_provider_route | No eligible provider route is available for the model under your policy |
403 | no_provider_route_meets_privacy | No route meets the privacy floor |
403 | privacy_boundary_not_met | No route of the model meets the workspace's privacy boundary. Refused before any provider sees the request |
400 | no_provider_route_meets_price | No route is within max_price |
409 | ticket_provider_policy_mismatch | The body's provider differs from the policy the ticket was issued for |
503 | provider_fallback_exhausted | Every permitted attempt failed before output after a fallback, or a later attempt got a 400 |
503 | provider_fallback_not_supported_for_e2ee | An E2EE provider failed before output; E2EE cannot switch providers |
When only one provider was tried, you receive that provider's own sanitized
error instead of provider_fallback_exhausted.
Examples
These use a compatibility-mode key. In the ticket flow, put
the same provider value in the ticket request as well.
# Auto: strongest privacy, healthy providers first, then lowest price.
curl -i https://api.anonrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $ANONROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v3.2",
"messages": [{ "role": "user", "content": "Hi" }]
}'
# Exact pin: one provider, no fallback.
curl https://api.anonrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $ANONROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v3.2",
"provider": "venice",
"messages": [{ "role": "user", "content": "Hi" }]
}'
# Cheapest route at or above Private, at most two providers.
curl https://api.anonrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $ANONROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v3.2",
"provider": { "sort": "price", "minimum_privacy": "private", "max_attempts": 2 },
"messages": [{ "role": "user", "content": "Hi" }]
}'const response = await fetch("https://api.anonrouter.ai/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.ANONROUTER_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "deepseek/deepseek-v3.2",
// Try DeepInfra first, then Venice, then anything else Auto allows.
provider: { order: ["deepinfra", "venice"], allow_fallbacks: true },
messages: [{ role: "user", content: "Hello" }],
}),
});
console.log(response.headers.get("x-anonrouter-provider"));
console.log(response.headers.get("x-anonrouter-provider-fallback"));A provider policy can also constrain /auto. The model is chosen
only from models with an eligible route under the policy, and provider selection
then runs for the chosen model:
{
"model": "/auto",
"provider": { "only": ["venice"], "minimum_privacy": "private" },
"messages": [{ "role": "user", "content": "Hello" }]
}Or hand the whole thing to your agent
One prompt carries the entire setup. Give it to your agent, review what it configures, and you are done.