Routing policies
Write a routing policy, check what it would do with a dry run, activate a version, and watch its canaries, shadows and experiments.
A routing policy decides which endpoint serves a request and in what order the fallbacks are tried. Without one, a request goes to the cheapest endpoint that can serve the model. A policy is how you say something different: keep to a region, prefer latency, hold a provider back.
Open Routing in the sidebar. Creating, editing and activating policies needs the owner or admin role. Members and viewers can read policies and their dashboards.
Two rules hold throughout:
- Nothing is stored that does not compile. A policy with an error cannot be saved.
- A version never changes once it is activated. Editing means saving a new version, so the version list is a true history.
A policy can also be managed from your own code. The grammar is described in full in the policies API reference.
The screen
The policies are listed on the left, each with the status of its current version: active, draft or retired. Select one to see its five tabs: editor, versions, canaries, shadows and experiments.
Create a policy
- Press New policy. With no policy yet, the button reads Write your first policy.
- Enter a Name. It is lower-case and must be the same name the policy gives itself in its text.
- Edit the starter policy in the editor.
- Press Create policy.
The new policy is saved as a draft. It does nothing until you activate it.
What a policy can express
A policy is written as YAML in the editor. In words, it has these parts:
- Defaults. How candidates are sorted (for example by price or by latency), whether fallbacks are allowed, the maximum number of attempts, a time budget for the first token and for the whole request, fallback models, and requirements and exclusions that apply to every request.
- Requirements. Conditions an endpoint must meet: a data policy (no training, zero data retention), regions, a maximum price per million input or output tokens, quantization, a minimum context size, capabilities, a minimum 30-day uptime and a minimum quality score.
- Exclusions. Providers, endpoints or quantizations never to use.
- Rules. Each rule has a name and a condition on the request, for example an estimate of its input tokens. When the condition matches, the rule can change the sort, add requirements or exclusions, name the candidate endpoints or providers with weights, and set fallback models.
- Canary. Send a percentage of traffic to one endpoint, protected by a guard.
- Shadow. Send a copy of a sample of requests to a second endpoint for comparison.
- Experiments. Split traffic between named arms.
The editor
The editor tab shows the newest version of the policy.
Validation
The policy is compiled each time you pause typing. The line under the editor shows the result:
- When it compiles, the line counts the rules, canaries, shadows and experiments it found.
- When it does not, the line shows the path of the field that is wrong, followed by the message.
Conditions are type-checked, so a misspelled request field or a comparison between the wrong kinds of value is caught here and not in production.
Endpoint ids
The Endpoint ids panel beside the editor is a reference. Choose a model to list its endpoints with their region and status. These ids, or a provider id, are what a policy names in its requirements, exclusions and canary. They cannot be guessed, so copy them from here.
Save a new version
Press Save as new version. The button is disabled until you have changed something and the policy compiles. The text beside it says which version you are editing on top of.
The screen confirms the new version number and that it is a draft. Activate it on the versions tab when you are ready.
Dry run
The Dry run panel under the editor answers the question "if I activate this, what changes?". It runs the real router over the real catalog. No provider is called and nothing is spent.
- Give it a request, in one of two ways:
- Request id: a request your organization really made, from the requests screen.
- or Model and Est. input tokens: a request you describe.
- Press Run.
The result shows:
- The decision: the sort, the affinity, the maximum attempts, and any experiment arm, canary or shadows that would apply.
- Notes from the policy, when there are any.
- The plan: each attempt in order, with its endpoint, region, estimated cost and the reason it was chosen.
- The endpoints that were removed, and at which stage and why. A removal caused by your policy is described in the policy's own words.
- The full Route trace.
If no endpoint could serve the request, the result shows the error the request would have received.
Canary guards are deliberately ignored in a dry run. A guard is a fact about production at one moment, and it would make the same dry run answer differently from minute to minute.
Versions
The versions tab lists every version with its status, when it was created and when it was activated. The version that is serving reads live.
Compare two versions
With more than one version, choose two in the selectors under the table. The diff shows the lines added and removed between them.
Activate a version
- Press Activate on the version's row. The panel says which version stops serving and is retired, or that nothing is active yet.
- Choose the Scope: Whole org or one project.
- In Models, enter a pattern over model ids. An asterisk alone matches every model. A vendor followed by a slash and an asterisk matches one vendor.
- Set the Weight, from 1 to 100.
- Press the Activate button that names the version.
The previous active version is retired.
Scope and weight
A policy can be bound to the whole organization or to one project. The narrowest binding sets the structure, and every wider binding's requirements, exclusions and price limits still apply. A narrower scope can tighten what a wider policy allows, never widen it.
When two bindings sit at one scope, traffic is split between them by weight. The split is stable, so a given API key always gets the same one.
Canaries
A canary sends a slice of traffic to a new endpoint and holds it back automatically when the endpoint misbehaves. Add a canary section to the policy, and the canaries tab shows one row per canary:
- Endpoint and the policy version.
- Guard:
green,amber,redorunknown, with the part of the configured share it is serving. - Error rate: the measured rate against the limit you set.
- Samples: the number of requests measured against the minimum you set.
- Decided: when the guard was last evaluated, and the reason.
How the guard decides:
- green: inside the error limit with enough samples. The canary gets its full share.
- amber: not enough evidence yet, in either direction. The canary gets a tenth of its share.
- red: over the limit with enough samples, or no traffic at all. The canary gets nothing.
- unknown: not yet measured. It is treated as red.
Guards are re-evaluated every 5 minutes over the last hour of traffic, so check the Decided column to see how fresh a state is.
Shadows
A shadow sends a copy of a request to a second endpoint and throws the answer away. You can compare cost and latency without changing what your user received. Add a shadow section to the policy, and the shadows tab shows one row per shadow endpoint:
- Shadow endpoint, and the primary endpoint it was compared against.
- Runs and Errors.
- Stop differs: how often the shadow stopped for a different reason than the primary.
- Cost vs primary: the difference, with both costs below it.
- TTFT: the time to the first token for the shadow and for the primary.
- Billed to:
orgorplatform.
A shadow runs after the primary request has finished and never slows it down. You pay for a shadow only when the policy says to bill it to your organization. A request under zero data retention is never shadowed. The outputs are not shown here.
Experiments
An experiment splits traffic between arms and stamps each request with the arm it took. Add an experiment section to the policy and activate it, and the experiments tab compares the arms.
Choose the window: Last hour, Last 24 hours or Last 7 days. Each arm shows:
- Requests and Errors, with the error rate.
- Cost and Cost / 1k req.
- TTFT p50 and TTFT p95.
These figures come from analytics. When analytics are running behind, a banner says by how much. This is the only routing tab that can show it.
The three dashboards are read-only. Looking at them never changes a routing decision.
Related
- Models: endpoints, prices and measured latency per model.
- Evals: quality scores on your own prompts, which the quality sort can use.
- Data policy: organization-wide limits on providers and regions.