Evals
Grade your own prompts against a set of models and endpoints, keep a run inside a budget, and track regressions between runs.
An eval measures models on your own prompts. You write a suite of graded prompts, run it against the models and endpoints you choose, and compare the scores with the previous run.
Open Evals in the sidebar. The screen has two tabs: Suites and Runs.
Creating suites, deleting them and starting runs needs the owner, admin or member role. Billing and viewer roles can read suites and runs.
Suites
A suite is a named set of cases. Each case has:
- an id;
- an input, which is the request to send;
- a grader, which decides how the output is scored;
- optionally a weight.
Graders
| Grader | Passes when |
|---|---|
exact | The output equals the expected text. |
contains | The output contains a given piece of text. |
regex | The output matches a pattern. |
json_schema | The output is valid against a JSON Schema. |
llm_judge | A judge model scores the output against your rubric, at or above a threshold. |
Create a suite
- On the Suites tab, press New suite.
- Enter a Name.
- Edit Cases (JSON). The field starts with a one-case template that you can change and extend.
- Press Create suite.
The cases must be valid JSON. If they are not, or if a case is refused, the drawer shows the error and nothing is created.
The table lists each suite with its id, its number of cases and when it was last updated.
Delete a suite
Press Delete on the suite's row, type delete to confirm, and press Delete suite. The suite's runs and results are deleted with it.
Start a run
- On the Suites tab, press Run on the suite's row.
- Enter the Targets, separated by commas. A target is a model id, or a model id followed by an @ sign and an endpoint id to test one particular endpoint. Endpoint ids are listed on each model's page; see models.
- Enter the Budget (USD).
- Press Start run.
The drawer confirms the run's id and an estimate of its cost.
A run works in the background. Each case is an ordinary request made on your organization's account, so it is billed like any other request and appears in your request history. See requests.
The budget
A run adds up the cost of its requests and stops once the budget is reached. Targets are worked through one at a time, so a run that stops early leaves whole targets unstarted and does not leave every target half measured.
Runs
The Runs tab lists every run with its id, status, cost and start time. A run that completed is shown in green, one that stopped early in amber and one that failed in red.
Click a run to open its detail.
Scores
One row per target, with:
- Mean: the weighted mean score across the cases.
- Passed: the number of cases passed out of the total, and the number of errors if there were any.
A case whose grader could not run, for example because of an invalid pattern or schema, is recorded as an error. It is not counted as a score of zero, and errors are left out of the mean. A grader that broke does not mean the model failed.
If the run stopped early or failed, the reason is shown at the top as Stopped.
Regressions
Regressions vs previous run appears when a target scored lower than in the previous completed run of the same suite. Each line shows the target, the previous mean, the current mean and the difference.
Per-case results
The last table lists every case for every target with its score, or err when the grader could not run.
How scores affect routing
A completed run's mean score per endpoint becomes that endpoint's quality score for your organization. A routing policy that sorts by quality, or that requires a minimum quality, uses it. Only targets that name an endpoint contribute. An endpoint you have not measured is not treated as bad: the quality sort falls back to price for it.
See routing policies.
Your suites, runs and results are private to your organization.
Related
- Evals API reference
- Playground: try a single prompt by hand before you write a suite.