aimldocs
Console

Models

Browse the catalog, compare models, and choose which host serves your requests.

Models in the console is the catalog: every model you can call, what it costs, what it can do, and which hosts serve it. From a model's page you can also choose where your requests are served: the fastest host, a host in one country, or one exact endpoint. Every role can use these screens; nothing here changes your organization.

Browse the catalog

Each model is a card showing its price per million tokens, context size, markup, uptime, time to first token and number of hosts. Click a card to open the model's page.

Narrow the list with the filter bar:

  • Search: type a name and press Enter.
  • Vendor and Modality (text, image, audio, video).
  • Capabilities: tools, json_schema, vision, audio, reasoning, cache. Select several to require all of them.
  • new this week, price changed, ZDR (zero data retention) and no-train.
  • sort: name, price, latency, throughput, popularity or newest.

Filters are kept in the page address, so a filtered view can be shared or bookmarked. If nothing matches, press Clear filters.

If your organization's data policy requires zero data retention or no-train providers, the catalog lists only models with a matching endpoint. A note at the top says so, and those two filters are switched on and cannot be turned off here.

Prices are in USD. If you have chosen a display currency in preferences, an approximate figure in that currency is shown as well. It is indicative only; you are billed in USD prices.

Compare models

  1. Tick compare on 2 to 4 model cards.
  2. Press Compare.

The comparison shows the models side by side: vendor, context, maximum output, input, output and cache read prices per million tokens, markup, 30-day uptime, time to first token, requests over 7 days, number of hosts, capabilities, modalities and status. Each model's name links to its page.

Press Open in playground compare to send the same prompt to all of them in the playground.

The model page

The top of the page gives the model's name and id, its input and output types, context size, maximum output, tags and aliases. Badges say when some of its hosts are in canary or paused.

Two notices can appear:

  • Deprecated: the model is being retired. The notice gives the sunset date and the suggested replacement, linked to its page. The model keeps serving until that date.
  • Price change scheduled: which endpoint changes, the new prices, and from when. A request is charged at the price in force when it starts.

Cost calculator

Enter your typical Input tokens, Output tokens and Cached share. The endpoint table then shows each host's cost per 1,000 requests of that shape, so you can compare hosts on what you would actually pay and not only on list price.

Endpoints

The table lists every host that serves the model, with:

  • provider, region and quantization;
  • context and maximum output on that host;
  • input, output and cache prices per million tokens, and markup;
  • cost per 1,000 requests from the calculator;
  • 30-day uptime and throughput in output tokens per second;
  • Last 24 h: success rate, time to first token and throughput over the last day. A host with no requests in that time shows no traffic, which is not a failure;
  • Fidelity: for open-weight models served by several hosts, how closely this host's answers agree with a reference host. Flags mark a quantized build that diverges, or a host that could not answer a long-context probe;
  • the host's data handling (retention, training), capabilities and status.

If your organization's data policy excludes some hosts, they are hidden and a line says how many. If it excludes all of them, the page says requests would be rejected and links to the policy.

Choose an endpoint

By default the router chooses a host for each request. Choose an endpoint lets you decide, per request, without changing any setting on your account.

  1. Under Sort by, order the hosts by Least latency, Throughput, Price or Uptime.
  2. Under Country, optionally keep only hosts in one country.
  3. To pin one host, press Use this on its row. The button changes to Selected; press it again to unpin. Paused and canary hosts are listed but cannot be pinned.
  4. Press Copy config and add the copied aiml.route settings to your request body, or pass the same route settings to the ai.ml SDK.

What the copied config does depends on your choices:

  • Sorted by least latency, no pin: the router picks the fastest matching host and falls back to another on failure. If you picked a country, only hosts there are used.
  • A country, no pin: requests stay on hosts in that country.
  • A pinned host: a hard pin. The request goes only to that endpoint, and fails if it is unavailable.
  • A pinned host with "allow fallback if it's down" ticked: the pinned endpoint is tried first, and the others are fallbacks.

Sorting by throughput, price or uptime reorders the table to help you pick a host to pin; on its own it produces no config.

A pin can only narrow where a request goes. The pinned host must still pass your key's and organization's rules, such as model lists and data policy. Changing the country clears a pin.

Charts, quick start and history

  • 30-day uptime and latency: attempts per day, succeeded and failed, and time to first token per day. Choose the endpoint to chart from the selector.
  • Quick start: a ready request for this model. Choose the format from the selector: curl, the OpenAI SDK in Python or TypeScript, the Anthropic SDK, or the ai.ml SDK. See the quickstarts for more.
  • Apps using it · 7 days: the apps sending the most requests to this model. See apps.
  • Changelog: price changes and other recorded changes to the model.

On this page