Best price-to-performance
…
ai& currently serves the open-weight models below. The list is dynamic — each organization sees only what it has access to, so the source of truth is always GET /v1/models.
Best price-to-performance
…
Multimodal
…
Largest context
…
| Model ID | Capabilities | Context | Input / 1M | Output / 1M |
|---|---|---|---|---|
| Loading the current models… | ||||
Sorted by output price, low to high. Prices are USD per million tokens — see Pricing for how the formula computes a per-request cost.
To have ai& pick one of these per request instead of choosing yourself, see Automatic Selection.
| Capability | Meaning |
|---|---|
reasoning | Emits internal reasoning tokens (charged as output) and accepts reasoning_effort — the values it takes are listed in that model’s reasoning_efforts, since models don’t share one vocabulary. |
tool_calling | Accepts tools / tool_choice. See Tool Calling. |
vision | Accepts image inputs (image_url or file_id). See Vision. |
video | Accepts video inputs by file_id. See Video Understanding. |
document | Accepts PDF / document inputs. |
A capability absent from the list means requests using that feature will be rejected at validation time.
curl https://api.aiand.com/v1/modelsfrom openai import OpenAI
client = OpenAI(base_url="https://api.aiand.com/v1", api_key="sk-...")for m in client.models.list().data: print(m.id, m.context_window, m.capabilities)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.aiand.com/v1", apiKey: "sk-..." });const { data } = await client.models.list();for (const m of data) console.log(m.id, m.context_window, m.capabilities);aiand modelsaiand models --capability vision --sort outputPrices are per 1M tokens in your billing currency. See ai& CLI.
{ "object": "list", "data": [ { "id": "deepseek-ai/deepseek-v4-flash", "name": "deepseek-ai/DeepSeek-V4-Flash", "object": "model", "created": 1780245232, "owned_by": "ai&", "provider": "deepseek-ai", "context_window": 1048576, "max_output_tokens": 384000, "capabilities": [ "reasoning", "tool_calling" ], "reasoning_efforts": [ "none", "high", "max" ], "reasoning_effort_default": "high", "description": "Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work", "currency": "usd", "input_per_1m": "0.150000", "output_per_1m": "0.250000", "cached_input_per_1m": "0.080000" } ]}| Field | Type | Meaning |
|---|---|---|
id | string | The value to pass as "model" in chat / responses / messages. |
provider | string | The lab that released the open weights (openai, google, qwen, …). All models are hosted on ai& infrastructure regardless of provider. |
context_window | int | Max combined input + output tokens. |
max_output_tokens | int | null | Recorded maximum generation length. null when no limit is recorded for the model. The list does not fill in a guess. |
capabilities | string[] | Supported features — see the table above. |
reasoning_efforts | string[] | null | The reasoning_effort values this model accepts; any other value is rejected with a 400. null on models that take no effort parameter. |
reasoning_effort_default | string | null | The level to default a UI to for this model. Not applied server-side — omitting reasoning_effort leaves the upstream’s own default in place. |
currency | string | Your organization’s billing currency (usd or jpy), set at org creation. usd when called without an API key. |
input_per_1m | string | Price per 1 million input tokens, in your billing currency. Numeric stored as a string for precision. |
output_per_1m | string | Price per 1 million output tokens, in your billing currency. |
GET /v1/models with an anthropic-version header returns the Anthropic shape — data[].display_name, created_at, has_more, first_id, last_id. Pricing and capabilities are only on the OpenAI surface.