Skip to content

AI Gateway Models

Varity Team Core Contributors Updated August 2026

The catalog is served live. This page is a snapshot of it; the endpoint is the source of truth:

Terminal window
curl https://varity.app/v1/models

GET /v1/models is public and needs no API key. It returns an OpenAI-shaped { "object": "list", "data": [...] } response. Varity adds fields OpenAI does not return: display_name, capabilities, context_window, max_completion_tokens, privacy, privacy_routes, availability, and customer_pricing.

42 models were listed at the time of writing, all reporting availability.status: "available" and chat_completions: true.

Pass the id column verbatim as the model field of a chat-completions request. Note that most ids are opaque model-<hash> strings rather than vendor-style names; the display_name is a label, not an identifier.

Prices are USD per million tokens, taken from each model’s customer_pricing object.

Model IDNameContext windowMax completion tokensTool callsStreamingInput / 1MOutput / 1M
model-23ceb717e8b9fbdf4170804cDeepSeek V3.2160,00032,768NoYes$0.3465$0.504
model-01e5cfb9cd019cee62a62b13Gemma 4 Uncensored256,0008,192YesYes$0.170625$0.525
model-e7a95bbb411e58c5b94bd949GLM 4.6198,00016,384NoYes$0.4515$1.8375
model-84999d7f31100ae1be7e4ee9GLM 4.7198,00016,384NoYes$0.5775$2.7825
model-4d0399841ad2104adac8673fGLM 4.7 Flash128,00016,384NoYes$0.063$0.42
model-8f0e33a63784a2569979827dGLM 4.7 Flash Heretic200,00024,000NoYes$0.0735$0.42
model-975fe730d6699badb8f13ca0GLM 5198,00032,000NoYes$1.05$3.36
glm-5-1GLM 5.1200,00080,000NoYes$1.617$5.082
glm-5-2GLM 5.21,000,000131,072NoYes$1.47$4.62
model-afb55f3c0379fce9fb050497Google Gemma 3 27B Instruct198,00016,384YesYes$0.126$0.21
model-592b84ed19d414fc58b3c1b8Google Gemma 4 26B A4B Instruct256,0008,192NoYes$0.1365$0.42
model-da12ebb256e118bd1b7f4c86Google Gemma 4 31B Instruct256,0008,192NoYes$0.126$0.378
model-1bcd5b1a9fc32ffb985afc6bGrok 4.202,000,000128,000YesYes$1.491$2.9715
model-13cb7872bbc7954e8ae400d8Grok 4.20 Multi-Agent2,000,000128,000NoYes$1.491$2.9715
model-4e27cf37044481968f24c858Grok 4.31,000,00032,000YesYes$1.491$2.9715
model-7ac496866ed390aa843e23eaGrok 4.5500,00032,000YesYes$2.3835$7.14
model-0bf4d9df75eb028b048a51edGrok Build 0.1256,00065,536YesYes$1.05$2.1
model-b5bcc7a2b398465798a0759aHermes 3 Llama 3.1 405b128,00016,384NoYes$1.155$3.15
model-35fec70a1356b5981dba26d9Inkling524,28865,536NoYes$1.3125$5.315625
model-547bc7d442bda226b1ac0d52Kimi K2.5256,00065,536NoYes$0.588$3.675
model-932fa0c5bc52d3ea2e193584Kimi K2.6256,00065,536NoYes$0.7875$3.675
model-28805f587470f2b9e31381f8Kimi K2.7 Code256,00065,536NoYes$0.7875$3.675
model-93e21d329a7b4dd9514d9c9fKimi K31,000,000131,072NoYes$3.9375$19.6875
model-c09137265edddbdec9dfda8bLlama 3.2 3B128,0004,096YesYes$0.1575$0.63
llama-3-3-70bLlama 3.3 70B128,0004,096YesYes$0.735$2.94
model-36acfa144b41de8ac1c27944MiniMax M2.5198,00032,768NoYes$0.2835$0.9975
model-c783b8f80c60041a340bdb07MiniMax M2.7198,00032,768NoYes$0.39375$1.575
model-ca8c32ffa3e6ec1adee98495Mistral Small 3.2 24B Instruct256,00016,384YesYes$0.0984375$0.2625
model-305938ab9ebc39e1beb95ac4Mistral Small 4256,00065,536NoYes$0.196875$0.7875
model-adeed85bab462c43614e8e1cNVIDIA Nemotron 3 Nano 30B128,00016,384NoYes$0.07875$0.315
model-2207a53b5a3f856264ce7753NVIDIA Nemotron 3 Ultra256,00032,768NoYes$0.65625$3.28125
model-94ab4b38ef8b46c8cdf5ffa0OpenAI GPT OSS 120B128,00016,384NoYes$0.0735$0.315
model-ab4d70a40590dc6ed07d5187Qwen 3 235B A22B Instruct 2507128,00016,384NoYes$0.1575$0.7875
model-366edcdf75d467e68639879bQwen 3 235B A22B Thinking 2507128,00016,384NoYes$0.4725$3.675
model-bbb8171a9ad936cd006e9640Qwen 3 Coder 480B Turbo256,00065,536NoYes$0.3675$1.575
model-2f3b522ed43cc37cb7cac581Qwen 3 Next 80b256,00016,384NoYes$0.3675$1.995
qwen-3-5-35b-a3bQwen 3.5 35B A3B256,00016,384NoYes$0.328125$1.3125
model-e838b0fde95335963c1bc9d4Qwen 3.5 9B256,00032,768NoYes$0.105$0.1575
model-f660e96789aa9606296abe39Qwen 3.6 27B256,00065,536NoYes$0.34125$3.4125
model-be42cfa83562024c857a01fbQwen3 VL 235B128,00016,384NoYes$0.2205$1.995
model-38c8e473186cf182a89ee7ffRole Play Uncensored128,0004,096YesYes$0.525$2.1
model-665e90a28adb3a74b27bb064Uncensored 1.2128,0008,192YesYes$0.21$0.945

Each entry in data carries these Varity fields:

FieldTypeNotes
idstringPass this as model in a request
display_namestringHuman label only, never an identifier
capabilities.chat_completionsbooleantrue on all 42 models
capabilities.tool_callsbooleantrue on 11 of 42 models
capabilities.streamingbooleantrue on all 42 models
capabilities.streaming_tool_callsbooleantrue on 7 of 42 models
context_windowintegerTotal tokens, input plus output
max_completion_tokensintegerCeiling on generated tokens
privacystring"private" on every model in the catalog
privacy_routesarray["private"] on every model in the catalog
availability.statusstring"available" on every model, with an observed_at timestamp
customer_pricingobjectSee below
refreshed_atstringWhen the gateway last refreshed the entry

Every model returns customer_pricing.status: "configured" with concrete numbers:

"customer_pricing": {
"status": "configured",
"version": "2026-07-28.v3.d74e3f5155b10314",
"currency": "usd",
"unit": "per_million_tokens",
"input": 1.47,
"output": 4.62,
"cached_input": 1.47
}

cached_input equals input on all 42 models. Cached input tokens are not discounted today. Always read unit and currency rather than assuming; the version string identifies the price set the numbers came from.

Across the catalog, input prices run from $0.063 to $3.9375 per million tokens and output prices from $0.1575 to $19.6875 per million tokens.

Context windows range from 128,000 to 2,000,000 tokens. max_completion_tokens is a separate, much smaller ceiling on any single response, and it goes as low as 4,096 on some models. A large context window does not imply a large answer.

  • Need tool calling? Only 11 models support it, and only 7 support tool calls while streaming. See Compatibility for the exact lists.
  • Need a very large context? model-13cb7872bbc7954e8ae400d8 (Grok 4.20 Multi-Agent) and model-1bcd5b1a9fc32ffb985afc6b (Grok 4.20) both report 2,000,000 tokens.
  • Cost sensitive? model-4d0399841ad2104adac8673f (GLM 4.7 Flash) is the cheapest input price in the catalog at $0.063 per million input tokens.