Model catalog

30 frontier models. 11 providers. One API.

Every id pinned and price-verified against the provider — never floating -latest aliases. The rate you see is the rate you pay, all-in. Retired models remap to successors — they never 404.

Routing aliases

Or skip model names entirely

Request a capability instead of a model. Each alias is a curated cross-provider chain with automatic failover, and your org's routing policy — data residency, provider and model blocks, PHI allowlists — is enforced per candidate. The response tells you what served via x-conduix-alias and x-conduix-model-served.

General

  • frontier-bestHighest-quality frontier model, cross-provider failover.

    GPT-5.5 → Claude Opus 4.8 → Gemini 3.1 Pro

  • frontier-fastFrontier-quality with production latency.

    Claude Sonnet 5 → GPT-5.4 → Gemini 3.5 Flash → Grok 4.3

  • frontier-budgetBest quality per dollar for high-volume workloads.

    GPT-5.4 mini → Claude Haiku 4.5 → Gemini 3.1 Flash-Lite → DeepSeek V4 Flash

  • enterpriseConservative chain: US frontier providers with enterprise SLAs.

    Claude Opus 4.8 → GPT-5.5 → GPT-5.4

Coding

  • code-premiumStrongest coding models available.

    GPT-5.3 Codex → Claude Opus 4.8 → Claude Sonnet 5

  • code-balancedProduction coding default — quality/cost balance.

    Claude Sonnet 5 → GPT-5.3 Codex → Codestral

  • code-fastLow-latency completions and quick edits.

    GPT-5.4 mini → Codestral → Claude Haiku 4.5

  • code-openOpen-weights coding models only.

    Kimi K2.7 Code (Together) → Qwen 3.7 Plus (Together)

  • agentic-codeLong-horizon agentic coding (tool loops, large diffs).

    GPT-5.3 Codex → Claude Sonnet 5 → Kimi K2.7 Code (Together)

Reasoning

  • reasoning-bestDeepest deliberate reasoning available.

    Claude Fable 5 → GPT-5.5 → DeepSeek V4 Pro

  • reasoning-fastThinking-capable models at interactive latency.

    Gemini 3.5 Flash → Grok 4.3 → GPT-5.4

  • reasoning-budgetChain-of-thought quality without frontier pricing.

    DeepSeek V4 Pro → DeepSeek V4 Flash → GPT-5.4 mini

Vision & documents

  • vision-bestStrongest image + document understanding.

    Gemini 3.1 Pro → GPT-5.5 → Claude Opus 4.8

  • vision-fastFast multimodal understanding.

    Gemini 3.5 Flash → GPT-5.4 mini → Gemma 4 31B (Together)

  • document-aiLong-document analysis (contracts, filings, PDFs).

    Gemini 3.1 Pro → Gemini 3.5 Flash → Claude Sonnet 5

Coming soon: ocr · speech-best · speech-fast · stt · tts · embedding-best · embedding-fast · image-best · image-fast · video-best · rerank-best — these aliases ship with their modality endpoints.

Browse the catalog

30 of 30 models

GPT-5.5

Frontier
$6.25 in · $37.50 out / 1M tokens
OpenAIChat1.05M contextgpt-5.5

OpenAI flagship. 1M context, adjustable reasoning effort, top-tier agentic work.

Tool callingVisionJSON modePrompt cache

Claude Fable 5

Frontier
$12.50 in · $62.50 out / 1M tokens
AnthropicReasoning1M contextclaude-fable-5

Anthropic’s most capable model. Thinking always on — the gateway strips sampling parameters this model rejects.

Tool callingVisionReasoningJSON modePrompt cache

Claude Opus 4.8

Frontier
$6.25 in · $31.25 out / 1M tokens
AnthropicChat1M contextclaude-opus-4-8

Anthropic flagship for complex writing, analysis, and agentic coding. 1M context.

Tool callingVisionJSON modePrompt cache

Gemini 3.1 Pro

Frontier
$2.50 in · $15.00 out / 1M tokens
Google GeminiReasoning1.05M contextgemini-3.1-pro-preview

Google’s top model — strongest multimodal + long-document reasoning. Google designates this tier "Preview".

Tool callingVisionReasoningJSON mode

Grok 4.3

Frontier
$1.56 in · $3.13 out / 1M tokens
xAIReasoning1M contextgrok-4.3

xAI flagship. 1M context; reasoning effort replaces the retired fast/mini tiers (set effort low for speed).

Tool callingVisionReasoningJSON modePrompt cache

GPT-5.4

Premium
$3.13 in · $18.75 out / 1M tokens
OpenAIChat1.05M contextgpt-5.4

OpenAI premium workhorse — 1M context at workhorse pricing.

Tool callingVisionJSON modePrompt cache

GPT-5.3 Codex

Premium
$2.19 in · $17.50 out / 1M tokens
OpenAICoding400K contextgpt-5.3-codex

OpenAI’s most capable agentic coding model.

Tool callingVisionJSON modePrompt cache

Claude Sonnet 5

Premium
$2.50 in · $12.50 out / 1M tokens
AnthropicChat1M contextclaude-sonnet-5

Anthropic workhorse — excellent coding + long-context. Introductory provider pricing until 2026-08-31 (operator repricing due 2026-09-01).

Tool callingVisionJSON modePrompt cache

DeepSeek V4 Pro

PremiumOpen weightsapac
$0.54 in · $1.09 out / 1M tokens
DeepSeekReasoning1M contextdeepseek-v4-pro

Frontier-class open-weights reasoning at budget pricing. 1M context.

Tool callingReasoningJSON mode

Mistral Medium 3.5

PremiumOpen weightseu
$1.88 in · $9.38 out / 1M tokens
MistralChat256K contextmistral-medium-2604

Mistral’s most capable model — EU residency with per-request reasoning effort.

Tool callingVisionJSON mode

Kimi K2.6 (Together)

PremiumOpen weights
$1.50 in · $5.63 out / 1M tokens
TogetherChat262K contextmoonshotai/Kimi-K2.6

Open-weights frontier chat — top-tier agentic tool use.

Tool callingJSON mode

Kimi K2.7 Code (Together)

PremiumOpen weights
$1.19 in · $5.00 out / 1M tokens
TogetherCoding262K contextmoonshotai/Kimi-K2.7-Code

Current open-weights coding flagship — agentic repo-scale work.

Tool callingJSON mode

GLM 5.2 (Together)

PremiumOpen weights
$1.75 in · $5.50 out / 1M tokens
TogetherChat262K contextzai-org/GLM-5.2

Open-weights frontier chat — strong tool calling + multilingual.

Tool callingJSON mode

GLM 5.2 (Fireworks)

PremiumOpen weights
$1.75 in · $5.50 out / 1M tokens
FireworksChat1M contextaccounts/fireworks/models/glm-5p2

Same GLM 5.2 via Fireworks — 1M-context alternate route.

Tool callingJSON mode

DeepSeek V4 Pro (Azure)

PremiumOpen weights
$2.17 in · $4.35 out / 1M tokens
Azure AI FoundryReasoning128K contextdeepseek-v4-pro-azure

DeepSeek V4 Pro hosted on Azure AI Foundry — Global Standard tier, US region. Distinct from the direct `deepseek-v4-pro` route (APAC region, direct DeepSeek): same model, hosted in Microsoft-managed US infrastructure.

ReasoningJSON mode

GPT-5.4 mini

Mid
$0.94 in · $5.63 out / 1M tokens
OpenAIChat400K contextgpt-5.4-mini

Cheap and capable — default for high-volume routine tasks.

Tool callingVisionJSON modePrompt cache

Claude Haiku 4.5

Mid
$1.25 in · $6.25 out / 1M tokens
AnthropicChat200K contextclaude-haiku-4-5-20251001

Fast, cheap Claude with 200k context. Good fit for classification + summarisation.

Tool callingVisionJSON modePrompt cache

Gemini 3.5 Flash

Mid
$1.88 in · $11.25 out / 1M tokens
Google GeminiChat1.05M contextgemini-3.5-flash

Fast thinking-capable Gemini. Great for long-context, high-throughput multimodal pipelines.

Tool callingVisionJSON mode

Mistral Large 3

MidOpen weightseu
$0.63 in · $1.88 out / 1M tokens
MistralChat256K contextmistral-large-2512

Open-weights (Apache 2.0) EU flagship — strong general chat at mid-tier pricing.

Tool callingVisionJSON mode

Codestral

Mideu
$0.38 in · $1.13 out / 1M tokens
MistralCoding256K contextcodestral-2508

EU-resident coding specialist — completion + fill-in-the-middle.

Tool callingJSON mode

Qwen 3.7 Plus (Together)

MidOpen weights
$0.40 in · $1.60 out / 1M tokens
TogetherChat1M contextQwen/Qwen3.7-Plus

Open-weights 1M-context workhorse — strong multilingual + tool use.

Tool callingJSON mode

Amazon Nova Pro

Mid
$1.00 in · $4.00 out / 1M tokens
Amazon BedrockChat300K contextamazon.nova-pro-v1:0

Amazon Nova Pro on AWS Bedrock (Converse API), served from AWS us-east-1 at standard on-demand pricing. Text in/out with streaming and tool calling.

Tool callingJSON mode

GPT-5.4 nano

Budget
$0.25 in · $1.56 out / 1M tokens
OpenAIChat400K contextgpt-5.4-nano

Lowest-cost OpenAI tier — classification, routing, bulk processing.

Tool callingVisionJSON modePrompt cache

Gemini 3.1 Flash-Lite

Budget
$0.31 in · $1.88 out / 1M tokens
Google GeminiChat1.05M contextgemini-3.1-flash-lite

Cheapest Gemini tier — bulk multimodal processing with 1M context.

Tool callingVisionJSON mode

DeepSeek V4 Flash

BudgetOpen weightsapac
$0.17 in · $0.35 out / 1M tokens
DeepSeekChat1M contextdeepseek-v4-flash

Cheapest strong general-purpose model — 1M context, dual thinking/non-thinking modes.

Tool callingJSON mode

Mistral Small 4

BudgetOpen weightseu
$0.19 in · $0.75 out / 1M tokens
MistralChat256K contextmistral-small-2603

EU-resident budget tier with built-in vision. Classification + light generation.

Tool callingVisionJSON mode

GPT-OSS 120B (Groq)

BudgetOpen weights
$0.19 in · $0.75 out / 1M tokens
GroqChat131K contextopenai/gpt-oss-120b

Open-weights workhorse at ~500 tokens/sec — Groq’s mainline model.

Tool callingJSON mode

GPT-OSS 20B (Groq)

BudgetOpen weights
$0.094 in · $0.38 out / 1M tokens
GroqChat131K contextopenai/gpt-oss-20b

Fastest cheap inference (~1,000 tokens/sec) — drafts, classification, agent loops.

Tool callingJSON mode

Llama 4 Scout (Groq)

BudgetOpen weights
$0.14 in · $0.42 out / 1M tokens
GroqChat131K contextmeta-llama/llama-4-scout-17b-16e-instruct

Multimodal open-weights Llama at ~600 tokens/sec — vision + tools at budget price.

Tool callingVisionJSON mode

Gemma 4 31B (Together)

BudgetOpen weights
$0.49 in · $1.21 out / 1M tokens
TogetherChat262K contextgoogle/gemma-4-31B-it

Light open-weights multimodal — cost-sensitive vision tasks.

Tool callingVisionJSON mode

Coming soon — new modalities

Embeddings, image, speech, video, OCR, and rerank models are staged in the catalog but excluded from the API until their endpoints ship — we don't sell dead SKUs. They become routable the day the endpoint goes live.

Embeddings

Coming soon
  • text-embedding-3-largeOpenAI
  • text-embedding-3-smallOpenAI
  • Gemini Embedding 2Google Gemini
  • Mistral EmbedMistral
  • Voyage 4 LargeVoyage AI
  • Jina Embeddings v4Jina AI
  • Cohere Embed 4Cohere

Rerank

Coming soon
  • Cohere Rerank 3.5Cohere
  • Voyage Rerank 2.5Voyage AI
  • Jina Reranker v3Jina AI

Image

Coming soon
  • GPT Image 2OpenAI
  • Gemini 3 Pro ImageGoogle Gemini
  • FLUX.2 [pro]Black Forest Labs
  • Ideogram 4.0Ideogram
  • Grok Imagine (Quality)xAI

Speech

Coming soon
  • Whisper Large v3 (Groq)Groq
  • GPT-4o TranscribeOpenAI
  • ElevenLabs Scribe v2ElevenLabs
  • ElevenLabs Eleven v3ElevenLabs
  • Voxtral SmallMistral

OCR

Coming soon
  • Mistral OCR 4Mistral

Video

Coming soon
  • Veo 3.1Google Gemini
  • Grok Imagine Video 1.5xAI

Realtime

Coming soon
  • GPT Realtime 2OpenAI
  • Gemini 3.1 Flash LiveGoogle Gemini

Chat

Coming soon
  • Kimi K2.6 (first-party)Moonshot AI
  • Cohere Command ACohere

Coding

Coming soon
  • Grok Build 0.1xAI

Route everything through one API.

Point your existing OpenAI or Anthropic SDK at Conduix and get this whole catalog — with governance, failover, and spend controls built in.