Closed models
Models whose makers do not publish the weights, so an API like this one is the only way in.
Anthropic’s most capable model, built for work that runs for hours and checks its own results.
The previous Opus, still available for work that was built and tested on it.
Anthropic’s recommended default: near the top on hard work, cheaper than the Opus it replaced.
Anthropic’s fast mid-tier model, now within a whisker of Opus on most work.
Anthropic’s smallest, fastest model, for chat, support and work split across many small agents.
OpenAI’s flagship, for the hardest maths, science and long agent work.
OpenAI’s flagship from July 2026, kept for long coding work built on it.
OpenAI’s current middle tier: close to Astra on agent work, at a Sol price.
The first GPT-6 middle tier, kept for work already built on it.
OpenAI’s coding specialist from early 2026, made for agents that work in a terminal.
The fast, low-cost GPT-5.6, for high-volume work.
OpenAI’s small model for classifying, extracting and ranking at scale.
OpenAI’s cheapest GPT-6, with the same million-token memory as Astra.
Google’s Pro model: reads video, audio, images and PDFs in one million-token window.
Google’s newest Flash: near-flagship coding at Flash speed.
xAI’s flagship from August 2026, strongest at graduate-level science questions.
Meta’s closed flagship for long coding sessions, with near-perfect recall across a million tokens.
The same Muse Spark 1.3 at a far lower price — because Meta may train on what you send it.
Alibaba’s fast, low-cost Qwen, and an early look at the Qwen4 design.
Upstage’s model that says when it cannot verify something, rather than filling the gap.
Open-weight models
Models whose weights are public. Often a fraction of the price for everyday work, and one of them runs in Colombo.
The small gpt-oss: o3-mini-class reasoning in a model that fits on a laptop.
OpenAI’s larger open-weight model: o4-mini-class reasoning you could also run yourself.
Google’s largest open Gemma: a 31-billion-parameter model that reads images and video, under Apache 2.0.
Meta’s dependable mid-size open model, close to Llama 3.1 405B at a sixth of the size.
Meta’s small Llama: quick and very cheap for simple, high-volume text work.
DeepSeek’s open flagship: a 1.6-trillion-parameter model under the plain MIT licence.
DeepSeek’s current default: fast, low-cost, reads images, and ahead of the bigger V4-Pro on DeepSeek’s own tests.
Alibaba’s flagship, and the first Max-tier Qwen with downloadable weights.
Alibaba’s small dense Qwen: open, multimodal and unusually strong at coding for 27 billion parameters.
Qwen3.8-27B on our own server in Colombo: your data is processed in Sri Lanka and nowhere else.
Moonshot AI’s 2.8-trillion-parameter open model, first at web research when it launched.
Z.ai’s open flagship for coding agents and security work.
Z.ai’s fast, low-cost GLM, and the first in the series to see images and video.
MiniMax’s open flagship: strong coding, a million-token context and image and video input in one model.
Xiaomi’s open coding model, built to keep going through a thousand tool calls.
Xiaomi’s open model that reads, looks and listens — text, images, audio and video.
Tencent’s open flagship, switching between quick answers and deep reasoning.
Speech, embeddings and evaluation
Transcribe audio, read text aloud, power search over your own documents, and grade answers.
Quick, low-cost speech from text, for real-time voice.
Fast speech-to-text in 99 languages, Sinhala and Tamil among them.
OpenAI’s most accurate embedding model, and the better one for languages other than English.
Low-cost embeddings for search, recommendations and finding duplicates.
A judge model: typed answers about your data, each with a probability, in well under a second.
Head to head
The pairs people actually weigh against each other, with prices side by side and a straight answer on which to use.
