More than chat,
on the same key
Everything below is the OpenAI shape you already write against. Swap the base URL and your existing code keeps working.
Give it to the team
without giving up the reins
The part that decides whether a company can actually let people use this: what leaves, what it costs, and who can do what.
Three cheaper ways to run the same model
None of these is a different model or a worse one. They are the same call, bought differently.
Repeat a long prefix — a system prompt, a contract, a codebase — and those tokens bill at the cached rate. Nothing to switch on.
Send a file of requests, collect results within 24 hours, at roughly half the rate. It runs on the provider’s own batch tier, so the discount is a real saving rather than a gift.
Faster and steadier on the models that offer it, at a premium. Worth it when somebody is sitting in front of the answer.
Not every model offers every tier. The rate card marks the ones that do.
Agents can use it directly
Point an MCP client here with the same key. It only sees the tools that key is allowed to use.
