Hosting
Charged per day, for as long as they exist, plus what they use.
AI models
Per million tokens. Open the rows tagged cache, batch or priority to see the cheaper ways to run the same one.
Repeat a long prefix — a system prompt, a contract, a codebase — and those tokens bill at the cached rate instead of the full one. It applies automatically; there is nothing to switch on.
Send a file of requests and collect the results within 24 hours, for roughly half the usual rate. Suited to work that doesn’t need an answer while someone waits — overnight classification, backfills, evaluations.
Faster, more consistent service on the models that offer it, at a premium over the standard rate. Worth it for something a person is sitting in front of.
Model rates are in USD per million units and are billed to your monthly invoice in rupees at the prevailing rate. Embeddings and transcription bill on input only. Last updated 7 September 2026.
Need volume pricing, a model that isn’t listed, or a rate held for a budget cycle? Talk to us — we’re in Colombo.
