Roar AI
DocsRequest access
All models
Zhipu AI · Open weights

GLM-5.3-Flash

Z.ai’s fast, low-cost GLM, and the first in the series to see images and video.

Input$0.195per 1M tokens
Output$0.65per 1M tokens
BillingOne invoice, in LKRwith every other model
Request early accessSee every model’s rate

About GLM-5.3-Flash

GLM-5.3-Flash is the volume tier of GLM-5.3 and the first GLM-5 model to take images and video as well as text. Before Z.ai put its name on it, it ran anonymously as “Ox Alpha” and drew more than 50 trillion tokens of traffic in five days. Z.ai reports 63.4% on DeepSWE, well up on GLM-5.2’s 46.2%.

Its weights are open under MIT. Like its bigger sibling, it always thinks before answering, which costs some speed on simple calls.

Where it’s strong

Cheap, high-volume coding agents.

Reading screenshots, documents and video inside agent loops.

Long tool-use sessions where cost per task matters.

Where to be careful

Thinking cannot be switched off.

Its scores are Z.ai’s own, apart from one independent index.

Price in rupees

Input$0.195 per 1M tokens
Output$0.65 per 1M tokens
Cached input$0.039 per 1M tokens

For scale: 1,000 requests of about 3,000 tokens in and 1,000 out cost $1.24 — a long prompt and a page of answer each. Every rate is published and stays the same from one request to the next. See every model.

Specs

Made byZhipu AI
ReleasedAugust 2026
TypeChat
Context window1M tokens
Longest answer128K tokens
ReadsText, Images, Video
Thinks before answeringAlways on
WeightsOpen (MIT)
Size320B total, 18B active
Model idglm-5.3-flash

Using it from Sri Lanka

GLM-5.3-Flash is billed in rupees, on the same monthly invoice as every other model your team uses — no foreign card, no separate Zhipu AI account, and one key for all of it.

Prompts sent to GLM-5.3-Flash are processed outside Sri Lanka. If a workload has to stay in the country, Qwen3.8-27B (Colombo) runs on our own servers in Colombo.

Call it

Use the OpenAI SDK you already have — point it at Roar AI and name the model. Read the docs.

from openai import OpenAI

client = OpenAI(base_url="https://api.roar-ai.com/v1", api_key="roar_live_…")

reply = client.chat.completions.create(
    model="glm-5.3-flash",
    messages=[{"role": "user", "content": "Summarise this contract in plain English."}],
)

Questions

How much does GLM-5.3-Flash cost in Sri Lanka?

Input is $0.195 per 1M tokens; output is $0.65 per 1M tokens. For scale, 1,000 requests of about 3,000 tokens in and 1,000 out cost $1.24. The rate is published and does not change from one request to the next.

Can I pay for GLM-5.3-Flash in Sri Lankan rupees?

Yes. Usage is billed to one monthly invoice in rupees, together with every other model your team uses, and you can also pay by card. There is no separate Zhipu AI account to open and no foreign card needed.

Does GLM-5.3-Flash keep data in Sri Lanka?

No — prompts sent to GLM-5.3-Flash are processed outside Sri Lanka. If your data has to stay in the country, use Qwen3.8-27B (Colombo), which runs on our servers in Colombo.

Is GLM-5.3-Flash open source?

Its weights are open under the MIT licence, so you could run it yourself. Through Roar AI you call it like any other model, with no servers to manage.