GLM-5.3-Flash
Z.ai’s fast, low-cost GLM, and the first in the series to see images and video.
About GLM-5.3-Flash
GLM-5.3-Flash is the volume tier of GLM-5.3 and the first GLM-5 model to take images and video as well as text. Before Z.ai put its name on it, it ran anonymously as “Ox Alpha” and drew more than 50 trillion tokens of traffic in five days. Z.ai reports 63.4% on DeepSWE, well up on GLM-5.2’s 46.2%.
Its weights are open under MIT. Like its bigger sibling, it always thinks before answering, which costs some speed on simple calls.
Where it’s strong
Cheap, high-volume coding agents.
Reading screenshots, documents and video inside agent loops.
Long tool-use sessions where cost per task matters.
Where to be careful
Thinking cannot be switched off.
Its scores are Z.ai’s own, apart from one independent index.
Price
For scale: 1,000 requests of about 3,000 tokens in and 1,000 out cost $1.24 — a long prompt and a page of answer each. Every rate is published and stays the same from one request to the next. See every model.
Specs
glm-5.3-flashCall it
Use the OpenAI SDK you already have — point it at Roar AI and name the model. Read the docs.
from openai import OpenAI
client = OpenAI(base_url="https://api.roar-ai.com/v1", api_key="roar_live_…")
reply = client.chat.completions.create(
model="glm-5.3-flash",
messages=[{"role": "user", "content": "Summarise this contract in plain English."}],
)Compare GLM-5.3-Flash
Questions
How much does GLM-5.3-Flash cost?
Input is $0.195 per 1M tokens; output is $0.65 per 1M tokens. For scale, 1,000 requests of about 3,000 tokens in and 1,000 out cost $1.24. The rate is published and does not change from one request to the next.
Can I use GLM-5.3-Flash alongside other models?
Yes. Every model on Roar AI is on the same API key and the same monthly invoice, so switching from GLM-5.3-Flash to another model is a change to one word in your code.
Is GLM-5.3-Flash open source?
Its weights are open under the MIT licence, so you could run it yourself. Through Roar AI you call it like any other model, with no servers to manage.
