Roar AI
DocsRequest access
All models
Alibaba · Chat

Qwen3.8-Flash

Alibaba’s fast, low-cost Qwen, and an early look at the Qwen4 design.

Input$0.1658per 1M tokens
Output$0.5193per 1M tokens
BillingOne invoice, in LKRwith every other model
Request early accessSee every model’s rate

About Qwen3.8-Flash

Qwen3.8-Flash is Alibaba’s volume model — 125 billion parameters with only 6 billion active per token, which is what makes it cheap and quick. It previews the design planned for Qwen4, and Alibaba says it cost about a ninth as much to train as Qwen3.7-Plus while beating it. Alibaba reports 62.5% on SWE-bench Pro and 84.5% on AndroidWorld, a phone-automation test.

It reads images and video, holds a million tokens, and thinks by default with adjustable effort. Reviewers’ main complaint is length: it writes far more than most models, which eats into its price advantage.

Where it’s strong

High-volume coding help at low cost.

Long video and document analysis.

Automating phone and desktop apps.

Office task automation.

Where to be careful

Very wordy, so tasks can cost more than its rates suggest.

Turning reasoning down does not reliably make long multi-turn sessions faster.

Its scores are Alibaba’s own.

Price in rupees

Input$0.1658 per 1M tokens
Output$0.5193 per 1M tokens
Cached input$0.0187 per 1M tokens
Cache write$0.26 per 1M tokens

For scale: 1,000 requests of about 3,000 tokens in and 1,000 out cost $1.02 — a long prompt and a page of answer each. Every rate is published and stays the same from one request to the next. See every model.

Specs

Made byAlibaba
ReleasedAugust 2026
TypeChat
Context window1M tokens
Longest answer131K tokens
ReadsText, Images, Video
Thinks before answeringYes, adjustable
WeightsClosed
Model idqwen3.8-flash

Using it from Sri Lanka

Qwen3.8-Flash is billed in rupees, on the same monthly invoice as every other model your team uses — no foreign card, no separate Alibaba account, and one key for all of it.

Prompts sent to Qwen3.8-Flash are processed outside Sri Lanka. If a workload has to stay in the country, Qwen3.8-27B (Colombo) runs on our own servers in Colombo.

Call it

Use the OpenAI SDK you already have — point it at Roar AI and name the model. Read the docs.

from openai import OpenAI

client = OpenAI(base_url="https://api.roar-ai.com/v1", api_key="roar_live_…")

reply = client.chat.completions.create(
    model="qwen3.8-flash",
    messages=[{"role": "user", "content": "Summarise this contract in plain English."}],
)

Questions

How much does Qwen3.8-Flash cost in Sri Lanka?

Input is $0.1658 per 1M tokens; output is $0.5193 per 1M tokens. For scale, 1,000 requests of about 3,000 tokens in and 1,000 out cost $1.02. The rate is published and does not change from one request to the next.

Can I pay for Qwen3.8-Flash in Sri Lankan rupees?

Yes. Usage is billed to one monthly invoice in rupees, together with every other model your team uses, and you can also pay by card. There is no separate Alibaba account to open and no foreign card needed.

Does Qwen3.8-Flash keep data in Sri Lanka?

No — prompts sent to Qwen3.8-Flash are processed outside Sri Lanka. If your data has to stay in the country, use Qwen3.8-27B (Colombo), which runs on our servers in Colombo.

Is Qwen3.8-Flash open source?

No. Alibaba does not publish the weights, so it is only available through an API like this one.