Roar AI
DocsRequest access
All models
Google · Chat

Gemini 3.8 Flash

Google’s newest Flash: near-flagship coding at Flash speed.

Input$0.8625per 1M tokens
Output$4.3125per 1M tokens
BillingOne invoice, in LKRwith every other model
Request early accessSee every model’s rate

About Gemini 3.8 Flash

Gemini 3.8 Flash is the third Flash release in six weeks and Google’s most capable so far. On the DeepSWE coding test it scored 73.7%, just under Claude Opus 5 and ahead of GPT-5.6 Sol, and Google reports wins over Opus 5 on finance and legal agent tests. It reads the same mix as Gemini Pro — text, images, audio, video and PDFs — in a million-token window.

Once it starts, it is fast: about 300 tokens a second in independent tests. Starting is the slow part — about 13 seconds to the first token, against a median of 3 — and it is wordy, so tasks can cost more than its rates suggest.

Where it’s strong

Coding agents at a mid-range price.

Working through large volumes of video, audio and PDFs.

Finance and legal analysis agents.

Batch jobs where total throughput matters more than the first word.

Where to be careful

Slow to begin answering, which makes it a poor fit for live chat.

Verbose: one review measured tasks costing about 40% more than on Gemini 3.7 Flash at the same rates.

Google’s model card notes its safety behaviour in languages other than English slipped compared with 3.7 Flash.

Price in rupees

Input$0.8625 per 1M tokens
Output$4.3125 per 1M tokens
Cached input$0.0862 per 1M tokens
Batch input$0.4313 per 1M tokens
Batch output$2.1563 per 1M tokens

For scale: 1,000 requests of about 3,000 tokens in and 1,000 out cost $6.90 — a long prompt and a page of answer each. Every rate is published and stays the same from one request to the next. See every model.

Specs

Made byGoogle
ReleasedSeptember 2026
TypeChat
Context window1.05M tokens
Longest answer66K tokens
ReadsText, Images, Audio, Video, PDFs
Thinks before answeringAlways on
WeightsClosed
Model idgemini-3.8-flash

Using it from Sri Lanka

Gemini 3.8 Flash is billed in rupees, on the same monthly invoice as every other model your team uses — no foreign card, no separate Google account, and one key for all of it.

Prompts sent to Gemini 3.8 Flash are processed outside Sri Lanka. If a workload has to stay in the country, Qwen3.8-27B (Colombo) runs on our own servers in Colombo.

Google’s model card records weaker safety behaviour outside English than the previous Flash, and Google publishes no Sinhala or Tamil results. If you serve Sinhala- or Tamil-speaking users, test with real messages and keep your own checks in place.

Call it

Use the OpenAI SDK you already have — point it at Roar AI and name the model. Read the docs.

from openai import OpenAI

client = OpenAI(base_url="https://api.roar-ai.com/v1", api_key="roar_live_…")

reply = client.chat.completions.create(
    model="gemini-3.8-flash",
    messages=[{"role": "user", "content": "Summarise this contract in plain English."}],
)

Questions

How much does Gemini 3.8 Flash cost in Sri Lanka?

Input is $0.8625 per 1M tokens; output is $4.3125 per 1M tokens. For scale, 1,000 requests of about 3,000 tokens in and 1,000 out cost $6.90. The rate is published and does not change from one request to the next.

Can I pay for Gemini 3.8 Flash in Sri Lankan rupees?

Yes. Usage is billed to one monthly invoice in rupees, together with every other model your team uses, and you can also pay by card. There is no separate Google account to open and no foreign card needed.

Does Gemini 3.8 Flash keep data in Sri Lanka?

No — prompts sent to Gemini 3.8 Flash are processed outside Sri Lanka. If your data has to stay in the country, use Qwen3.8-27B (Colombo), which runs on our servers in Colombo.

Is Gemini 3.8 Flash open source?

No. Google does not publish the weights, so it is only available through an API like this one.