Roar AI
DocsRequest access
All models
Google · Open weights

Gemma 4 31B

Google’s largest open Gemma: a 31-billion-parameter model that reads images and video, under Apache 2.0.

Input$0.182per 1M tokens
Output$0.52per 1M tokens
BillingOne invoice, in LKRwith every other model
Request early accessSee every model’s rate

About Gemma 4 31B

Gemma 4 31B is the biggest model in Google’s open Gemma family, built from the same research as Gemini 3. It is the first Gemma released under the standard Apache 2.0 licence rather than Google’s own terms, and at launch it ranked third among open models on the LMArena text leaderboard. Google’s model card reports 85.2% on MMLU-Pro and 84.3% on GPQA Diamond.

It reads images and video as well as text, with a 256K-token context, and thinks before answering when you switch that on. As a dense 31B model it is small next to the trillion-parameter open models, which shows on long, hard agent tasks — but it is quick and inexpensive for document and image work.

Where it’s strong

Reading documents, forms and images.

Everyday chat and drafting at a low price.

Maths for its size — 89.2% on AIME 2026, by Google’s count.

Where to be careful

Its knowledge stops at January 2025, older than most 2026 models.

Behind the large open models on hard, multi-step coding agents.

The 31B version does not take audio.

Price in rupees

Input$0.182 per 1M tokens
Output$0.52 per 1M tokens
Batch input$0.091 per 1M tokens
Batch output$0.26 per 1M tokens

For scale: 1,000 requests of about 3,000 tokens in and 1,000 out cost $1.07 — a long prompt and a page of answer each. Every rate is published and stays the same from one request to the next. See every model.

Specs

Made byGoogle
ReleasedApril 2026
TypeChat
Context window256K tokens
ReadsText, Images, Video
Thinks before answeringYes, adjustable
WeightsOpen (Apache 2.0)
Size31B, dense
Model idgemma-4-31b-it

Using it from Sri Lanka

Gemma 4 31B is billed in rupees, on the same monthly invoice as every other model your team uses — no foreign card, no separate Google account, and one key for all of it.

Prompts sent to Gemma 4 31B are processed outside Sri Lanka. If a workload has to stay in the country, Qwen3.8-27B (Colombo) runs on our own servers in Colombo.

Google says Gemma 4 was trained on more than 140 languages, but Sinhala and Tamil are not among the 35 it names as supported out of the box. Test it on your own messages first.

Call it

Use the OpenAI SDK you already have — point it at Roar AI and name the model. Read the docs.

from openai import OpenAI

client = OpenAI(base_url="https://api.roar-ai.com/v1", api_key="roar_live_…")

reply = client.chat.completions.create(
    model="gemma-4-31b-it",
    messages=[{"role": "user", "content": "Summarise this contract in plain English."}],
)

Questions

How much does Gemma 4 31B cost in Sri Lanka?

Input is $0.182 per 1M tokens; output is $0.52 per 1M tokens. For scale, 1,000 requests of about 3,000 tokens in and 1,000 out cost $1.07. The rate is published and does not change from one request to the next.

Can I pay for Gemma 4 31B in Sri Lankan rupees?

Yes. Usage is billed to one monthly invoice in rupees, together with every other model your team uses, and you can also pay by card. There is no separate Google account to open and no foreign card needed.

Does Gemma 4 31B keep data in Sri Lanka?

No — prompts sent to Gemma 4 31B are processed outside Sri Lanka. If your data has to stay in the country, use Qwen3.8-27B (Colombo), which runs on our servers in Colombo.

Is Gemma 4 31B open source?

Its weights are open under the Apache 2.0 licence, so you could run it yourself. Through Roar AI you call it like any other model, with no servers to manage.