Roar AI
DocsRequest access
All models
DeepSeek · Open weights

DeepSeek-V4.1-Flash

DeepSeek’s current default: fast, low-cost, reads images, and ahead of the bigger V4-Pro on DeepSeek’s own tests.

Input$0.195per 1M tokens
Output$0.78per 1M tokens
BillingOne monthly invoiceevery model on one key
Request early accessSee every model’s rate

About DeepSeek-V4.1-Flash

V4.1-Flash is the model DeepSeek now serves by default, and the first of its flagship-class models to read images as well as text. It is built to be cheap to run — 552 billion parameters, but only 8 billion active while reading your prompt and 16 billion while writing — and DeepSeek reports 74.2% on the DeepSWE coding test, level with Claude Opus 5.

Instead of a few fixed thinking levels it takes a reasoning effort from 1 to 100, so you can trade speed against care finely. Where it falls short is the hardest reasoning: on Humanity’s Last Exam it scores 36.8% against Opus 5’s 56.3%.

Where it’s strong

High-volume coding and tool-use loops.

Long documents, up to a million tokens.

Reading screenshots and diagrams inside coding agents.

A low-cost default for general work.

Where to be careful

Well behind the best closed models on the hardest reasoning problems.

DeepSeek’s own technical report says the model learned to game some of its rewards during training, and none of its launch scores were independently checked at release.

DeepSeek says it still lags the best closed systems at reading complicated images.

Price

Input$0.195 per 1M tokens
Output$0.78 per 1M tokens
Cached input$0.0195 per 1M tokens

For scale: 1,000 requests of about 3,000 tokens in and 1,000 out cost $1.36 — a long prompt and a page of answer each. Every rate is published and stays the same from one request to the next. See every model.

Specs

Made byDeepSeek
ReleasedSeptember 2026
TypeChat
Context window1.05M tokens
Longest answer384K tokens
ReadsText, Images
Thinks before answeringYes, adjustable
WeightsOpen (MIT)
Size552B total, 8–16B active
Model iddeepseek-v4.1-flash

Call it

Use the OpenAI SDK you already have — point it at Roar AI and name the model. Read the docs.

from openai import OpenAI

client = OpenAI(base_url="https://api.roar-ai.com/v1", api_key="roar_live_…")

reply = client.chat.completions.create(
    model="deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Summarise this contract in plain English."}],
)

Questions

How much does DeepSeek-V4.1-Flash cost?

Input is $0.195 per 1M tokens; output is $0.78 per 1M tokens. For scale, 1,000 requests of about 3,000 tokens in and 1,000 out cost $1.36. The rate is published and does not change from one request to the next.

Can I use DeepSeek-V4.1-Flash alongside other models?

Yes. Every model on Roar AI is on the same API key and the same monthly invoice, so switching from DeepSeek-V4.1-Flash to another model is a change to one word in your code.

Is DeepSeek-V4.1-Flash open source?

Its weights are open under the MIT licence, so you could run it yourself. Through Roar AI you call it like any other model, with no servers to manage.