Roar AI
DocsRequest access
All models
Meta · Open weights

Llama 3.3 70B

Meta’s dependable mid-size open model, close to Llama 3.1 405B at a sixth of the size.

Input$0.3809per 1M tokens
Output$2.9289per 1M tokens
BillingOne invoice, in LKRwith every other model
Request early accessSee every model’s rate

About Llama 3.3 70B

Llama 3.3 70B was Meta’s answer to an expensive problem: it comes close to the much larger Llama 3.1 405B on key tests at about a sixth of the size. It is a plain, predictable model — no reasoning step, so answers start quickly — that follows instructions very well (92.1% on IFEval) and writes clean code for its class (88.4% on HumanEval).

It has been around since December 2024, so it is well understood and widely tested. That makes it a steady choice for drafting, summarising and answering from your own documents, where consistent behaviour matters more than the newest benchmark score.

Where it’s strong

Drafting and editing.

Answering questions from your own documents.

Assistants that need fast, predictable replies.

Where to be careful

Its knowledge stops at December 2023.

No reasoning mode: far behind 2026 reasoning models on hard science and maths (50.5% on GPQA Diamond).

Text only, under Meta’s own licence.

Price in rupees

Input$0.3809 per 1M tokens
Output$2.9289 per 1M tokens

For scale: 1,000 requests of about 3,000 tokens in and 1,000 out cost $4.07 — a long prompt and a page of answer each. Every rate is published and stays the same from one request to the next. See every model.

Specs

Made byMeta
ReleasedDecember 2024
TypeChat
Context window128K tokens
ReadsText
Thinks before answeringNo
WeightsOpen (Llama 3.3 Community License)
Size70B, dense
Model idllama-3.3-70b

Using it from Sri Lanka

Llama 3.3 70B is billed in rupees, on the same monthly invoice as every other model your team uses — no foreign card, no separate Meta account, and one key for all of it.

Prompts sent to Llama 3.3 70B are processed outside Sri Lanka. If a workload has to stay in the country, Qwen3.8-27B (Colombo) runs on our own servers in Colombo.

Meta’s supported languages for Llama 3.3 include Hindi and Thai but not Sinhala or Tamil. Try it on your own documents before relying on it for either.

Call it

Use the OpenAI SDK you already have — point it at Roar AI and name the model. Read the docs.

from openai import OpenAI

client = OpenAI(base_url="https://api.roar-ai.com/v1", api_key="roar_live_…")

reply = client.chat.completions.create(
    model="llama-3.3-70b",
    messages=[{"role": "user", "content": "Summarise this contract in plain English."}],
)

Questions

How much does Llama 3.3 70B cost in Sri Lanka?

Input is $0.3809 per 1M tokens; output is $2.9289 per 1M tokens. For scale, 1,000 requests of about 3,000 tokens in and 1,000 out cost $4.07. The rate is published and does not change from one request to the next.

Can I pay for Llama 3.3 70B in Sri Lankan rupees?

Yes. Usage is billed to one monthly invoice in rupees, together with every other model your team uses, and you can also pay by card. There is no separate Meta account to open and no foreign card needed.

Does Llama 3.3 70B keep data in Sri Lanka?

No — prompts sent to Llama 3.3 70B are processed outside Sri Lanka. If your data has to stay in the country, use Qwen3.8-27B (Colombo), which runs on our servers in Colombo.

Is Llama 3.3 70B open source?

Its weights are open under the Llama 3.3 Community License licence, so you could run it yourself. Through Roar AI you call it like any other model, with no servers to manage.