Llama 3.1 8B
Meta’s small Llama: quick and very cheap for simple, high-volume text work.
About Llama 3.1 8B
Llama 3.1 8B is the smallest model in Meta’s Llama 3.1 family and one of the most widely used open models anywhere. It has no reasoning mode and makes no pretence of one: it is fast, very inexpensive, follows instructions well for its size (80.4% on IFEval) and handles a 128K-token context.
Use it where the task is simple and the volume is high — routing messages, tagging, short summaries, pulling out a few fields. For anything that needs real reasoning, a 2026 small model will do much better.
Where it’s strong
Routing, tagging and extraction at very low cost.
Short summaries and simple chat.
Steps in a pipeline where speed matters more than depth.
Where to be careful
Its knowledge stops at December 2023.
Weak at maths and multi-step reasoning next to newer small models.
Meta’s licence carries use restrictions; it is not open source in the OSI sense.
Price in rupees
For scale: 1,000 requests of about 3,000 tokens in and 1,000 out cost $0.30 — a long prompt and a page of answer each. Every rate is published and stays the same from one request to the next. See every model.
Specs
llama-3.1-8bUsing it from Sri Lanka
Llama 3.1 8B is billed in rupees, on the same monthly invoice as every other model your team uses — no foreign card, no separate Meta account, and one key for all of it.
Prompts sent to Llama 3.1 8B are processed outside Sri Lanka. If a workload has to stay in the country, Qwen3.8-27B (Colombo) runs on our own servers in Colombo.
Meta officially supports eight languages for Llama 3.1, including Hindi, but not Sinhala or Tamil.
Call it
Use the OpenAI SDK you already have — point it at Roar AI and name the model. Read the docs.
from openai import OpenAI
client = OpenAI(base_url="https://api.roar-ai.com/v1", api_key="roar_live_…")
reply = client.chat.completions.create(
model="llama-3.1-8b",
messages=[{"role": "user", "content": "Summarise this contract in plain English."}],
)Compare Llama 3.1 8B
Questions
How much does Llama 3.1 8B cost in Sri Lanka?
Input is $0.065 per 1M tokens; output is $0.104 per 1M tokens. For scale, 1,000 requests of about 3,000 tokens in and 1,000 out cost $0.30. The rate is published and does not change from one request to the next.
Can I pay for Llama 3.1 8B in Sri Lankan rupees?
Yes. Usage is billed to one monthly invoice in rupees, together with every other model your team uses, and you can also pay by card. There is no separate Meta account to open and no foreign card needed.
Does Llama 3.1 8B keep data in Sri Lanka?
No — prompts sent to Llama 3.1 8B are processed outside Sri Lanka. If your data has to stay in the country, use Qwen3.8-27B (Colombo), which runs on our servers in Colombo.
Is Llama 3.1 8B open source?
Its weights are open under the Llama 3.1 Community License licence, so you could run it yourself. Through Roar AI you call it like any other model, with no servers to manage.
