Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
CoreWeave logo

Model details

Llama 3.3 70B

Llama 3.3 70B Instruct is Meta's December 2024 update to the Llama 3.1 70B line, designed as a text-only, auto-regressive transformer that has been instruction-tuned for conversational and assistant-style use. The 70-billion-parameter model uses an optimized transformer architecture with Grouped-Query Attention to keep inference efficient, and it was built as a multilingual generative system supporting English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai. Meta positioned the release as a way to deliver performance comparable to the much larger Llama 3.1 405B at a more accessible 70B scale, focusing on dialogue quality, tool calling, math, coding, and instruction following rather than pushing toward a new flagship size. The broad availability across GitHub Models, watsonx.ai, Amazon Bedrock, NVIDIA NIM, OpenRouter, and Fireworks AI signals an intent for the model to be widely deployed in both research and commercial product settings, including enterprise assistants and developer workflows that need a balance of capability and cost.

Under the hood, the model is pretrained on a new mix of publicly available online data totaling more than 15 trillion tokens, and the tuned release combines supervised fine-tuning with reinforcement learning from human feedback to align outputs with human preferences for helpfulness and safety. Source descriptions call out industry-leading results on common reasoning, math, and instruction-following benchmarks relative to other open and closed chat models of similar scale, alongside improved tool calling and multilingual support over its Llama 3.1 70B predecessor. Because the weights are openly released under the Llama 3.3 Community License, the model is well suited to self-hosted fine-tuning, distillation into smaller variants, and integration into retrieval, structured-output, or agent pipelines where teams want full control over data and deployment. Its static, offline-trained nature and broad serving footprint make it a practical choice for organizations looking to standardize on an open-weights mid-to-large model that can handle long-context reasoning tasks while remaining economical to run compared with frontier-scale alternatives.

CoreWeavemeta-llama/Llama-3.3-70B-Instructllama

Quick Info

Powered by
Provider
CoreWeave
Model key
meta-llama/Llama-3.3-70B-Instruct
Release date
Dec 1, 2024
Last updated
Dec 1, 2024
Knowledge cutoff
2023-12
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.71
Output token cost
$0.71

Limits

Output tokens
128,000 tokens
Context window
128,000 tokens

Transparent token rates

Compare Llama 3.3 70B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Llama 3.3 70B

Weights & Biases

CoverageRelease Notes

Meta has released a new model, Llama 3.3 70B Instruct, now available in GitHub Models. It provides similar performance to Llama 3.1 405B, but at a...

CoreWeave

CoverageBenchmark

InferenceX provides a detailed architectural and evaluation breakdown of Meta's Llama 3.3 70B Instruct, released December 6, 2024. The page documents it as a dense 70B-parameter auto-regressive transformer with 80 layers, Grouped Query Attention, RMSNorm, an output head over a 128,256-token vocabulary, and a 128K conte The same InferenceX page positions Llama 3.3 70B Instruct as Meta's instruction-tuned multilingual text-only model trained on more than 15 trillion tokens with a December 2023 knowledge cutoff, supporting English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai. Meta's own instruction-tuned comparison tab

Weights & Biases

CoverageBenchmark

The PickModel aggregator describes Llama 3.3 70B as Meta's high-performance open-source model, positioning it as near-frontier quality that is free to download, self-host, and fine-tune. It reports a 131,072-token context window, a maximum output of 8,192 tokens, and a knowledge cutoff of December 2023, with a release Benchmark figures on the PickModel page — MMLU 86%, GPQA Diamond 50.5%, MATH 73.5%, GSM8K 93.7%, and HumanEval 78.9%, with AIME 2025, LiveCodeBench, SWE-bench Verified, and HellaSwag marked as not yet published — offer an at-a-glance performance view sourced from provider technical reports and independent evaluations.

Weights & Biases

CoverageBenchmark

The n8n AI Benchmark page explicitly names "Llama 3.3 70B Instruct" and describes it as a pretrained and instruction-tuned generative model in 70B (text in/text out) optimized for multilingual dialogue. The page documents a context window of 131,072 tokens, support for English, German, French, Italian, Portuguese, Hind Per-token pricing listed ($0.00000013 prompt / $0.00000038 completion per 1K tokens) reflects a specific third-party serving gateway rather than Meta's model pricing, so those figures should not be attributed to Meta. The page provides directly relevant model-spec evidence for the exact Instruct variant and is useful a

Weights & Biases

Coverage

Llama 3.3 70B Instruct is the December update of Llama 3.1 70B. The model improves upon Llama 3.1 70B (released July 2024) with advances in tool calling, multilingual text support, math and coding. The model achieves industry leading results in reasoning, math and instruction following and provides similar performance

Videos about Llama 3.3 70B

More models around Llama 3.3 70B