CloudFerro Sherlock
Meta has released a new model, Llama 3.3 70B Instruct, now available in GitHub Models. It provides similar performance to Llama 3.1 405B, but at a...
Model details
Llama 3.3 70B Instruct is Meta's instruction-tuned generative model in the 70B parameter class, designed for text-in and text-out workloads such as multilingual dialogue, reasoning, math, and structured instruction following. Meta positioned the release as offering performance comparable to the much larger Llama 3.1 405B while delivering better results than both Llama 3.1 70B and Llama 3.2 90B on common text benchmarks, reflecting continued gains in instruction tuning efficiency for the Llama family. As an openly licensed model, it gives organizations a middle-ground option between smaller open chat models and frontier-scale closed systems, particularly for production assistants and analytical workflows where reasoning quality matters more than extreme context length.
In practice, Llama 3.3 70B Instruct fits well into retrieval-augmented generation, agent pipelines, and enterprise chat applications that need reliable instruction adherence without paying for the largest models. The model is deployable across multiple ecosystems, appearing as a managed offering on Oracle Cloud Infrastructure for on-demand inferencing, dedicated clusters, and fine-tuning, and as a "Direct from Azure" option on Microsoft Foundry with unified billing and pay-as-you-go scaling. Third-party routes such as OpenRouter additionally expose a free tier with a roughly 65k token context window, which is well suited to multi-document summarization and tool-augmented reasoning where the full chat history and retrieved passages need to fit within a single prompt.
CloudFerro Sherlock
Meta has released a new model, Llama 3.3 70B Instruct, now available in GitHub Models. It provides similar performance to Llama 3.1 405B, but at a...
CloudFerro Sherlock
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). $0 per million input tokens, $0 per million output tokens. 65,536 token context window. Higher uptime with 14 providers. Includes independent benchmarks from Artificial Analysis.
This exact model name is also listed by 25 other providers.