Currently listed through these providers:
Model details
Llama 3.3 70B Instruct fp8 Fast
Llama 3.3 70B Instruct fp8 Fast belongs to Meta's Llama 3.3 family of decoder-only transformer models, carrying forward the architecture and training advances established in earlier Llama releases. The fp8 quantization is a deliberate engineering choice: reducing numerical precision to 8-bit floating point shrinks the model's memory footprint and enables faster inference without sacrificing the core capabilities of the 70-billion parameter base. This makes the model substantially more accessible for teams that want the power of a large language model without the overhead of full-precision deployment.
The model ships with built-in tool calling, making it well-suited for agentic workflows where a language model needs to invoke external functions or APIs as part of a reasoning loop. As an open-weights release, it invites community fine-tuning and customization, enabling developers to adapt the base model to specialized tasks in coding, analysis, and structured generation. The combination of a large parameter count, fp8 speed optimization, and function-calling support positions this model for developers building interactive AI applications that require both depth and responsiveness.
Quick Info
Powered by- Provider
- Cloudflare Workers AI
- Model key
- @cf/meta/llama-3.3-70b-instruct-fp8-fast
- Release date
- Dec 6, 2024
- Last updated
- Dec 6, 2024
- Knowledge cutoff
- 2023-12
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.293
- Output token cost
- $2.253
Limits
- Output tokens
- 24,000 tokens
- Context window
- 24,000 tokens