Sulat.com
AI models
Cloudflare Workers AI logo

Model details

Llama 3.3 70B Instruct fp8 Fast

Llama 3.3 70B Instruct fp8 Fast belongs to Meta's Llama 3.3 family of decoder-only transformer models, carrying forward the architecture and training advances established in earlier Llama releases. The fp8 quantization is a deliberate engineering choice: reducing numerical precision to 8-bit floating point shrinks the model's memory footprint and enables faster inference without sacrificing the core capabilities of the 70-billion parameter base. This makes the model substantially more accessible for teams that want the power of a large language model without the overhead of full-precision deployment.

The model ships with built-in tool calling, making it well-suited for agentic workflows where a language model needs to invoke external functions or APIs as part of a reasoning loop. As an open-weights release, it invites community fine-tuning and customization, enabling developers to adapt the base model to specialized tasks in coding, analysis, and structured generation. The combination of a large parameter count, fp8 speed optimization, and function-calling support positions this model for developers building interactive AI applications that require both depth and responsiveness.

Cloudflare Workers AI@cf/meta/llama-3.3-70b-instruct-fp8-fastllama

Quick Info

Powered by
Provider
Cloudflare Workers AI
Model key
@cf/meta/llama-3.3-70b-instruct-fp8-fast
Release date
Dec 6, 2024
Last updated
Dec 6, 2024
Knowledge cutoff
2023-12
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.293
Output token cost
$2.253

Limits

Output tokens
24,000 tokens
Context window
24,000 tokens

Latest news about Llama 3.3 70B Instruct fp8 Fast

Videos about Llama 3.3 70B Instruct fp8 Fast

Recent tweets and retweets from Cloudflare Workers AI

More models around Llama 3.3 70B Instruct fp8 Fast