Nebius Token Factory
Explore the results of our LLM quantization benchmark where we compared 4 precision formats of Qwen3-32B on a single H100 GPU.
Model details
Qwen3-32B is a mid-sized open-weight language model in the Qwen3 family, positioned as a practical balance between capability and computational efficiency for production text tasks. Independent benchmarking studies have examined its behavior through quantization experiments, comparing BF16, FP8, and INT4 precision formats on high-end accelerators to assess deployment trade-offs. Additional performance profiling has evaluated the model using inference frameworks such as Llama.cpp across different hardware platforms, helping practitioners understand throughput characteristics on both data-center and workstation-class GPUs.
The model's open-weight availability makes it suitable for teams that need to self-host for data sovereignty, cost control, or fine-tuning experiments, while its served endpoint provides a managed option for those preferring API access. Its mid-range parameter count offers a workable compromise for workloads that exceed what smaller models can handle but where larger frontier models would be cost-prohibitive, such as structured document analysis, multi-turn agent workflows, and instruction-following applications. The combination of open release, competitive pricing, and independent third-party validation through community benchmarks positions it as a versatile option in the current open-model landscape.
Nebius Token Factory
Explore the results of our LLM quantization benchmark where we compared 4 precision formats of Qwen3-32B on a single H100 GPU.
Nebius Token Factory
First up was Qwen3 32B with Llama.cpp on these two impressive NVIDIA platforms. For TG128, the GH200 platform was around 10.6x the performance of the GB10.
Nebius Token Factory
Qwen3 32B pricing: $0.08/M input, $0.24/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.