Helicone
I am attempting to serve Mistral-Large-3-Instruct-2512-NVFP4 on a node with 8x NVIDIA B300 (Blackwell) GPUs using vLLM. The engine initializes the workers but hangs indefinitely at the Waiting for ...
Model details
Mistral Large was designed as Mistral AI's flagship text generation model, engineered to deliver top-tier reasoning across demanding workloads. The model emphasizes complex multilingual reasoning, spanning text understanding, transformation, and code generation tasks, with native fluency in English, French, Spanish, German, and Italian that extends beyond simple translation to a nuanced grasp of grammar and cultural context. Its architecture supports a 32K token context window, enabling precise information retrieval from lengthy documents, and includes native function calling paired with constrained output modes that allow developers to build moderation policies and scale application development directly on the platform.
Positioned early on as the world's second-ranked generally available API model next to GPT-4, Mistral Large achieved strong results on commonly used benchmarks, signaling a deliberate push into the top tier of language model performance. The model was released through la Plateforme and Microsoft Azure, reflecting a strategy of broad cloud distribution from launch. Its instruction-following precision and tool-calling capabilities make it well-suited for developers building automated workflows, tech stack modernization projects, and AI-native applications that require reliable function execution alongside high-quality text generation.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Helicone
I am attempting to serve Mistral-Large-3-Instruct-2512-NVFP4 on a node with 8x NVIDIA B300 (Blackwell) GPUs using vLLM. The engine initializes the workers but hangs indefinitely at the Waiting for ...
This exact model name is also listed by 8 other providers.