Currently listed through these providers:
Model details
Granite 4.0 H Micro
Granite-4.0-H-Micro is a 3 billion parameter instruct model built on a Hybrid H Mamba architecture, positioning it within IBM's broader Granite 4.0 family alongside smaller Nano and larger Small variants. The architecture choice reflects a deliberate engineering balance: Mamba-based designs can offer faster inference and lower memory consumption compared to traditional transformer-only designs, which matters significantly for a model deployed at the edge or in high-throughput enterprise environments. The model is explicitly designed as a long-context capable instruction follower with strong tool-calling capabilities, making it suitable for building AI assistants that need to maintain conversation history and take structured actions across business workflows.
The training lineage of this model draws from a diverse pipeline combining supervised finetuning, reinforcement learning-based alignment, and model merging techniques, built upon a base model trained on approximately 15 trillion tokens. This multi-stage approach reflects a modern post-training philosophy where raw capability is developed first, then refined through instruction datasets with permissive licenses alongside internally generated synthetic data. An October 2025 update further introduced a default system prompt to steer the model toward more professional and consistent responses, a targeted refinement for enterprise deployment scenarios. The model supports twelve languages out of the box and targets business-domain applications like summarization, text classification, and extraction, with Apache 2.0 licensing making it accessible for organizations wanting to self-host or fine-tune without licensing overhead.
Quick Info
Powered by- Provider
- Cloudflare Workers AI
- Model key
- @cf/ibm-granite/granite-4.0-h-micro
- Release date
- Oct 7, 2025
- Last updated
- Oct 7, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.017
- Output token cost
- $0.112
Limits
- Output tokens
- 131,000 tokens
- Context window
- 131,000 tokens