Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

Granite 4.2 8B

IBM Granite 4.2 8B is positioned within the Granite family as a mid-sized language model from IBM. On the composite LLM Stats leaderboard, the model ranks 207 overall with a score of 19.7 and a reported blended price of $0.069 per million tokens, placing it in the same broad cost band as Gemma 4 E4B and GPT OSS 120B while delivering substantially lower composite performance than either. The name suggests an 8-billion-parameter scale class, though no technical report or model card excerpt was available to confirm parameter count or architecture details.

Capability-tier placement indicates a clear emphasis on mathematical reasoning relative to other tasks. IBM Granite 4.2 8B earns a tier B (Average) rating in the top-half tier for Math, where it ranks 151 of 329 tracked models, while landing in tier C (Below top half) across Long Context, Legal, Finance, Healthcare, Tool Calling, Reasoning, and Coding categories, with the weakest standings in Reasoning (216 of 370) and Coding (233 of 274). This profile points to fit in math-assisted analytics and structured problem decomposition, with less reliable performance for open-ended reasoning chains, long-context document work, or domain-specific professional tasks in legal, financial, or healthcare settings.

DevPass (LLM Gateway)granite-4.2-8bgranite

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
granite-4.2-8b
Release date
Sep 1, 2026
Last updated
Sep 1, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.25

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Transparent token rates

Compare Granite 4.2 8B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Granite 4.2 8B

DevPass (LLM Gateway)

CoverageBenchmark

Granite 4.2 introduces native reasoning to IBM's open model line, enabling step-by-step planning and self-correction within think tags before answering. The family ships in three sizes, 3B, 8B, and 30B, all released under Apache 2.0, with reasoning switchable across full thinking, non-thinking direct answer, and low-effort brief reasoning modes. This lets a single model skip chain-of-thought on queries where latency matters, paying for reasoning only when needed. Compared to Granite 4.0's fast instruct-only design, the 4.2 family adds reasoning depth aimed at math, multi-step logic, and tool-selection decisions. The model's benchmark gains are most visible in competition math and code tasks, where 8B-level dense reasoning competes effectively. Developers can toggle reasoning effort per request, making Granite 4.2 adaptable for both high-volume low-latency serving and deeper analytical workflows.

DevPass (LLM Gateway)

CoverageBenchmark

The eesel AI explainer positions Granite 4.2 as IBM's reasoning-focused open-model release consisting of three checkpoints: granite-4.2-3b, granite-4.2-8b, and granite-4.2-30b, each post-trained from the matching Granite 4.1 base. Granite 4.2 8B is described as a dense 8B-parameter decoder-only transformer with 40 layers and a 128K context window, post-trained from granite-4.1-8b-base. The article frames IBM's pitch as competitive reasoning capability at lower cost than larger thinking models. Release logistics include Apache 2.0 weights, cryptographic signing, ISO certification, and a top rating on Stanford's Foundation Model Transparency Index, targeting regulated on-premises deployment. The piece notes native toggleable reasoning and reasoning-augmented tool calling as the key new capabilities versus Granite 4.0, alongside a shift from hybrid to purely dense architectures. Use-case guidance positions 8B for general-purpose enterprise work, with 3B targeting laptops/edge and 30B for heavier workloads.

DevPass (LLM Gateway)

CoverageBenchmark

IBM Granite 4.2 8B was released on August 31, 2026, as an 8-billion-parameter dense reasoning model with a 131K-token context window. It is positioned for mathematics, code generation, multilingual dialogue, and agentic workflows requiring multi-step reasoning. The model achieves 99.0% accuracy on Email Classification, 91.0% on Coding, and a 99% reliability success rate, with perfect accuracy in acknowledging uncertainty under Hallucinations evaluation. Granite 4.2 8B supports structured outputs, tools, response format, logprobs, reasoning inclusion, and seed parameters via its API, with reported pricing of $0.10 per 1M prompt tokens and $0.15 per 1M completion tokens. Its Reasoning accuracy of 26.0% and Instruction Following of 68.7% show room for improvement relative to its dense reasoning model positioning. The model is available on Hugging Face and is listed across multiple inference providers including CoreWeave and DeepInfra.

DevPass (LLM Gateway)

CoverageBenchmark

IBM's Granite 4.2 8B is the mid-size variant in IBM's reasoning-focused Granite 4.2 family, a dense Apache 2.0 model that completes IBM's full agentic-RL training block across SWE-agent, Terminal-agent, and Search-agent stages. It reports 47.7% on SWE-bench Verified, 86.7% on AIME25, and 50.3% on BFCL-v4, with a native 128K-token context window. The model shares a five-phase pre-training pipeline (15T tokens) with its 3B and 30B siblings, plus 1 trillion tokens of synthetic code from IBM's CodeAlchemy pipeline. Granite 4.2 8B supports native tool calling in OpenAI function-calling format, integrating directly with OpenHands, OpenCode, and Pi harnesses, and includes a speculative-decoding layer for faster inference. It trails the 30B sibling on every reported metric, with the largest gaps on SWE-bench Pro (19.1% vs 33.3%) and Terminal-Bench 2.1 at 20.6%. The model is text-only and not natively multimodal, targeting enterprise agentic coding, terminal, and search workflows on mid-size hardware where teams want agentic-RL-trained tool use without the 30B's compute footprint.

DevPass (LLM Gateway)

CoverageBenchmark

The LLM Stats page names IBM as the creator and reports a blended price of $0.069 per 1M tokens against an LLM Stats Score of 19.7, ranking the 8B variant at 207 overall. Per-category standing places the 8B in the B Average tier for Math (rank 151 of 329) and top-half Long Context (rank 91 of 119), while Coding (233 of 274) and Reasoning (216 of 370) sit below top half. The page explicitly cautions that benchmark scores are self-reported by the model provider. Individual benchmark rows include AIME 2025 (rank 62, score 0.87/1), RULER 64k (rank 3, score 0.81/1), IFBench (rank 10, score 0.79/1), and HMMT25 (rank 14, score 0.78/1), all sourced to huggingface.co. Cost-efficiency placement sits between GPT OSS 120B ($0.043, score 28.6) and DeepSeek-V4-Flash-0731 ($0.066, score 43.8). The page was generated on Tue Sep 29 2026.

DevPass (LLM Gateway)

CoverageBenchmark

BenchLM's page for Granite 4.2 8B lists a capability score of 30.2/100 against a field median of 50.2, ranking 159 of 194 tracked models, and frames the model as self-hosted with infrastructure cost varying rather than a first-party API price. Speed is reported at 102 tok/s against a 91 tok/s field median, with a 20.48-second first-token latency and a 128K token context window. The page does not explicitly attribute creation to IBM or any provider. Across 14 published benchmark rows, Granite 4.2 8B ranks 97 of 105 in Agentic (8th percentile), 124 of 135 in Coding (8th percentile), and 141 of 158 in Knowledge (11th percentile), while Instruction Following is comparatively stronger at rank 47 of 124 (63rd percentile, score 82.8). Reasoning, Multimodal, and Multilingual categories have no measured benchmarks published. The Decision snapshot describes it as a well-rounded choice across tasks with validation recommended given the limited evidence base.

DevPass (LLM Gateway)

CoverageBenchmark

Granite 4.2 8B is an open-weights language model attributed to IBM, released August 2026 under Apache 2.0 with 8B total parameters and text input/output support. The Artificial Analysis page reports an Intelligence Index score of 11, placing it above the open-weights median of 8, alongside a 131k token context window and reasoning capability. Pricing on the page is listed at $0.06 per 1M input tokens and $0.25 per 1M output tokens with a 75% cache discount, though these reflect Artificial Analysis's tracked host rates rather than creator-set pricing. Speed is logged at 71.8 output tokens per second, which is slower than the 90 tok/s median for comparable open-weights models in its size class. The model generated 130M tokens during Intelligence Index evaluation, described as somewhat verbose versus the 82M median. The page also notes a non-reasoning sibling variant may exist, and the model is positioned for small-class (4B–40B parameters) open-weights comparison, supporting deployment context for enterprise and edge use.

Videos about Granite 4.2 8B

More models around Granite 4.2 8B