Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Pioneer logo

Model details

GLiNER-2.5-Decide

GLiNER-2.5-Decide is a compact 340M-parameter open-weight encoder built by Fastino Labs for schema-defined decision making. Rather than generating free-form text, the model takes a passage and a set of typed questions and returns valid answers together with probabilities, confidence scores, and constraint-feasibility metadata, so callers can see not only what the model decided but how strongly and whether the answer respects the supplied schema. Alongside these constrained decisions it can extract spans and relations and enforce rules across related outputs, making it well suited to classification, routing, triage, and other structured content-understanding tasks where consistency across fields matters more than open-ended generation.

Because the model is released under the Apache 2.0 license and runs locally on CPUs, it can be deployed in air-gapped or otherwise restricted environments and supports both full and LoRA-based fine-tuning for downstream adaptation. Fastino reports end-to-end p50 latency of 38.3 ms on an NVIDIA V100 and 167.3 ms on a 48-vCPU Intel Xeon Platinum 8581C for short-document inputs, and on its unseen 17-dataset benchmark spanning classification, routing, triage, and content understanding the model reaches the highest reported average score, outperforming both a much larger decoder-based decision baseline and a comparable encoder baseline. The combination of small footprint, local deployment, structured outputs, and fine-tuning hooks positions GLiNER-2.5-Decide as a practical choice when teams need predictable, schema-aware predictions with low latency.

Pioneerfastino/GLiNER-2.5-Decide

Quick Info

Powered by
Provider
Pioneer
Model key
fastino/GLiNER-2.5-Decide
Release date
Sep 24, 2026
Last updated
Sep 24, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.15

Limits

Output tokens
8,192 tokens
Context window
8,192 tokens

Latest news about GLiNER-2.5-Decide

Pioneer

Coverage

Fastino Labs released GLiNER2.5-Decide on September 24, 2026, as an open-weight encoder-based decision model designed to run on CPU hardware rather than GPUs. According to Fastino's release materials and Hugging Face model card, the model carries 340 million parameters and ships under the Apache 2.0 license at the path fastino/GLiNER2.5-Decide. Fastino described it on X as built for fast, deterministic classification. Rather than generating open-ended text, GLiNER2.5-Decide accepts a passage plus a set of user-defined questions or rules and returns a structured decision with probability distributions and confidence scores. Fastino's own benchmark reports a p50 latency of 167.3 milliseconds at batch size 1 and 64 tokens on a 48-vCPU Intel Xeon Platinum 8581C, requiring no GPU. Use cases cited include compliance screening and structured classification workflows where generative LLM nondeterminism is a liability.

Pioneer

Coverage

Fastino Labs announced GLiNER-2.5-Decide on September 24, 2026, a 340-million-parameter open-weight encoder model designed to replace LLM calls with deterministic, auditable decisions. It targets routing, triage, classification, sentiment, and LLM-as-judge tasks where generative models are overkill. Weights are released under Apache 2.0 on Hugging Face as fastino/GLiNER2.5-Decide. Instead of generating labels, users define a schema of typed questions with permitted answers and cardinality rules; a constrained decoder then resolves contradictions and emits probabilities and confidence metadata. The model is CPU-friendly, reporting 167.3 ms p50 latency at 64 tokens on a 48-vCPU Xeon, and achieves a 60.1% average across 17 datasets, leading on 9.

Videos about GLiNER-2.5-Decide