Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vultr logo

Model details

Muse Glimmer 30B

Muse Glimmer 30B is a dense causal transformer from Meta Superintelligence Labs, paired with a perception encoder for handling visual and video inputs alongside text. The architecture is not a mixture-of-experts design, keeping it suitable for single-machine deployment rather than distributed serving. Meta supplies calibrated GGUF quantizations along with separate vision and speculative-decoding components, letting the model run on consumer hardware with roughly 24GB to 32GB of memory. A 131K context window configuration allows the model to manage long agent sessions, code repositories, and extended document reasoning without truncation.

The model is positioned for local agent use, tool-augmented reasoning, and structured output generation rather than as a general-purpose flagship. Meta's launch evaluation covers an unusually broad agent benchmark suite, and third-party analysis notes that competing dense models like Qwen3.6-27B still edge ahead on some practical agent and multimodal tests in the published comparison. Glimmer's real differentiator is the open-weights release that lets teams self-host, fine-tune, and run speculative decoding locally, which makes it attractive for organizations that need on-premise inference without relying on a hosted endpoint.

Vultrmuse-glimmer-30bmuse

Quick Info

Powered by
Provider
Vultr
Model key
muse-glimmer-30b
Release date
Aug 10, 2026
Last updated
Aug 10, 2026
Knowledge cutoff
2026-01-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.25
Output token cost
$1.00

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Transparent token rates

Compare Muse Glimmer 30B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Muse Glimmer 30B

Vultr

CoverageBenchmark

Meta released Muse Glimmer 30B on 10 August 2026 as a 29.6-billion-parameter dense causal transformer plus a roughly 1.8B perception encoder, targeting 24 GB and 32 GB local hardware. The model accepts interleaved text and images, produces text, runs at a 131,072-token context, and ships under Apache 2.0 weights with Meta's Usage Policy on Hugging Face as meta-models/Muse-Glimmer-30B. Configuration evidence shows a 52-layer decoder with 6,656 hidden size and a hybrid attention pattern of three 2,048-token sliding-window layers followed by one global layer, repeating. Released artifacts include BF16 weights, two 4-bit quantizations, a DFlash drafter, and a frozen ViT-G/14 vision encoder, with reasoning levels controllable as low, medium, high, or xhigh.

Videos about Muse Glimmer 30B

More models around Muse Glimmer 30B