DevPass (LLM Gateway)
Meta released Muse Glimmer, a 30-billion-parameter multimodal model distilled from Muse Spark, tuned for always-on local agent workflows and shipped under Apache 2.0. To make the 30B parameter model practical for local execution, Meta compresses it to roughly 4-bit and adds block-level speculative decoding so it can an According to the article, the model is aimed at solo developers and startups running it on a 24 GB GPU or an M4/M5 Max Mac, mid-market teams needing on-prem inference without per-token billing, and regulated enterprises that require an air-gappable agent. Meta recommends adding system-level guardrails rather than shipp