Muse Spark 1.2 is the third Muse release in roughly four months, positioned as a step-change in agentic capability rather than a wholesale redesign. On the Artificial Analysis Intelligence Index, the xhigh variant scores 54, up from 51 on the prior release and 43 at the April debut, placing Meta in a tie with SpaceXAI for third among US labs and within a point of GPT-5.5 xhigh. The lift is most visible on agentic knowledge work, the area Artificial Analysis flagged as the clearest weakness at launch of the previous version. It also makes measured gains on agentic coding and terminal use, suggesting the same underlying family is being tuned toward longer, tool-using workflows rather than raw chat quality.
In practical terms, Muse Spark 1.2 is best understood as a stronger agent for realistic multi-step tasks such as preparing presentations and analyses, browsing, and shell-based work, where it now ranks among the leaders Artificial Analysis has benchmarked. Its relative weakness, framed by the headline "Improved Agentic Performance at Higher Cost per Task," is economics: getting those agentic gains costs more per task than earlier Muse Spark versions, which matters most for budget-sensitive, high-volume deployments. Teams that already run agent loops and want stronger real-world task completion will see the most value, while cost-optimized or simple single-turn use cases are likely better served by lighter or older options.