Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Novita AI logo

Model details

Macaron V1 Venti

Macaron V1 Venti is a 748B-parameter flagship from the Macaron-V1 family, designed around personal intelligence, tool use, coding workflows, and code-native generative UI. Its architecture is a Mixture-of-LoRA (MoL) setup built on top of GLM-5.2, combining a 744B-parameter base model with four 1B-parameter LoRA specialists that cover chat, personal-agent tasks, coding, and GenUI. An L0 router picks the most suitable specialist for each incoming request, which lets a single deployment switch behavior between conversational assistance and code-oriented generation without retraining the base weights. This specialist design points to a practical focus on adaptability across everyday assistance and developer-style workloads in one model.

On the serving side, the model is exposed through Novita AI's serverless API with OpenAI-compatible endpoints, making it straightforward to drop into existing toolchains. It supports a million-token context window with outputs up to 128K tokens, well suited to long-running coding sessions, multi-file reasoning, and rich generative UI assembly. Function calling and reasoning are positioned as core capabilities, aligning with the model's intent of acting as both a personal agent and a coding assistant. For teams already invested in OpenAI-style SDKs, the combination of a large MoL backbone and a long context window makes Venti a strong fit for agentic pipelines and complex code generation tasks.

Novita AImindai/macaron-v1-venti

Quick Info

Powered by
Provider
Novita AI
Model key
mindai/macaron-v1-venti
Release date
Jul 21, 2026
Last updated
Jul 21, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.825
Output token cost
$2.475

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Latest news about Macaron V1 Venti

Novita AI

Coverage

Novita AI's blog announces Macaron V1 Venti availability through its hosted API, describing a 748B-parameter Mixture-of-LoRA model built on GLM-5.2 aimed at software development, autonomous agents, and generated user interfaces. The post frames Venti as worth evaluating when applications benefit from a very large model with a long context window, while noting that smaller models can be faster and more predictable for tightly scoped tasks. The listing specifies a 1,048,576-token context window, a model ID of mindai/macaron-v1-venti, and the 744B base plus four 1B LoRA specialists with an L0 router selecting the best specialist per request. Developers are pointed to the Novita quick-start documentation and the Macaron V1 Tall sibling for family comparison, emphasizing context-heavy use cases like large repositories, long agent traces, and multi-document sessions.

Novita AI

Coverage

Mind Lab released Macaron-V1, introducing the flagship Macaron-V1-Venti as a 748B-parameter model combining a 744B GLM-5.2 base with four 1B-parameter LoRA specialists. The model retains the Mixture-of-LoRA architecture from the Preview release while advancing the base from GLM-5.1 to GLM-5.2, broadening benchmarking, and adding the LongStraw long-context infrastructure that enables up to 2M-token training contexts. Venti is positioned for personal intelligence, tool use, coding workflows, and code-native Generative UI, with the four LoRA specialists covering chat, personal-agent tasks, coding, and GenUI. V1 expands evaluation across personal-intelligence benchmarks designed in-house as well as general agent and coding suites, reflecting its dual bet on recursive self-improvement and multi-agent collaboration over monolithic scaling.

Novita AI

CoverageBenchmark

Benchmark Atlas aggregates Macaron-V1-Venti's self-reported scores across twelve benchmarks, noting a parameter-count discrepancy: Mind Lab markets 748B (744B base plus four 1B specialists) while the published safetensors contain 784,082,110,464 parameters. The model is sourced from Mind Lab with a July 21, 2026 release date and described as an open-weight personal-agent model built on a frozen GLM-5.2 base. Reported headline scores include SWE-bench Verified at 85.6, Terminal-Bench 2.1 at 87.6, UI4A-Bench at 87.8, PinchBench Best at 94.0, DeepSWE v1 at 58.4, and Macaron ChatBench at 58.3, with all twelve entries flagged as self-reported from the Macaron-V1 launch appendix. Rankings are provided relative to comparable models, giving developers an at-a-glance view of Venti's standing across agent and coding evaluations.

Novita AI

Coverage

The updated arXiv report (v2, 24 Aug 2026) restates Macaron-V1's dual design goals of adaptation through recursive improvement and collaboration through the Mixture-of-LoRA architecture with one LoRA selected per turn. Venti is documented as a 748B model combining a 744B GLM-5.2 base with four LoRAs covering chat, agent, coding, and GenUI workloads. V2 reinforces the system's co-design across architecture, algorithms, and infrastructure, including UI4A, MindForge, MinT, and LongStraw, plus stability techniques for sparse MoE and DSA base models. The authors frame ongoing gains from continual learning and collective intelligence as open research questions, situating Macaron-V1 within a broader shift toward post-training advances over raw pre-training scale.

Novita AI

Coverage

The ModelScope model card describes Macaron-V1-Venti as a 748B-parameter flagship using a Mixture-of-LoRA design on a 744B GLM-5.2 base with four 1B-parameter LoRA specialists, each handling chat, personal-agent, coding, or GenUI tasks. An L0 router selects the best specialist per request, and the card lists the published safetensors checkpoint at 753.33B parameters under an MIT license. The card highlights Venti as the first model post-trained on GLM-5.2 and notes the co-designed training stack using MinT and MindForge for production-aligned routing, tool use, UI4A Generative UI, and agent workflows. It also documents LongStraw-based long-context post-training supporting Venti's 1M-token context, with reported leading scores on ChatBench, LivingBench, PinchBench, TerminalBench 2.1, and UI4ABench.

Videos about Macaron V1 Venti