Sulat.com
AI models

AI news, without the noise.

Model releases, provider updates, and independent coverage. Clear summaries, original sources, fewer repeated stories.

AI-generated summaries.
Original sources always linked.

23 stories · Past 30 days · Newest first

Follow the latest AI news via RSS

One feed for the stories here. Add it to your favourite reader—no account needed.

RSS feed

DataNorth AI

Liquid AI d1: decision model with zero output tokens

Liquid AI released d1 on 29 September 2026, a decision model which classifies, routes, and scores inputs without producing any output tokens. Instead of generating text, it outputs numeric answers and bills only for input tokens, targeting use cases such as content moderation, ticket triage, safety guardrails, and agent tool-call approvals.

The model supports Noul yes/no questions, Choice selections across named options, and Score ratings on an ordered scale, all combinable in a single call. It offers a 32,000-token context window, proprietary hosted API access, no downloadable weights, and no fine-tuning. Access is available through Liquid's API and Vercel AI Gateway under the model ID liquid/d1, with OpenRouter support planned.

Read original sourceIndependent coverage

Trending Topics

GPT-6.1 Sol: Nearly as Smart as Astra, Far Cheaper | Trending Topics

OpenAI unveiled GPT-6.1 Sol at DevDay, positioning it as delivering near-Astra intelligence for coding, computer use, and professional work at roughly one-fifth of GPT-6 Astra's API cost. The model is attributed to OpenAI; the article does not address Vercel AI Gateway specifics. Family-level evidence applies to the GPT-6.1 Sol Fast alias.

According to Artificial Analysis's Intelligence Index, GPT-6.1 Sol reaches 52 points at maximum reasoning effort — four points above its GPT-6 Sol predecessor and one point below GPT-6 Astra, placing it tenth among 222 tested models. API pricing is $2 per million input tokens and $10 per million output tokens, with cached input at $0.10 per million tokens and a context window exceeding one million tokens.

Read original sourceIndependent coverage

The Rundown AI

Claude Sonnet 5.5 nears Opus ahead of OpenAI’s DevDay | The Rundown AI

The Rundown AI reports Anthropic launched Claude Sonnet 5.5 on September 28, 2026, positioning the middle tier close to Opus 5.5 on several tests at half the flagship's input and output token prices. At maximum effort, Artificial Analysis gave Sonnet 5.5 a score of 56 on its Intelligence Index, behind only Opus 5.5 at 58, and 64% on Terminal-Bench 4.0 versus roughly 60% for Opus 5.5 and GPT-6 Astra. AA's results were provisional pending reruns due to a prerelease structured-output bug.

Sonnet 5.5 is priced at $2 per million input tokens and $10 per million output tokens, with output at least 30% faster than Sonnet 5 and task costs up to 30% lower. At maximum effort, AA measured roughly $7.60 per task, about 50% more than Sonnet 5. Anthropic gives Opus the edge on complex work requiring judgment, while Sonnet 5.5 suits well-scoped everyday tasks like bug fixes and document production.

Read original sourceIndependent coverage

remio

Fireworks AI Ember-1 Cuts Kimi K3's Token Load, but the Evidence Is Still Vendor-Run

Fireworks announced Ember-1 on September 23, 2026, as a specialized model post-trained from Moonshot AI's Kimi K3 open weights, training reasoning length as a behavior rather than leaving it as a fixed cost of model quality. The objective was narrower than building a new foundation model: teach K3 to use shorter reasoning traces without abandoning useful analysis.

Fireworks reports completing more than 50 training experiments and over 200 evaluations across mathematics, coding, instruction following, conversation, search, tool use, and software engineering, using task and environment feedback for on-policy learning. However, the evidence is primarily vendor-run, and Ember-1 is API-only with no released weights or reproducible training pipeline, leaving outside researchers unable to independently inspect or reproduce the result.

Read original sourceIndependent coverage

Medium

MiniMax M3.1-Flash-Preview, tested outside MiniMax Code | by JP Caparas | Sep, 2026 | AI @ Sulat.com

MiniMax-M3.1-Flash-Preview debuted on September 27, 2026, initially inside MiniMax Code, positioned for everyday development work spanning bug fixes through full feature delivery. A first-hand test wired it into OpenCode via a custom provider entry and confirmed it answers on the public API, softening the walled-garden impression from the launch note. The author collected a day of speed and pass-rate numbers against the model directly.

Beyond benchmarks, the testing surfaced one consistent failure mode so repeatable it could be used as a timing signal. MiniMax also layered promotional sweeteners on launch: Token Plan quotas reset at launch and daily check-in credits doubled from September 28 through October 7, applicable across the catalog including H3 video generation. No published comparison to MiniMax M3 was supplied at launch.

Read original sourceIndependent coverage

Shattered

Fastino GLiNER2.5-Decide: 340M CPU AI Model [2026]

Fastino Labs released GLiNER2.5-Decide on September 24, 2026, as an open-weight encoder-based decision model designed to run on CPU hardware rather than GPUs. According to Fastino's release materials and Hugging Face model card, the model carries 340 million parameters and ships under the Apache 2.0 license at the path fastino/GLiNER2.5-Decide. Fastino described it on X as built for fast, deterministic classification.

Rather than generating open-ended text, GLiNER2.5-Decide accepts a passage plus a set of user-defined questions or rules and returns a structured decision with probability distributions and confidence scores. Fastino's own benchmark reports a p50 latency of 167.3 milliseconds at batch size 1 and 64 tokens on a 48-vCPU Intel Xeon Platinum 8581C, requiring no GPU. Use cases cited include compliance screening and structured classification workflows where generative LLM nondeterminism is a liability.

Read original sourceIndependent coverage

zavino.co

Meituan LongCat 2.5: 1.6T Parameters, Free on OpenCode | ZAVINO

Meituan's LongCat team has released LongCat-2.5-Preview, a mixture-of-experts model with 1.6 trillion total parameters and roughly 48 billion active per pass, designed for long-horizon agent workloads across browsers, terminals, GUIs and spreadsheets. The post also corrects a circulating "16 trillion parameter" rumor, noting both the HuggingNews page and a Chinese spec summary list 1.6T.

The preview keeps a one-million-token native context window and adds image understanding for multi-app workflows. OpenCode is offering the model free for two weeks under a zero-data-retention policy, though no published benchmarks, weights, or post-promotion pricing were available from the sources reviewed.

Read original sourceIndependent coverage

tpsreport.news

Perceptron Mk1.5: New Perception Model for Robotics AI | TPS

Perceptron released Mk1.5 on September 25, 2026, as a multimodal perception model built for physical agents. It accepts text, image, video, and audio inputs, and outputs natural-language text alongside optional structured spatial annotations including points, bounding boxes, polygons, and temporal clips. Pricing is set at $0.15 per million input tokens and $1.50 per million output tokens.

Mk1.5 succeeds Perceptron's prior flagship vision-language model Mk1, extending its support to four input modalities and structured annotation outputs. The model supports graded reasoning via standard reasoning controls, function tool calling, and structured outputs via JSON Schema. Audio analysis of video soundtracks is opt-in and only runs when explicitly enabled per request.

Read original sourceIndependent coverage

AI×AI

Space Bunny Alpha Stealth Model: Benchmarks, MiniMax M3 & Specs | AI×AI

On September 23, 2026, an anonymous stealth AI model called Space Bunny Alpha appeared on OpenRouter, Cline, and Command Code, offered free for a one-week testing window under a zero-data-retention policy. Within twelve hours, community researchers matched its tokenizer signatures, Chinese error traces, and prompt defaults to Shanghai-based MiniMax, suggesting the model is an unannounced public test of MiniMax M3.

Space Bunny Alpha features a 1,000,000-token context window, a 524,288-token output ceiling, and native text, image, audio, and video processing without separate perception encoders. Developers reported strong procedural coding results, including a single-file Three.js Minecraft clone, but noted formatting inconsistencies and verbosity versus Claude Opus 5.5. The article argues the deployment reflects stealth crowdsourced testing trends and rising Chinese frontier competitiveness in long-output, omnimodal systems.

Read original sourceIndependent coverage

Basic Tutorials

Claude Opus 5.5: Anthropic Launches Faster AI Model at a Competitive Price - Basic Tutorials

Anthropic has released Claude Opus 5.5, the first model in its new 5.5 generation, on September 22, 2026, just two months after Opus 5. The company describes it as its most powerful model yet, with major gains in coding and agent-based tasks while costing roughly 40 percent less than its predecessor.

Pricing now stands at $4 per million input tokens and $20 per million output tokens, with cache reads reduced to $0.20 per million. Anthropic reports text generation over 30 percent faster than Opus 5, benchmark gains on Terminal Bench 4.0, FrontierCode v1.1, and GDPval-AA v2.1, and 85 percent fewer attempts to circumvent security boundaries. Sonnet 5.5 and Haiku 5.5 are expected in the coming weeks.

Read original sourceIndependent coverage

myclaw.ai

GPT-6 Sol Review: Pricing, Benchmarks, and Best Uses

OpenAI released GPT-6 Sol on September 22, 2026, positioning it as the middle-tier option in the GPT-6 lineup for demanding coding and agent workflows that need judgment but not the highest price. The model carries the API ID gpt-6-sol, sits between Astra for hard tasks and Luna for narrow high-volume work, and is presented as a sensible upgrade target from the older GPT-5.6 Sol.

Standard API pricing is $2 per million input tokens and $10 per million output tokens, with cached input at $0.20 and cache writes at $2.50, and requests above 272,000 tokens shifting to long-context rates. OpenAI reports 68.8% on DeepSWE v1.1 at max effort, 33.2% on AutomationBench 1.0.6 at xhigh, and 60.5% on OSWorld 2.0 offline at xhigh, though the article notes these are vendor results without independent verification. The model accepts text and images, supports a 1,050,000-token context with 128,000-token output, and exposes web search, code interpreter, computer use, and other Responses API tools, but fine-tuning is unavailable and native audio or video input is not supported.

Read original sourceIndependent coverage

Google

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS

Google has introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, its most expressive audio generation models yet, expanding the Gemini Audio family with tools for creators, developers, and enterprises. The models enable users to design custom voices, replicate existing ones from short samples, and direct scenes line by line across surfaces including Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.

Gemini 3.8 Flash TTS is built for creative direction and character design with over 100 languages and access to 2,000+ production-ready voices, while Flash-Lite targets high-volume dubbing and voice agents. On Hume AI's Voice Design Benchmark, Flash TTS took the top spot with 71.4, and both models ranked first and second on the Overall Quality Index, with Flash-Lite also leading in accent modeling at 60.8. Both are rolling out to developers via the Gemini API and AI Studio starting September 23, 2026, with enterprise API access coming soon, while flash is in Gemini Notebook and Flash-Lite in Google Vids. Safety features include a mandatory verbal consent recording for replication, SynthID audio watermarking, and C2PA credentials.

Read original sourceOfficial source

This is a digest of recent coverage, not original reporting. Summaries are generated from saved source articles; follow the original links for full context. Prefer a feed of newly listed models? Subscribe to model releases.