Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

Perceptron: Perceptron Mk1.5

Perceptron Mk1.5 is a multimodal model built by Perceptron, Inc. and released on September 25, 2026 as the company's most performant public model to date. According to Perceptron's documentation and launch blog, the model accepts text, images, video, and audio as input, with audio handled in WAV, MP3, or FLAC formats and subject to a per-item audio token limit that shares the overall context window with other inputs. Beyond plain text generation, Mk1.5 is designed to emit spatial and temporal annotations such as points, boxes, polygons, clips, and timestamped object tracks, making it suited to applications that require both language reasoning and grounded visual or audiovisual output.

Perceptron positions Mk1.5 around embodied-agent use cases, describing deployments on drones, quadrupeds, smart glasses, and smartphones without platform-specific retraining, and highlighting new capabilities added over earlier Perceptron models, including native audio support, video tracking, web search, sub-agent calls, and more complex visual reasoning. Practically, the model supports configurable reasoning effort, function calling on chat completions, constrained responses via JSON Schema and regex, and a "multilook" mode that runs independent prompts over shared context, giving developers flexible ways to structure agentic workflows that combine perception and action across physical embodiments.

Kilo Gatewayperceptron/perceptron-mk1.5

Quick Info

Powered by
Provider
Kilo Gateway
Model key
perceptron/perceptron-mk1.5
Release date
Sep 25, 2026
Last updated
Sep 25, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$1.50

Limits

Output tokens
8,192 tokens
Context window
36,864 tokens

Latest news about Perceptron: Perceptron Mk1.5

Kilo Gateway

Coverage

Perceptron released Mk1.5 on September 25, 2026, as a multimodal perception model built for physical agents. It accepts text, image, video, and audio inputs, and outputs natural-language text alongside optional structured spatial annotations including points, bounding boxes, polygons, and temporal clips. Pricing is set at $0.15 per million input tokens and $1.50 per million output tokens. Mk1.5 succeeds Perceptron's prior flagship vision-language model Mk1, extending its support to four input modalities and structured annotation outputs. The model supports graded reasoning via standard reasoning controls, function tool calling, and structured outputs via JSON Schema. Audio analysis of video soundtracks is opt-in and only runs when explicitly enabled per request.

Kilo Gateway

CoverageBenchmark

A third-party model profile confirms Perceptron Mk1.5 as an embodied reasoning model for physical agents, released September 25, 2026. It accepts text, image, video, and audio inputs and returns text plus optional structured annotations such as points, boxes, polygons, and tracks. API pricing is listed at $0.15 per million input tokens and $1.50 per million output tokens. The listing reports a 36,900-token context window with 8.2K maximum output, along with capability flags for reasoning, vision, function calling, JSON mode, and streaming. These specifications align with the model's positioning as a multimodal perception system for robotics and physical-AI applications.

Videos about Perceptron: Perceptron Mk1.5