OpenRouter
Perceptron released Mk1.5 on September 25, 2026 as an update to its perception model line aimed at physical/embodied agents. According to the TPS Report write-up, Mk1.5 accepts four input modalities — text, image, video, and audio — and responds with natural-language text plus, when explicitly requested via an annotation format parameter, structured spatial outputs including points, bounding boxes, polygons, and temporal video clips. The model also supports graded reasoning via standard reasoning controls, function tool calling, and structured outputs via JSON Schema, with audio analysis on video soundtracks treated as an opt-in, per-request feature. Per the TPS article and corroborated by aggregator catalog data, Mk1.5 is priced at $0.15 per 1M input tokens and $1.50 per 1M output tokens, with a reported context window of roughly 36K–37K tokens (sources differ between 36K, 36.9K, and 37K). Mk1.5 succeeds Perceptron's prior flagship vision-language model Mk1, retaining its focus on video and embodied reasoning while adding audio input and structured spatial annotations. No third-party benchmark results were available at release.