Claudinio
Claudinio published a first-party blog post on 1 August 2026 describing the internal coding-model evaluation suite it uses to decide which model to route traffic through. The suite contains 108 tasks and is run against Claudinio's live proxy rather than directly against a provider's API, so each score reflects the full Tasks split into 21 execution items scored by running the model's generated code against reference pytest tests in a plain subprocess (with a stripped environment and timeout, explicitly not Docker) and 87 judge items scored by another model against a behavioural rubric covering planning, verification, language adheren