Registry collection cards used to be a paragraph of plain text.
Now they pull the same rich, interactive blocks as Reports, so a collection actually reads like a model card.
We also shipped artifact panel grids to compare metrics across versions. 📊
Let's say your robotic policy fails a task. Your success metrics look fine.
What would you actually trust to tell you what went wrong?
Poll
20% — 📊 Scalar metrics
12% — 🎥 Rollout video
14% — 🗂️ Failure taxonomy
54% — 🦾 Real-world test
59 votes • 3 days
Discuss this model
Add corrections, implementation notes, pricing changes, or usage caveats for other readers.