Add corrections, implementation notes, pricing changes, or usage caveats for other readers.
Last updated
Aug 2, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities
1,048,576 tokens
Recent tweets and retweets from Chutes
Week 12. A shipping week for the confidential-compute stack:
→ Next TEE release is staged and in testing: mutual-TLS on attestation, hardware-attested registry access, and measurements third parties can verify themselves
→ Every GPU node on Chutes is now TEE-verified. The…
Article
Last Week in Chutes | July 29 to August 4, 2026
Last Week in Chutes
July 29 to August 4, 2026
Week twelve. A shipping week for the confidential-compute stack. The next TEE release is staged and in testing, the legacy non-TEE verification path is on
Recurrent attention would not sync. Jon Durbin tried the existing decentralized algorithms, the whole DiLoco family included, and none of them translate.
So he built a new sync method into Parallax, exact and fully non-blocking, within 0.6% of the centralized DDP baseline at…
165,335 tokens per second. That is Parallax pre-training a 5B test model on eight consumer RTX 5090s.
Optimized DDP training on eight RTX Pro 6000s managed 147,708. Hardware that costs roughly 3x more.
Parallax won even on the cheaper hardware.
Discuss this model
Add corrections, implementation notes, pricing changes, or usage caveats for other readers.