Add corrections, implementation notes, pricing changes, or usage caveats for other readers.
Jun 22, 2026
Input modalities
Output modalities
Capabilities
1,048,576 tokens
Recent tweets and retweets from Wafer
Flash Attention on blackwell is the kernel that BROKE Triton.
so the triton team dropped down a level, built a language called gluon, and wrote it by hand.
here's what changed, and what gluon gives the kernel programmer back:
triton's deal was almost magical when it came…
in speculative decoding, the draft sits idle during verification. SSD eliminates the wait, pre-computing speculations for every likely verification outcome in parallel.
here's how it works at the lowest level:
the problem being tackled is that in standard speculative…
I'm in Vegas for @Ai4Conferences from August 4th - 6th. @wafer_ai is a proud sponsor this year!
Booth K615
If you're building voice agents or running performance-sensitive LLM workloads in production, i'd love to chat!
Discuss this model
Add corrections, implementation notes, pricing changes, or usage caveats for other readers.