Tempr
A revised Markaicode benchmark article focuses on GPT-OSS-20B as the model Groq and Ollama users run in 2026, citing Groq's published specification of 1,000 tokens/sec with a 128K context window and pricing of $0.075 input / $0.30 output per million tokens. The piece notes GPT-OSS-120B cannot fit on a single RTX 4090, making 20B the only fair cloud-vs-local matchup. Independent RTX 4090 benchmarks via Ollama report throughput ranging from 45 to 225 tok/s depending on context length and llama.cpp build, with no single authoritative local number. The article also documents Groq's mid-2026 model lineup, confirming GPT-OSS-20B and GPT-OSS-120B as Groq's mid-tier workhorses replacing the deprecated Mixtral endpoint.