RunInfra
KGP Talkie runs a hands-on benchmark of ornith-ai's Ornith-1.5-9B-GGUF and Ornith-1.5-35B-A3B-GGUF on a single RTX 5090 with 32 GB VRAM, using llama-server build 10448 on Windows 11, Q4_K_M quantization for all models, f16 KV cache with flash attention, and 512-token decodes via the raw /completion endpoint with ignore The article walks through reading Ornith's hybrid block layout, why decode speed tracks active parameters (and why that isn't unique to Ornith), KV-cache cost per token across the three architectures, and why 'max context that loads' on Windows is a misleading figure. It also notes two practical pitfalls: the --jinja f