bookmarks
3 bookmarks

Today we are releasing our speculative decoding implementation in our inference engine uzu. Initially for Qwen3.6 27B, with support for Qwen3.8 27B and Muse Glimmer coming soon. On Apple M5-series chips, we outperform MTPLX (MLX + speculative decoding) by almost 2x, and llama.c…↗

After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements. - 15%+ better token-generation efficiency from improved speculative decoding.↗

·@OpenAI·Jul 29, 2026·news·efficiency·gpt-5-6·gpu-kernel·sol

Excited to share a major milestone from Mirai Labs. We've just published our latest research on state-of-the-art speculative decoding for LLM interactivity. Local LLM inference runs at batch size = 1, so speculative decoding must scale to large draft budgets. Existing factorize…↗