speculative decoding
3 saves·entity·peaked week of Jul 6, 2026
Weekly volume
Shows up alongside
Topics tagged on the same saves.
Who’s driving it
Most-saved authors on this topic.
Recent saves
The items behind the chart, newest first.
Today we are releasing our speculative decoding implementation in our inference engine uzu. Initially for Qwen3.6 27B, with support for Qwen3.8 27B and Muse Glimmer coming soon. On Apple M5-series …
@trymirai·Sep 3, 2026↗
After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements…
@OpenAI·Jul 29, 2026↗
Excited to share a major milestone from Mirai Labs. We've just published our latest research on state-of-the-art speculative decoding for LLM interactivity. Local LLM inference runs at batch size =…
@dmitrshvets·Jul 10, 2026↗