Excited to share a major milestone from Mirai Labs.
We've just published our latest research on state-of-the-art speculative decoding for LLM interactivity.
Local LLM inference runs at batch size = 1, so speculative decoding must scale to large draft budgets. Existing factorize…↗