bookmarks
1 bookmark

Excited to share a major milestone from Mirai Labs. We've just published our latest research on state-of-the-art speculative decoding for LLM interactivity. Local LLM inference runs at batch size = 1, so speculative decoding must scale to large draft budgets. Existing factorize…↗