Today we are releasing our speculative decoding implementation in our inference engine uzu.
Initially for Qwen3.6 27B, with support for Qwen3.8 27B and Muse Glimmer coming soon.
On Apple M5-series chips, we outperform MTPLX (MLX + speculative decoding) by almost 2x, and llama.c…↗