Perplexity Benchmarks Its Open-Source Engine Lily: 1.23× Prefill and 1.35× Decode Gains vs MLX-LM
Key Info
Perplexity benchmarked its open-source Apple-silicon inference engine Lily on an M5 Max MacBook Pro and found it consistently outperforms MLX-LM across ten prompt lengths and ten decode contexts.
Highlights
- Lily averaged 1.23× higher prefill throughput and 1.35× higher decode throughput than MLX-LM.
- Output quality remained effectively unchanged in the comparison.
- The benchmark focused on Qwen3.6-35B-A3B, the model Lily was built to specialize on.
- Lily is designed for hybrid compute in Perplexity Computer and is now available as an open-source local inference engine.