Perplexity Benchmarks Its Open-Source Engine Lily: 1.23× Prefill and 1.35× Decode Gains vs MLX-LM

Perplexity ·

Key Info

Perplexity benchmarked its open-source Apple-silicon inference engine Lily on an M5 Max MacBook Pro and found it consistently outperforms MLX-LM across ten prompt lengths and ten decode contexts.

Highlights

  • Lily averaged 1.23× higher prefill throughput and 1.35× higher decode throughput than MLX-LM.
  • Output quality remained effectively unchanged in the comparison.
  • The benchmark focused on Qwen3.6-35B-A3B, the model Lily was built to specialize on.
  • Lily is designed for hybrid compute in Perplexity Computer and is now available as an open-source local inference engine.
Loading...