Perplexity Makes Lily, Its Apple Silicon AI Inference Engine, Generally Available
Key Info
Perplexity has announced that Lily, the local inference engine it built for hybrid compute, is now available. Lily treats Apple silicon as a distinct inference platform and maps Qwen's operations directly to its compute and memory architecture.
Highlights
- Lily is specialized for running Qwen3.6-35B-A3B on Apple silicon, using the chip's architecture as the basis for inference.
- Designed for on-device compute in Perplexity Computer, Lily is now publicly accessible via the link in the announcement.