Perplexity Unveils Custom Serving Infrastructure to Cut Embedding Search Costs
Key Info
Perplexity shared details of its custom serving infrastructure combining Ivy, Tulip, and ROSE, which cuts latency and cost while improving throughput for online and batch embedding workloads.
Highlights
- Ivy acts as the HTTP gateway handling CPU-side request prep, tokenization, and batch splitting before sending work to Tulip via gRPC.
- The full pipeline improves embedding search speed and reduces cost relative to off-the-shelf solutions.
- Perplexity published research detailing how it built state-of-the-art serving infrastructure behind its answers.