Tencent's Hy4 Preview: Open-Source 770B MoE Model Lifts Inference Throughput by 31.8%
Key Info
Tencent's Hy4 preview, a 770B-parameter open-source MoE model with 49B active parameters and 1M context, is now available. The team reports that Hy4 independently identified inference bottlenecks and improved end-to-end throughput by 31.8% through operator fusion and communication optimizations, with gains stable across context lengths and concurrency levels.
Highlights
- Hy4 preview: 770B total parameters, 49B active, 1M context, open-source and positioned as an affordable frontier model.
- Self-discovered inference bottlenecks led to a 31.8% end-to-end throughput lift via operator fusion and communication optimizations.
- Performance gains remain stable across different context lengths and concurrency settings.
- Blog, Hugging Face, and GitHub resources are available for testing and feedback.