Tencent's Sherry Quantization Shrinks Hy4 Preview from 1.5TB to 214GB with Minimal Accuracy Loss

Tencent Hy ·

Key Info

Tencent's Hy4 preview model can be compressed from 1.5TB to just 214GB using the Sherry quantization method, which packs weights down to 1.25 bits each while keeping accuracy almost intact. The quantized GGUF and original model are both available for download.

Highlights

  • Model size drops about 7x (1.5TB → 214GB) with minimal impact on performance.
  • Benchmark deltas are small: MCP Atlas 83.7→83.2, SWE-Bench multi 82.9→81.3, MRCR 81.3→81.1, IFBench 73.5→72.5.
  • Sherry also enables distributed inference: GPUs already available across machines can be stitched together to run as one system.
  • Tencent has released both the quantized GGUF and the original model weights.
Loading...