DeepSeek-V4.1-Flash Arrives on SiliconFlow: Faster, Cheaper, and Vision-Ready

SiliconFlow ·

Key Info

DeepSeek-V4.1-Flash is now live on SiliconFlow on day one, touted as a flagship performer: a 552B-parameter MoE model with ~8B active parameters during prefill and ~16B during decode, native vision, a 1M context window, and an MIT license.

Highlights

  • Faster inference and higher throughput, keeping Flash fast even at production scale.
  • Native vision support extends the model to multimodal tasks.
  • 1M context window with a roughly 4× smaller KV cache footprint than V4 Flash.
  • Production-ready on SiliconFlow today, available to build with immediately.
Loading...