DeepSeek-V4.1-Flash
DeepSeek released the 多模态 model · Launched on 2026-09
DeepSeek-V4.1-Flash is the smallest member of DeepSeek's new architecture family, released on September 10, 2026. It is a 552B-parameter MoE model built on the new Causal Encoder-Decoder (CED) asymmetric architecture, activating only 8B parameters for input and 16B for output, designed for the 'input-heavy, output-light' workloads of the Agent era. It supports a 1M-token context with up to 384K output, native multimodal vision understanding, and a KV cache that needs just 1/4 of the HBM of the previous generation. Its benchmarks comprehensively surpass the flagship V4 Pro, and API pricing is as low as ¥0.02/million tokens on cache hits, making it the most cost-effective open-source model of the Agent era.
核心特性
- CED 非对称架构(输入 8B / 输出 16B)
- 100万 tokens 上下文 + 384K 最大输出
- 原生多模态视觉理解
- KV Cache 压缩(HBM 降至 1/4、SSD 降至 1/8)
核心优势
- 基准全面超越 V4 Pro
- API 成本极低(缓存命中 ¥0.02/M)
- 输入侧轻量化高效读入
- 开源可商用
缓存命中输入 ¥0.02/百万 tokens、未命中 ¥1、输出 ¥4(空闲时段,高峰翻倍);开源免费
GPQA Diamond 90.9, Codeforces 3471, Terminal-Bench 2.1 90.6, CyberGym 88.1,全面超越 V4 Pro