DeepSeek-V4 Technical Report
arXiv: 2606.19348
Domains
TLDR (English)
A trillion-parameter-class MoE flagship using sparse and hybrid attention with native ~1M-token context. The report claims roughly an order-of-magnitude reduction in long-context inference FLOPs and KV cache footprint versus V3.2, with weights released under MIT.
TLDR(中文)
万亿参数级 MoE 旗舰,采用稀疏注意力与混合注意力架构,原生支持约 1M token 上下文;技术报告称长上下文推理 FLOPs 与 KV cache 占用相比 V3.2 下降约一个数量级,权重以 MIT 协议开放。
Appears in These Articles
Co-cited Papers
These papers appear in the same articles as this one
Related Papers
Other papers in the same domain