Skip to content

DeepSeek-V4 Technical Report

Authors: DeepSeek-AI (2026)

arXiv: 2606.19348

Domains

ArchitectureLong ContextMixture of Experts

TLDR (English)

A trillion-parameter-class MoE flagship using sparse and hybrid attention with native ~1M-token context. The report claims roughly an order-of-magnitude reduction in long-context inference FLOPs and KV cache footprint versus V3.2, with weights released under MIT.

TLDR(中文)

万亿参数级 MoE 旗舰,采用稀疏注意力与混合注意力架构,原生支持约 1M token 上下文;技术报告称长上下文推理 FLOPs 与 KV cache 占用相比 V3.2 下降约一个数量级,权重以 MIT 协议开放。

Appears in These Articles

Co-cited Papers

These papers appear in the same articles as this one

Related Papers

Other papers in the same domain