DeepSeek-V4 Technical Report
arXiv: 2606.19348
领域
TLDR(中文)
万亿参数级 MoE 旗舰,采用稀疏注意力与混合注意力架构,原生支持约 1M token 上下文;技术报告称长上下文推理 FLOPs 与 KV cache 占用相比 V3.2 下降约一个数量级,权重以 MIT 协议开放。
TLDR (English)
A trillion-parameter-class MoE flagship using sparse and hybrid attention with native ~1M-token context. The report claims roughly an order-of-magnitude reduction in long-context inference FLOPs and KV cache footprint versus V3.2, with weights released under MIT.
出现在这些文章里
同被引用
这些论文与本文出现在同一篇文章中
相关论文
同一领域的其他论文