跳转到内容

DeepSeek-V4 Technical Report

作者: DeepSeek-AI (2026)

arXiv: 2606.19348

领域

架构长上下文混合专家

TLDR(中文)

万亿参数级 MoE 旗舰,采用稀疏注意力与混合注意力架构,原生支持约 1M token 上下文;技术报告称长上下文推理 FLOPs 与 KV cache 占用相比 V3.2 下降约一个数量级,权重以 MIT 协议开放。

TLDR (English)

A trillion-parameter-class MoE flagship using sparse and hybrid attention with native ~1M-token context. The report claims roughly an order-of-magnitude reduction in long-context inference FLOPs and KV cache footprint versus V3.2, with weights released under MIT.

出现在这些文章里

同被引用

这些论文与本文出现在同一篇文章中

相关论文

同一领域的其他论文