Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
公开更多
AI 研究摘要
Structured state space models (SSMs) and attention mechanisms have evolved largely as separate sequence modeling paradigms. This work introduces Structured State Space Duality (SSD), establishing theoretical connections between structured SSMs and attention variants through semiseparable matrices and dual contraction representations. Exploiting this duality, the authors develop a block-decomposed matrix multiplication algorithm that computes 1-semiseparable selective SSMs utilizing matrix multiplication units on modern accelerators. Incorporating SSD into a refined parallel neural network block yields Mamba-2, which achieves substantial compute throughput improvements over Mamba-1 while matching or outperforming Transformers on autoregressive language modeling benchmarks.
由 CiteArk 生成
这意味着什么
Structured State Space Duality unifies state space models and attention mechanisms under a common structured matrix framework. By showing that 1-semiseparable SSMs correspond directly to structured masked attention, Mamba-2 enables SSM computation via tensor cores and matrix multiplication units rather than custom memory-bound recurrence scans. This bridges algorithmic principles between Transformer scaling and linear-time recurrent architectures, facilitating large-scale training optimizations like tensor parallelism while preserving constant-memory inference efficiency.
Run · Match · Repeat
Run · Match · Repeat 表示仓库中证据最充分的一条结论推进到哪一步,不代表论文整体复现覆盖度。
只有论文作者或可信机构对 Artifact 完成签名确认后,三环才会出现外圈。
尚未复现成功
0/11
当前还没有可验证结论获得成功复现证据。
研究结论
本次从论文中抽取的主要经验性结论,以及它们的计划覆盖和实时证据状态。点击结论就地展开详情。
Mamba-2-780M trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 61.7% on LAMBADA, 54.9% on HellaSwag, 72.0% on PIQA, 61.0% on Arc-Easy, 28.5% on Arc-Challenge, 60.2% on WinoGrande, 36.2% on OpenbookQA, and 53.5% average accuracy across tasks.等待复现尚未安排实验,未证实不可复现报告 61.7%
报告
61.7%
观测
—
Official public checkpoint state-spaces/mamba2-780m exists; deferred to prioritize testing smallest official checkpoint (130M) and large checkpoint (2.7B).
lambada_acc · 尚未评估
报告: 61.7%
hellaswag_acc · 尚未评估
报告: 54.9%
piqa_acc · 尚未评估
报告: 72%
arc_easy_acc · 尚未评估
报告: 61%
arc_challenge_acc · 尚未评估
报告: 28.5%
winogrande_acc · 尚未评估
报告: 60.2%
openbookqa_acc · 尚未评估
报告: 36.2%
average_acc · 尚未评估
报告: 53.5%
Mamba-2-370M trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 55.8% on LAMBADA, 46.9% on HellaSwag, 70.5% on PIQA, 54.9% on Arc-Easy, 26.9% on Arc-Challenge, 55.7% on WinoGrande, 32.4% on OpenbookQA, and 49.0% average accuracy across tasks.等待复现尚未安排实验,未证实不可复现报告 55.8%
报告
55.8%
观测
—
Official public checkpoint state-spaces/mamba2-370m exists; deferred to prioritize testing smallest official checkpoint (130M) and large checkpoint (2.7B) within compute envelope.
lambada_acc · 尚未评估
报告: 55.8%
hellaswag_acc · 尚未评估
报告: 46.9%
piqa_acc · 尚未评估
报告: 70.5%
arc_easy_acc · 尚未评估
报告: 54.9%
arc_challenge_acc · 尚未评估
报告: 26.9%
winogrande_acc · 尚未评估
报告: 55.7%
openbookqa_acc · 尚未评估
报告: 32.4%
average_acc · 尚未评估
报告: 49%
On an NVIDIA A100-80GB PCIe GPU with state dimension N=64, the SSD layer implementation achieves 2x to 8x speedup over Mamba-1's optimized fused associative scan across sequence lengths from 512 to 512K, and outperforms FlashAttention-2 at sequence lengths 2K and above.等待复现尚未安排实验,未证实不可复现报告 2 factor
报告
2 factor
观测
—
Hardware-sensitive throughput benchmark measured on NVIDIA A100-80GB PCIe GPU; available compute context offers GCP L4 (24GB VRAM) which allows cross-hardware profiling but cannot replicate exact A100 absolute execution timings.
speedup_factor_vs_mamba1 · 尚未评估
报告: 2 factor
speedup_factor_vs_mamba1 · 尚未评估
报告: 8 factor
Mamba-2-130M trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 43.9% on LAMBADA, 35.3% on HellaSwag, 64.9% on PIQA, 47.4% on Arc-Easy, 24.2% on Arc-Challenge, 52.1% on WinoGrande, 30.6% on OpenbookQA, and 42.6% average accuracy across tasks.等待复现报告 43.9%
报告
43.9%
观测
—
实验方案已生成,还没有运行记录
lambada_acc · 尚未评估
报告: 43.9%
hellaswag_acc · 尚未评估
报告: 35.3%
piqa_acc · 尚未评估
报告: 64.9%
arc_easy_acc · 尚未评估
报告: 47.4%
arc_challenge_acc · 尚未评估
报告: 24.2%
winogrande_acc · 尚未评估
报告: 52.1%
openbookqa_acc · 尚未评估
报告: 30.6%
average_acc · 尚未评估
报告: 42.6%
When pretrained on 300B tokens of the Pile, Mamba-2 achieves validation perplexities of 10.48 at 130M parameters, 8.21 at 370M parameters, 7.26 at 780M parameters, 6.66 at 1.3B parameters, and 6.09 at 2.7B parameters.等待复现尚未安排实验,未证实不可复现报告 10.48 perplexity
报告
10.48 perplexity
观测
—
Full Pile validation split evaluation across all 5 model scales requires large-scale dataset tokenization and forward passes; zero-shot downstream task evaluations test the same public checkpoints more directly.
pile_validation_perplexity · 尚未评估
报告: 10.48 perplexity
pile_validation_perplexity · 尚未评估
报告: 8.21 perplexity
pile_validation_perplexity · 尚未评估
报告: 7.26 perplexity
pile_validation_perplexity · 尚未评估
报告: 6.66 perplexity
pile_validation_perplexity · 尚未评估
报告: 6.09 perplexity
Mamba-2-1.3B trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 65.7% on LAMBADA, 59.9% on HellaSwag, 73.2% on PIQA, 64.3% on Arc-Easy, 33.3% on Arc-Challenge, 60.9% on WinoGrande, 37.8% on OpenbookQA, and 56.4% average accuracy across tasks.等待复现尚未安排实验,未证实不可复现报告 65.7%
报告
65.7%
观测
—
Official public checkpoint state-spaces/mamba2-1.3b exists; deferred to prioritize 130M and 2.7B evaluations.
lambada_acc · 尚未评估
报告: 65.7%
hellaswag_acc · 尚未评估
报告: 59.9%
piqa_acc · 尚未评估
报告: 73.2%
arc_easy_acc · 尚未评估
报告: 64.3%
arc_challenge_acc · 尚未评估
报告: 33.3%
winogrande_acc · 尚未评估
报告: 60.9%
openbookqa_acc · 尚未评估
报告: 37.8%
average_acc · 尚未评估
报告: 56.4%
复现与技术信息2
实验运行论文明确声明、且经 CiteArk 核验的代码仓库与固定版本。
不使用作者代码;CiteArk 根据论文独立实现实验协议,并记录所有生成文件和假设。
要形成可证伪实验,仍缺少关键方法细节、输入、数据划分、指标或评估规则。