Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
PublicMore
AI research summary
Structured state space models (SSMs) and attention mechanisms have evolved largely as separate sequence modeling paradigms. This work introduces Structured State Space Duality (SSD), establishing theoretical connections between structured SSMs and attention variants through semiseparable matrices and dual contraction representations. Exploiting this duality, the authors develop a block-decomposed matrix multiplication algorithm that computes 1-semiseparable selective SSMs utilizing matrix multiplication units on modern accelerators. Incorporating SSD into a refined parallel neural network block yields Mamba-2, which achieves substantial compute throughput improvements over Mamba-1 while matching or outperforming Transformers on autoregressive language modeling benchmarks.
Generated by CiteArk.
What this means
Structured State Space Duality unifies state space models and attention mechanisms under a common structured matrix framework. By showing that 1-semiseparable SSMs correspond directly to structured masked attention, Mamba-2 enables SSM computation via tensor cores and matrix multiplication units rather than custom memory-bound recurrence scans. This bridges algorithmic principles between Transformer scaling and linear-time recurrent architectures, facilitating large-scale training optimizations like tensor parallelism while preserving constant-memory inference efficiency.
Run · Match · Repeat
Run · Match · Repeat shows the deepest evidence stage reached by any claim in this repository, not paper-wide reproduction coverage.
An outer circle appears only after a paper author or trusted institution signs the Artifact.
Not reproduced yet
0/11
No verifiable claim has successful reproduction evidence yet.
Research claims
The main empirical claims extracted in this research plan, with their plan coverage and live evidence state. Click a claim to expand it inline.
Mamba-2-780M trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 61.7% on LAMBADA, 54.9% on HellaSwag, 72.0% on PIQA, 61.0% on Arc-Easy, 28.5% on Arc-Challenge, 60.2% on WinoGrande, 36.2% on OpenbookQA, and 53.5% average accuracy across tasks.Awaiting reproductionExperiment not scheduled; feasibility remains unprovenReported 61.7%
Reported
61.7%
Observed
—
Official public checkpoint state-spaces/mamba2-780m exists; deferred to prioritize testing smallest official checkpoint (130M) and large checkpoint (2.7B).
lambada_acc · Not assessed
Reported: 61.7%
hellaswag_acc · Not assessed
Reported: 54.9%
piqa_acc · Not assessed
Reported: 72%
arc_easy_acc · Not assessed
Reported: 61%
arc_challenge_acc · Not assessed
Reported: 28.5%
winogrande_acc · Not assessed
Reported: 60.2%
openbookqa_acc · Not assessed
Reported: 36.2%
average_acc · Not assessed
Reported: 53.5%
Mamba-2-370M trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 55.8% on LAMBADA, 46.9% on HellaSwag, 70.5% on PIQA, 54.9% on Arc-Easy, 26.9% on Arc-Challenge, 55.7% on WinoGrande, 32.4% on OpenbookQA, and 49.0% average accuracy across tasks.Awaiting reproductionExperiment not scheduled; feasibility remains unprovenReported 55.8%
Reported
55.8%
Observed
—
Official public checkpoint state-spaces/mamba2-370m exists; deferred to prioritize testing smallest official checkpoint (130M) and large checkpoint (2.7B) within compute envelope.
lambada_acc · Not assessed
Reported: 55.8%
hellaswag_acc · Not assessed
Reported: 46.9%
piqa_acc · Not assessed
Reported: 70.5%
arc_easy_acc · Not assessed
Reported: 54.9%
arc_challenge_acc · Not assessed
Reported: 26.9%
winogrande_acc · Not assessed
Reported: 55.7%
openbookqa_acc · Not assessed
Reported: 32.4%
average_acc · Not assessed
Reported: 49%
On an NVIDIA A100-80GB PCIe GPU with state dimension N=64, the SSD layer implementation achieves 2x to 8x speedup over Mamba-1's optimized fused associative scan across sequence lengths from 512 to 512K, and outperforms FlashAttention-2 at sequence lengths 2K and above.Awaiting reproductionExperiment not scheduled; feasibility remains unprovenReported 2 factor
Reported
2 factor
Observed
—
Hardware-sensitive throughput benchmark measured on NVIDIA A100-80GB PCIe GPU; available compute context offers GCP L4 (24GB VRAM) which allows cross-hardware profiling but cannot replicate exact A100 absolute execution timings.
speedup_factor_vs_mamba1 · Not assessed
Reported: 2 factor
speedup_factor_vs_mamba1 · Not assessed
Reported: 8 factor
Mamba-2-130M trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 43.9% on LAMBADA, 35.3% on HellaSwag, 64.9% on PIQA, 47.4% on Arc-Easy, 24.2% on Arc-Challenge, 52.1% on WinoGrande, 30.6% on OpenbookQA, and 42.6% average accuracy across tasks.Awaiting reproductionReported 43.9%
Reported
43.9%
Observed
—
Experiment plan ready; no runs yet.
lambada_acc · Not assessed
Reported: 43.9%
hellaswag_acc · Not assessed
Reported: 35.3%
piqa_acc · Not assessed
Reported: 64.9%
arc_easy_acc · Not assessed
Reported: 47.4%
arc_challenge_acc · Not assessed
Reported: 24.2%
winogrande_acc · Not assessed
Reported: 52.1%
openbookqa_acc · Not assessed
Reported: 30.6%
average_acc · Not assessed
Reported: 42.6%
When pretrained on 300B tokens of the Pile, Mamba-2 achieves validation perplexities of 10.48 at 130M parameters, 8.21 at 370M parameters, 7.26 at 780M parameters, 6.66 at 1.3B parameters, and 6.09 at 2.7B parameters.Awaiting reproductionExperiment not scheduled; feasibility remains unprovenReported 10.48 perplexity
Reported
10.48 perplexity
Observed
—
Full Pile validation split evaluation across all 5 model scales requires large-scale dataset tokenization and forward passes; zero-shot downstream task evaluations test the same public checkpoints more directly.
pile_validation_perplexity · Not assessed
Reported: 10.48 perplexity
pile_validation_perplexity · Not assessed
Reported: 8.21 perplexity
pile_validation_perplexity · Not assessed
Reported: 7.26 perplexity
pile_validation_perplexity · Not assessed
Reported: 6.66 perplexity
pile_validation_perplexity · Not assessed
Reported: 6.09 perplexity
Mamba-2-1.3B trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 65.7% on LAMBADA, 59.9% on HellaSwag, 73.2% on PIQA, 64.3% on Arc-Easy, 33.3% on Arc-Challenge, 60.9% on WinoGrande, 37.8% on OpenbookQA, and 56.4% average accuracy across tasks.Awaiting reproductionExperiment not scheduled; feasibility remains unprovenReported 65.7%
Reported
65.7%
Observed
—
Official public checkpoint state-spaces/mamba2-1.3b exists; deferred to prioritize 130M and 2.7B evaluations.
lambada_acc · Not assessed
Reported: 65.7%
hellaswag_acc · Not assessed
Reported: 59.9%
piqa_acc · Not assessed
Reported: 73.2%
arc_easy_acc · Not assessed
Reported: 64.3%
arc_challenge_acc · Not assessed
Reported: 33.3%
winogrande_acc · Not assessed
Reported: 60.9%
openbookqa_acc · Not assessed
Reported: 37.8%
average_acc · Not assessed
Reported: 56.4%
Reproduction and technical details2
The experiment runs a repository and commit explicitly declared by the paper and verified by CiteArk.
No author code is used. CiteArk independently rebuilds the paper-grounded protocol and records every generated file and assumption.
A falsifiable execution still needs a missing method detail, input, split, metric, or evaluation rule.
Signed research planVerifiedThe Research Plan CAP fixes paper sources, complete Claim coverage, planned Experiments, and source-declared values.Download research plan
cef77c4073