Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Public
Authors:Tri Dao, Albert Gu
More
Copy repository link

AI research summary

Structured state space models (SSMs) and attention mechanisms have evolved largely as separate sequence modeling paradigms. This work introduces Structured State Space Duality (SSD), establishing theoretical connections between structured SSMs and attention variants through semiseparable matrices and dual contraction representations. Exploiting this duality, the authors develop a block-decomposed matrix multiplication algorithm that computes 1-semiseparable selective SSMs utilizing matrix multiplication units on modern accelerators. Incorporating SSD into a refined parallel neural network block yields Mamba-2, which achieves substantial compute throughput improvements over Mamba-1 while matching or outperforming Transformers on autoregressive language modeling benchmarks.

Generated by CiteArk.

What this means

Structured State Space Duality unifies state space models and attention mechanisms under a common structured matrix framework. By showing that 1-semiseparable SSMs correspond directly to structured masked attention, Mamba-2 enables SSM computation via tensor cores and matrix multiplication units rather than custom memory-bound recurrence scans. This bridges algorithmic principles between Transformer scaling and linear-time recurrent architectures, facilitating large-scale training optimizations like tensor parallelism while preserving constant-memory inference efficiency.

Reproduction progress
RunMatchRepeat

Not reproduced yet

0/11

No verifiable claim has successful reproduction evidence yet.

Awaiting author or institution signature.

Research claims

The main empirical claims extracted in this research plan, with their plan coverage and live evidence state. Click a claim to expand it inline.

Mamba-2-780M trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 61.7% on LAMBADA, 54.9% on HellaSwag, 72.0% on PIQA, 61.0% on Arc-Easy, 28.5% on Arc-Challenge, 60.2% on WinoGrande, 36.2% on OpenbookQA, and 53.5% average accuracy across tasks.Awaiting reproductionExperiment not scheduled; feasibility remains unprovenReported 61.7%

Reported

61.7%

Observed

Official public checkpoint state-spaces/mamba2-780m exists; deferred to prioritize testing smallest official checkpoint (130M) and large checkpoint (2.7B).

lambada_acc · Not assessed

Reported: 61.7%

hellaswag_acc · Not assessed

Reported: 54.9%

piqa_acc · Not assessed

Reported: 72%

arc_easy_acc · Not assessed

Reported: 61%

arc_challenge_acc · Not assessed

Reported: 28.5%

winogrande_acc · Not assessed

Reported: 60.2%

openbookqa_acc · Not assessed

Reported: 36.2%

average_acc · Not assessed

Reported: 53.5%

Mamba-2-370M trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 55.8% on LAMBADA, 46.9% on HellaSwag, 70.5% on PIQA, 54.9% on Arc-Easy, 26.9% on Arc-Challenge, 55.7% on WinoGrande, 32.4% on OpenbookQA, and 49.0% average accuracy across tasks.Awaiting reproductionExperiment not scheduled; feasibility remains unprovenReported 55.8%

Reported

55.8%

Observed

Official public checkpoint state-spaces/mamba2-370m exists; deferred to prioritize testing smallest official checkpoint (130M) and large checkpoint (2.7B) within compute envelope.

lambada_acc · Not assessed

Reported: 55.8%

hellaswag_acc · Not assessed

Reported: 46.9%

piqa_acc · Not assessed

Reported: 70.5%

arc_easy_acc · Not assessed

Reported: 54.9%

arc_challenge_acc · Not assessed

Reported: 26.9%

winogrande_acc · Not assessed

Reported: 55.7%

openbookqa_acc · Not assessed

Reported: 32.4%

average_acc · Not assessed

Reported: 49%

On an NVIDIA A100-80GB PCIe GPU with state dimension N=64, the SSD layer implementation achieves 2x to 8x speedup over Mamba-1's optimized fused associative scan across sequence lengths from 512 to 512K, and outperforms FlashAttention-2 at sequence lengths 2K and above.Awaiting reproductionExperiment not scheduled; feasibility remains unprovenReported 2 factor

Reported

2 factor

Observed

Hardware-sensitive throughput benchmark measured on NVIDIA A100-80GB PCIe GPU; available compute context offers GCP L4 (24GB VRAM) which allows cross-hardware profiling but cannot replicate exact A100 absolute execution timings.

speedup_factor_vs_mamba1 · Not assessed

Reported: 2 factor

speedup_factor_vs_mamba1 · Not assessed

Reported: 8 factor

Mamba-2-130M trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 43.9% on LAMBADA, 35.3% on HellaSwag, 64.9% on PIQA, 47.4% on Arc-Easy, 24.2% on Arc-Challenge, 52.1% on WinoGrande, 30.6% on OpenbookQA, and 42.6% average accuracy across tasks.Awaiting reproductionReported 43.9%

Reported

43.9%

Observed

Experiment plan ready; no runs yet.

lambada_acc · Not assessed

Reported: 43.9%

hellaswag_acc · Not assessed

Reported: 35.3%

piqa_acc · Not assessed

Reported: 64.9%

arc_easy_acc · Not assessed

Reported: 47.4%

arc_challenge_acc · Not assessed

Reported: 24.2%

winogrande_acc · Not assessed

Reported: 52.1%

openbookqa_acc · Not assessed

Reported: 30.6%

average_acc · Not assessed

Reported: 42.6%

When pretrained on 300B tokens of the Pile, Mamba-2 achieves validation perplexities of 10.48 at 130M parameters, 8.21 at 370M parameters, 7.26 at 780M parameters, 6.66 at 1.3B parameters, and 6.09 at 2.7B parameters.Awaiting reproductionExperiment not scheduled; feasibility remains unprovenReported 10.48 perplexity

Reported

10.48 perplexity

Observed

Full Pile validation split evaluation across all 5 model scales requires large-scale dataset tokenization and forward passes; zero-shot downstream task evaluations test the same public checkpoints more directly.

pile_validation_perplexity · Not assessed

Reported: 10.48 perplexity

pile_validation_perplexity · Not assessed

Reported: 8.21 perplexity

pile_validation_perplexity · Not assessed

Reported: 7.26 perplexity

pile_validation_perplexity · Not assessed

Reported: 6.66 perplexity

pile_validation_perplexity · Not assessed

Reported: 6.09 perplexity

Mamba-2-1.3B trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 65.7% on LAMBADA, 59.9% on HellaSwag, 73.2% on PIQA, 64.3% on Arc-Easy, 33.3% on Arc-Challenge, 60.9% on WinoGrande, 37.8% on OpenbookQA, and 56.4% average accuracy across tasks.Awaiting reproductionExperiment not scheduled; feasibility remains unprovenReported 65.7%

Reported

65.7%

Observed

Official public checkpoint state-spaces/mamba2-1.3b exists; deferred to prioritize 130M and 2.7B evaluations.

lambada_acc · Not assessed

Reported: 65.7%

hellaswag_acc · Not assessed

Reported: 59.9%

piqa_acc · Not assessed

Reported: 73.2%

arc_easy_acc · Not assessed

Reported: 64.3%

arc_challenge_acc · Not assessed

Reported: 33.3%

winogrande_acc · Not assessed

Reported: 60.9%

openbookqa_acc · Not assessed

Reported: 37.8%

average_acc · Not assessed

Reported: 56.4%

Reproduction and technical details2
Implementation path
Official implementation2

The experiment runs a repository and commit explicitly declared by the paper and verified by CiteArk.

CiteArk reconstruction0

No author code is used. CiteArk independently rebuilds the paper-grounded protocol and records every generated file and assumption.

Information insufficient9

A falsifiable execution still needs a missing method detail, input, split, metric, or evaluation rule.

Signed research planVerifiedThe Research Plan CAP fixes paper sources, complete Claim coverage, planned Experiments, and source-declared values.Download research plan
Generated byCiteArk Official
Modelgoogle/gemini-3.8-flash
Completed
Artifact digestcef77c4073
Attestation statusTrusted signature verified
Output
View independent verification JSON