Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Not assessedPlannedFindingclaim-downstream-zero-shot-130m

Mamba-2-130M trained on 300B tokens of the Pile achieves zero-shot downstream task performance of 43.9% on LAMBADA, 35.3% on HellaSwag, 64.9% on PIQA, 47.4% on Arc-Easy, 24.2% on Arc-Challenge, 52.1% on WinoGrande, 30.6% on OpenbookQA, and 42.6% average accuracy across tasks.

Source: source-paper:Table 10, page 52 (and Table 1, page 29)

Reported and observed measurements

lambada_acc

rm-130m-lambada-acc

Reported 43.9 percent

Observed — percent

hellaswag_acc

rm-130m-hellaswag-acc

Reported 35.3 percent

Observed — percent

piqa_acc

rm-130m-piqa-acc

Reported 64.9 percent

Observed — percent

arc_easy_acc

rm-130m-arc-e-acc

Reported 47.4 percent

Observed — percent

arc_challenge_acc

rm-130m-arc-c-acc

Reported 24.2 percent

Observed — percent

winogrande_acc

rm-130m-winogrande-acc

Reported 52.1 percent

Observed — percent

openbookqa_acc

rm-130m-openbookqa-acc

Reported 30.6 percent

Observed — percent

average_acc

rm-130m-avg-acc

Reported 42.6 percent

Observed — percent

Assessments (0)

No immutable Assessment has been published for this Claim yet.

Runs (0)

No execution has been linked to this Claim yet.