Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Not assessedPlan blockedMeasurementclaim_t1_ntp

Standard Next-Token Prediction (NTP) baseline downstream evaluation performance after 60B mid-training tokens across 15 tasks covering Reasoning, Factual Recall, and Knowledge & Commonsense.

Source: source_paper:Table 1, Page 9

Reported and observed measurements

downstream_accuracy

t1_ntp_gsm8k

Reported 40.4 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_gsm_s

Reported 29.8 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_gsmplus

Reported 23.1 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_bbh

Reported 29.9 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_drop

Reported 29.8 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_math

Reported 3.8 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_tqa

Reported 56.7 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_nq

Reported 25.5 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_sqa

Reported 8.7 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_mmlu

Reported 43.6 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_mmlu_p

Reported 15.5 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_arc_c

Reported 51.1 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_obqa

Reported 51.8 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_wino

Reported 51.4 percentage_points

Observed — percentage_points

downstream_accuracy

t1_ntp_agi

Reported 34.1 percentage_points

Observed — percentage_points

Assessments (0)

No immutable Assessment has been published for this Claim yet.

Runs (0)

No execution has been linked to this Claim yet.