Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Not assessedPlan blockedFindingclaim_fig13_switch_distillation_acceleration

Switch Distillation substantially accelerates reasoning acquisition during mid-training, surpassing the final 60B-token NTP reasoning macro-average within 2.5B tokens (24x fewer tokens), while standard FKD requires 5.0B tokens (12x fewer).

Source: source_paper:Figure 13, Page 32

Reported and observed measurements

No structured measurement is attached to this Claim.

Assessments (0)

No immutable Assessment has been published for this Claim yet.

Runs (0)

No execution has been linked to this Claim yet.