Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Not assessedPlan blockedFindingclaim_fig13_switch_distillation_acceleration
Switch Distillation substantially accelerates reasoning acquisition during mid-training, surpassing the final 60B-token NTP reasoning macro-average within 2.5B tokens (24x fewer tokens), while standard FKD requires 5.0B tokens (12x fewer).
Source: source_paper:Figure 13, Page 32
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.
Runs (0)
No execution has been linked to this Claim yet.