Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Not assessedPlan blockedFindingclaim_fig12_gradient_alignment_fkl_rkl

Forward KL consistently exhibits higher CE-KL gradient cosine alignment than Reverse KL across pre-training and mid-training, with the gap widening at higher alpha and later training steps.

Source: source_paper:Figure 12, Page 31

Reported and observed measurements

No structured measurement is attached to this Claim.

Assessments (0)

No immutable Assessment has been published for this Claim yet.

Runs (0)

No execution has been linked to this Claim yet.