Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Public
Authors:Jacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li, Hanqing Zeng, Yinglong Xia, Benyu Zhang, Zhuokai Zhao, Qiang Zhang, Pang Wei Koh, Luke Zettlemoyer, Wen-tau Yih
More
Copy repository link

Reproduction

Start with coverage and measured results; open a run only when you need evidence or technical details.

RunMatchRepeat

0/76

claims supported by evidence

0

Supported

0

Challenged or mixed

0

Contradicted

1

Inconclusive

75

Not assessed

Supporting Assessment

Measurement confirmed

Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).

Reported

0.815 score

Observed

0.8079 score

Difference -0.0071

Run history

Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.

Failure log

Failed paths grouped by cause — check them before reproducing.

Data access1 failures

Claim–experiment reproduction matrix

See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions.

2 targets·1 with evidence·0 active·1 need attention

Mid-training ablation results for Always CE relative difference across Reasoning, Factual Recall, and Knowledge task groups.

Information insufficientclaim_t3_always_ce0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage RLVR1 for 13b_rkd.

Information insufficientclaim_t8_rlvr1_13b_rkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

The stage-dependent distillation tradeoff generalizes to the SmolLM family (SmolLM2 360M student, 1.7B Instruct teacher, SmolLM3 Stage-3 data), where mid-training KD exhibits the reasoning-recall deficit and Switch Distillation mitigates it.

Information insufficientclaim_fig7_smollm_tradeoff_generalizability0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage RLVR1 for 13b_fkd.

Information insufficientclaim_t8_rlvr1_13b_fkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage SFT for 13b_rkd.

Information insufficientclaim_t8_sft_13b_rkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage SFT for 13b_sd.

Information insufficientclaim_t8_sft_13b_sd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Factual acquisition stratification by teacher entropy quintile persists when using OLMo-2 1B Instruct and 13B Instruct teachers.

Information insufficientclaim_fig10_factual_acquisition_1b_13b0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Teacher entropy under OLMo-2 7B Instruct strongly predicts factual acquisition under NTP: by the end of pre-training, the student learns 67% of Q1 facts vs 5% of Q5 facts; by mid-training initialization (4T tokens), 80% of Q1 facts are learned vs 7% of Q5 facts.

Information insufficientclaim_fig4_factual_acquisition_quintiles0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage RLVR1 for ntp_shared.

Information insufficientclaim_t8_rlvr1_ntp_shared0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage SFT for 7b_rkd.

Information insufficientclaim_t8_sft_7b_rkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Switch Distillation (SD, q=20%, tau=2) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.

Information insufficientclaim_t1_13b_sd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q30.

Information insufficientclaim_t4_7b_q300 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Per-task downstream accuracy for mid-training ablation random_routing using OLMo-2 7B Instruct teacher.

Information insufficientclaim_t9_random_routing0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Per-task downstream accuracy for mid-training ablation teacher_top1 using OLMo-2 7B Instruct teacher.

Information insufficientclaim_t9_teacher_top10 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Mid-training ablation results for Switch Distillation absolute macro-averages across Reasoning, Factual Recall, and Knowledge task groups.

Information insufficientclaim_t3_sd_baseline0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage SFT for 7b_sd.

Information insufficientclaim_t8_sft_7b_sd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q10.

Information insufficientclaim_t4_13b_q100 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage DPO for 7b_trkd.

Information insufficientclaim_t8_dpo_7b_trkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Forward KL Distillation (FKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.

Information insufficientclaim_t1_7b_fkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Forward KL consistently exhibits higher CE-KL gradient cosine alignment than Reverse KL across pre-training and mid-training, with the gap widening at higher alpha and later training steps.

Information insufficientclaim_fig12_gradient_alignment_fkl_rkl0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B TRKD post-training performance.

Information insufficientclaim_t2_7b_trkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Teacher predictive entropy distinguishes procedural from knowledge-intensive domains across diverse instruction-tuned open-weight model families (OLMo-3 7B Instruct: 0.771, Qwen 3 8B: 0.705, Gemma-3 12B it: 0.707, Granite 3.3 8B Instruct: 0.696).

CiteArk reconstructionclaim_fig9_asymmetric_supervision_families1 plan0 runs
Scientific conclusionNot assessed

Evaluate Teacher Supervision Asymmetry (ROC AUC) Across Model Families

exp_fig9_asymmetric_supervision_families1 attempt
Execution failed

Next step

A person must review the failure and evidence integrity.

View local evidence pathExpand
  1. Research plan

    Claim and experiment binding established

  2. Execution target

    Execution failed

  3. Run attempt

    Experiment execution · Running

  4. CAP evidence

    No linked immutable Artifact yet

  5. Scientific judgment

    No Assessment yet

Legacy task without target-level resource requirements

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B FKD post-training performance.

Information insufficientclaim_t2_7b_fkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage DPO for 13b_sd.

Information insufficientclaim_t8_dpo_13b_sd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Mid-training ablation results for Oracle Domain Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.

Information insufficientclaim_t3_oracle_domain0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Gold-token probability, gradient attenuation relative to NTP, and factual recall deficits under KD hold consistently for OLMo-2 1B and 13B Instruct teachers.

Information insufficientclaim_fig11_factual_recall_kd_analysis_1b_13b0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage RLVR1 for 7b_rkd.

Information insufficientclaim_t8_rlvr1_7b_rkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Reverse KL Distillation (RKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.

Information insufficientclaim_t1_13b_rkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage SFT for 7b_trkd.

Information insufficientclaim_t8_sft_7b_trkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Mid-training ablation results for Teacher Top-1 Labels relative difference across Reasoning, Factual Recall, and Knowledge task groups.

Information insufficientclaim_t3_teacher_top10 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage SFT for 13b_fkd.

Information insufficientclaim_t8_sft_13b_fkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Per-task downstream accuracy for mid-training ablation oracle_domain using OLMo-2 7B Instruct teacher.

Information insufficientclaim_t9_oracle_domain0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B FKD post-training performance.

Information insufficientclaim_t2_13b_fkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q10.

Information insufficientclaim_t4_7b_q100 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B TRKD post-training performance.

Information insufficientclaim_t2_13b_trkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B RKD post-training performance.

Information insufficientclaim_t2_7b_rkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Switch Distillation substantially accelerates reasoning acquisition during mid-training, surpassing the final 60B-token NTP reasoning macro-average within 2.5B tokens (24x fewer tokens), while standard FKD requires 5.0B tokens (12x fewer).

Information insufficientclaim_fig13_switch_distillation_acceleration0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage RLVR1 for 7b_fkd.

Information insufficientclaim_t8_rlvr1_7b_fkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage DPO for 13b_fkd.

Information insufficientclaim_t8_dpo_13b_fkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Switch Distillation (SD, q=20%, tau=2) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher, demonstrating substantial gains on reasoning while maintaining factual recall.

Information insufficientclaim_t1_7b_sd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage DPO for 13b_rkd.

Information insufficientclaim_t8_dpo_13b_rkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Token-Routing Knowledge Distillation (TRKD) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.

Information insufficientclaim_t1_7b_trkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Forward KL Distillation (FKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.

Information insufficientclaim_t1_13b_fkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Standard Next-Token Prediction (NTP) baseline downstream evaluation performance after 60B mid-training tokens across 15 tasks covering Reasoning, Factual Recall, and Knowledge & Commonsense.

Information insufficientclaim_t1_ntp0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage SFT for ntp_shared.

Information insufficientclaim_t8_sft_ntp_shared0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Token-Routing Knowledge Distillation (TRKD) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.

Information insufficientclaim_t1_13b_trkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Lower teacher predictive entropy corresponds to substantially higher teacher top-1 agreement with the ground-truth token across all data domains and teacher sizes.

Information insufficientclaim_fig3_top1_agreement_entropy0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage DPO for 7b_rkd.

Information insufficientclaim_t8_dpo_7b_rkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B Switch Distillation post-training performance.

Information insufficientclaim_t2_13b_sd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Higher teacher entropy corresponds to strictly lower gold-token probability, attenuates the gold-token gradient relative to NTP (reaching ~0.5x NTP for Q5 facts), and produces larger downstream factual recall deficits.

Information insufficientclaim_fig5_gradient_attenuation_factual_supervision0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage DPO for 7b_sd.

Information insufficientclaim_t8_dpo_7b_sd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Across distillation strengths alpha in [0.0, 1.0], KL directions (forward and reverse), and teacher sizes (1B, 7B, 13B), mid-training distillation traces a reasoning-recall frontier that falls below NTP, which Switch Distillation mitigates.

Information insufficientclaim_fig2_reasoning_recall_tradeoff_sweep0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage DPO for ntp_shared.

Information insufficientclaim_t8_dpo_ntp_shared0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage SFT for 13b_trkd.

Information insufficientclaim_t8_sft_13b_trkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Knowledge distillation generally improves reasoning, but its effect on factual recall changes across training stages: during pre-training, distillation improves both reasoning and factual recall over NTP, whereas during mid-training it improves reasoning at the expense of factual recall.

Information insufficientclaim_fig1_reasoning_recall_stage_dependence0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q30.

Information insufficientclaim_t4_13b_q300 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage RLVR1 for 7b_sd.

Information insufficientclaim_t8_rlvr1_7b_sd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q20.

Information insufficientclaim_t4_13b_q200 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Mid-training ablation results for Switch Distillation with FKL relative difference across Reasoning, Factual Recall, and Knowledge task groups.

Information insufficientclaim_t3_sd_fkl0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage DPO for 13b_trkd.

Information insufficientclaim_t8_dpo_13b_trkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B Switch Distillation post-training performance.

Information insufficientclaim_t2_7b_sd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Reverse KL Distillation (RKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.

Information insufficientclaim_t1_7b_rkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage SFT for 7b_fkd.

Information insufficientclaim_t8_sft_7b_fkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage RLVR1 for 13b_sd.

Information insufficientclaim_t8_rlvr1_13b_sd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Per-task downstream accuracy for mid-training ablation sd_fkl using OLMo-2 7B Instruct teacher.

Information insufficientclaim_t9_sd_fkl0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B RKD post-training performance.

Information insufficientclaim_t2_13b_rkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Mid-training ablation results for Random Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.

Information insufficientclaim_t3_random_routing0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage RLVR1 for 7b_trkd.

Information insufficientclaim_t8_rlvr1_7b_trkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Per-task downstream accuracy for mid-training ablation teacher_correct using OLMo-2 7B Instruct teacher.

Information insufficientclaim_t9_teacher_correct0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q20.

Information insufficientclaim_t4_7b_q200 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for NTP post-training baseline performance.

Information insufficientclaim_t2_ntp_shared0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Mid-training ablation results for Teacher-Correct Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.

Information insufficientclaim_t3_teacher_correct0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Per-task downstream accuracy for mid-training ablation always_ce using OLMo-2 7B Instruct teacher.

Information insufficientclaim_t9_always_ce0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage RLVR1 for 13b_trkd.

Information insufficientclaim_t8_rlvr1_13b_trkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Downstream task evaluation after post-training stage DPO for 7b_fkd.

Information insufficientclaim_t8_dpo_7b_fkd0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).

CiteArk reconstructionclaim_fig8_asymmetric_supervision_stages1 plan1 run
Scientific conclusionInconclusive

Evaluate Teacher Supervision Asymmetry (ROC AUC) Across Training Stages

exp_fig8_asymmetric_supervision_stages2 attempts
Evidence published

Next step

The execution path is complete; inspect the scientific Assessment next.

View local evidence pathExpand
  1. Research plan

    Claim and experiment binding established

  2. Execution target

    Evidence published

  3. Run attempt

    Evidence publication · Completed

  4. CAP evidence

    1 immutable Artifact

  5. Scientific judgment

    Assessed but inconclusive

Legacy task without target-level resource requirements