Mid-training ablation results for Always CE relative difference across Reasoning, Factual Recall, and Knowledge task groups.
This claim has no executable experiment plan yet.
Start with coverage and measured results; open a run only when you need evidence or technical details.
0/76
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
1
Inconclusive
75
Not assessed
Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).
Reported
0.815 score
Observed
0.8079 score
Difference -0.0071
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
Failed paths grouped by cause — check them before reproducing.
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions.
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
Mid-training ablation results for Always CE relative difference across Reasoning, Factual Recall, and Knowledge task groups.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage RLVR1 for 13b_rkd.
This claim has no executable experiment plan yet.
The stage-dependent distillation tradeoff generalizes to the SmolLM family (SmolLM2 360M student, 1.7B Instruct teacher, SmolLM3 Stage-3 data), where mid-training KD exhibits the reasoning-recall deficit and Switch Distillation mitigates it.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage RLVR1 for 13b_fkd.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage SFT for 13b_rkd.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage SFT for 13b_sd.
This claim has no executable experiment plan yet.
Factual acquisition stratification by teacher entropy quintile persists when using OLMo-2 1B Instruct and 13B Instruct teachers.
This claim has no executable experiment plan yet.
Teacher entropy under OLMo-2 7B Instruct strongly predicts factual acquisition under NTP: by the end of pre-training, the student learns 67% of Q1 facts vs 5% of Q5 facts; by mid-training initialization (4T tokens), 80% of Q1 facts are learned vs 7% of Q5 facts.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage RLVR1 for ntp_shared.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage SFT for 7b_rkd.
This claim has no executable experiment plan yet.
Switch Distillation (SD, q=20%, tau=2) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.
This claim has no executable experiment plan yet.
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q30.
This claim has no executable experiment plan yet.
Per-task downstream accuracy for mid-training ablation random_routing using OLMo-2 7B Instruct teacher.
This claim has no executable experiment plan yet.
Per-task downstream accuracy for mid-training ablation teacher_top1 using OLMo-2 7B Instruct teacher.
This claim has no executable experiment plan yet.
Mid-training ablation results for Switch Distillation absolute macro-averages across Reasoning, Factual Recall, and Knowledge task groups.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage SFT for 7b_sd.
This claim has no executable experiment plan yet.
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q10.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage DPO for 7b_trkd.
This claim has no executable experiment plan yet.
Forward KL Distillation (FKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.
This claim has no executable experiment plan yet.
Forward KL consistently exhibits higher CE-KL gradient cosine alignment than Reverse KL across pre-training and mid-training, with the gap widening at higher alpha and later training steps.
This claim has no executable experiment plan yet.
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B TRKD post-training performance.
This claim has no executable experiment plan yet.
Teacher predictive entropy distinguishes procedural from knowledge-intensive domains across diverse instruction-tuned open-weight model families (OLMo-3 7B Instruct: 0.771, Qwen 3 8B: 0.705, Gemma-3 12B it: 0.707, Granite 3.3 8B Instruct: 0.696).
Evaluate Teacher Supervision Asymmetry (ROC AUC) Across Model Families
Next step
A person must review the failure and evidence integrity.
Claim and experiment binding established
Execution failed
Experiment execution · Running
No linked immutable Artifact yet
No Assessment yet
Legacy task without target-level resource requirements
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B FKD post-training performance.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage DPO for 13b_sd.
This claim has no executable experiment plan yet.
Mid-training ablation results for Oracle Domain Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.
This claim has no executable experiment plan yet.
Gold-token probability, gradient attenuation relative to NTP, and factual recall deficits under KD hold consistently for OLMo-2 1B and 13B Instruct teachers.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage RLVR1 for 7b_rkd.
This claim has no executable experiment plan yet.
Reverse KL Distillation (RKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage SFT for 7b_trkd.
This claim has no executable experiment plan yet.
Mid-training ablation results for Teacher Top-1 Labels relative difference across Reasoning, Factual Recall, and Knowledge task groups.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage SFT for 13b_fkd.
This claim has no executable experiment plan yet.
Per-task downstream accuracy for mid-training ablation oracle_domain using OLMo-2 7B Instruct teacher.
This claim has no executable experiment plan yet.
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B FKD post-training performance.
This claim has no executable experiment plan yet.
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q10.
This claim has no executable experiment plan yet.
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B TRKD post-training performance.
This claim has no executable experiment plan yet.
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B RKD post-training performance.
This claim has no executable experiment plan yet.
Switch Distillation substantially accelerates reasoning acquisition during mid-training, surpassing the final 60B-token NTP reasoning macro-average within 2.5B tokens (24x fewer tokens), while standard FKD requires 5.0B tokens (12x fewer).
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage RLVR1 for 7b_fkd.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage DPO for 13b_fkd.
This claim has no executable experiment plan yet.
Switch Distillation (SD, q=20%, tau=2) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher, demonstrating substantial gains on reasoning while maintaining factual recall.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage DPO for 13b_rkd.
This claim has no executable experiment plan yet.
Token-Routing Knowledge Distillation (TRKD) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.
This claim has no executable experiment plan yet.
Forward KL Distillation (FKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.
This claim has no executable experiment plan yet.
Standard Next-Token Prediction (NTP) baseline downstream evaluation performance after 60B mid-training tokens across 15 tasks covering Reasoning, Factual Recall, and Knowledge & Commonsense.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage SFT for ntp_shared.
This claim has no executable experiment plan yet.
Token-Routing Knowledge Distillation (TRKD) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.
This claim has no executable experiment plan yet.
Lower teacher predictive entropy corresponds to substantially higher teacher top-1 agreement with the ground-truth token across all data domains and teacher sizes.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage DPO for 7b_rkd.
This claim has no executable experiment plan yet.
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B Switch Distillation post-training performance.
This claim has no executable experiment plan yet.
Higher teacher entropy corresponds to strictly lower gold-token probability, attenuates the gold-token gradient relative to NTP (reaching ~0.5x NTP for Q5 facts), and produces larger downstream factual recall deficits.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage DPO for 7b_sd.
This claim has no executable experiment plan yet.
Across distillation strengths alpha in [0.0, 1.0], KL directions (forward and reverse), and teacher sizes (1B, 7B, 13B), mid-training distillation traces a reasoning-recall frontier that falls below NTP, which Switch Distillation mitigates.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage DPO for ntp_shared.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage SFT for 13b_trkd.
This claim has no executable experiment plan yet.
Knowledge distillation generally improves reasoning, but its effect on factual recall changes across training stages: during pre-training, distillation improves both reasoning and factual recall over NTP, whereas during mid-training it improves reasoning at the expense of factual recall.
This claim has no executable experiment plan yet.
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q30.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage RLVR1 for 7b_sd.
This claim has no executable experiment plan yet.
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q20.
This claim has no executable experiment plan yet.
Mid-training ablation results for Switch Distillation with FKL relative difference across Reasoning, Factual Recall, and Knowledge task groups.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage DPO for 13b_trkd.
This claim has no executable experiment plan yet.
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B Switch Distillation post-training performance.
This claim has no executable experiment plan yet.
Reverse KL Distillation (RKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage SFT for 7b_fkd.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage RLVR1 for 13b_sd.
This claim has no executable experiment plan yet.
Per-task downstream accuracy for mid-training ablation sd_fkl using OLMo-2 7B Instruct teacher.
This claim has no executable experiment plan yet.
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B RKD post-training performance.
This claim has no executable experiment plan yet.
Mid-training ablation results for Random Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage RLVR1 for 7b_trkd.
This claim has no executable experiment plan yet.
Per-task downstream accuracy for mid-training ablation teacher_correct using OLMo-2 7B Instruct teacher.
This claim has no executable experiment plan yet.
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q20.
This claim has no executable experiment plan yet.
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for NTP post-training baseline performance.
This claim has no executable experiment plan yet.
Mid-training ablation results for Teacher-Correct Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.
This claim has no executable experiment plan yet.
Per-task downstream accuracy for mid-training ablation always_ce using OLMo-2 7B Instruct teacher.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage RLVR1 for 13b_trkd.
This claim has no executable experiment plan yet.
Downstream task evaluation after post-training stage DPO for 7b_fkd.
This claim has no executable experiment plan yet.
Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).
Evaluate Teacher Supervision Asymmetry (ROC AUC) Across Training Stages
Next step
The execution path is complete; inspect the scientific Assessment next.
Claim and experiment binding established
Evidence published
Evidence publication · Completed
1 immutable Artifact
Assessed but inconclusive
Legacy task without target-level resource requirements