Mid-training ablation results for Always CE relative difference across Reasoning, Factual Recall, and Knowledge task groups.
这条结论尚无可执行实验方案。
优先展示覆盖情况与实测结果;需要审计时再打开运行记录和技术细节。
0/76
条结论获得证据支持
0
获得支持
0
受到挑战或冲突
0
遭到反驳
1
无法判定
75
尚未评估
Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).
论文报告值
0.815 score
CiteArk 实测值
0.8079 score
差异 -0.0071
每一行是一条真实执行记录;命令、日志、哈希和签名都收纳在详情中。
按原因聚类的失败路径,复现前先看看别人踩过的坑。
逐个实验展示当前执行状态、阻塞原因、恢复动作和证据去向;技术执行与科研结论始终分开。
执行成功不等于论文结论成立
目标状态回答平台有没有跑完;右侧科学结论只由不可变证据和 Assessment 决定。资源不足或平台故障不会被写成反驳论文的科研结论。
Mid-training ablation results for Always CE relative difference across Reasoning, Factual Recall, and Knowledge task groups.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage RLVR1 for 13b_rkd.
这条结论尚无可执行实验方案。
The stage-dependent distillation tradeoff generalizes to the SmolLM family (SmolLM2 360M student, 1.7B Instruct teacher, SmolLM3 Stage-3 data), where mid-training KD exhibits the reasoning-recall deficit and Switch Distillation mitigates it.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage RLVR1 for 13b_fkd.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage SFT for 13b_rkd.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage SFT for 13b_sd.
这条结论尚无可执行实验方案。
Factual acquisition stratification by teacher entropy quintile persists when using OLMo-2 1B Instruct and 13B Instruct teachers.
这条结论尚无可执行实验方案。
Teacher entropy under OLMo-2 7B Instruct strongly predicts factual acquisition under NTP: by the end of pre-training, the student learns 67% of Q1 facts vs 5% of Q5 facts; by mid-training initialization (4T tokens), 80% of Q1 facts are learned vs 7% of Q5 facts.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage RLVR1 for ntp_shared.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage SFT for 7b_rkd.
这条结论尚无可执行实验方案。
Switch Distillation (SD, q=20%, tau=2) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.
这条结论尚无可执行实验方案。
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q30.
这条结论尚无可执行实验方案。
Per-task downstream accuracy for mid-training ablation random_routing using OLMo-2 7B Instruct teacher.
这条结论尚无可执行实验方案。
Per-task downstream accuracy for mid-training ablation teacher_top1 using OLMo-2 7B Instruct teacher.
这条结论尚无可执行实验方案。
Mid-training ablation results for Switch Distillation absolute macro-averages across Reasoning, Factual Recall, and Knowledge task groups.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage SFT for 7b_sd.
这条结论尚无可执行实验方案。
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q10.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage DPO for 7b_trkd.
这条结论尚无可执行实验方案。
Forward KL Distillation (FKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.
这条结论尚无可执行实验方案。
Forward KL consistently exhibits higher CE-KL gradient cosine alignment than Reverse KL across pre-training and mid-training, with the gap widening at higher alpha and later training steps.
这条结论尚无可执行实验方案。
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B TRKD post-training performance.
这条结论尚无可执行实验方案。
Teacher predictive entropy distinguishes procedural from knowledge-intensive domains across diverse instruction-tuned open-weight model families (OLMo-3 7B Instruct: 0.771, Qwen 3 8B: 0.705, Gemma-3 12B it: 0.707, Granite 3.3 8B Instruct: 0.696).
Evaluate Teacher Supervision Asymmetry (ROC AUC) Across Model Families
下一步
需要人工核查失败原因和证据完整性。
结论与实验绑定已建立
执行失败
实验执行 · 进行中
尚无已关联的不可变 Artifact
尚无 Assessment
历史任务未记录目标级资源要求
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B FKD post-training performance.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage DPO for 13b_sd.
这条结论尚无可执行实验方案。
Mid-training ablation results for Oracle Domain Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.
这条结论尚无可执行实验方案。
Gold-token probability, gradient attenuation relative to NTP, and factual recall deficits under KD hold consistently for OLMo-2 1B and 13B Instruct teachers.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage RLVR1 for 7b_rkd.
这条结论尚无可执行实验方案。
Reverse KL Distillation (RKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage SFT for 7b_trkd.
这条结论尚无可执行实验方案。
Mid-training ablation results for Teacher Top-1 Labels relative difference across Reasoning, Factual Recall, and Knowledge task groups.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage SFT for 13b_fkd.
这条结论尚无可执行实验方案。
Per-task downstream accuracy for mid-training ablation oracle_domain using OLMo-2 7B Instruct teacher.
这条结论尚无可执行实验方案。
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B FKD post-training performance.
这条结论尚无可执行实验方案。
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q10.
这条结论尚无可执行实验方案。
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B TRKD post-training performance.
这条结论尚无可执行实验方案。
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B RKD post-training performance.
这条结论尚无可执行实验方案。
Switch Distillation substantially accelerates reasoning acquisition during mid-training, surpassing the final 60B-token NTP reasoning macro-average within 2.5B tokens (24x fewer tokens), while standard FKD requires 5.0B tokens (12x fewer).
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage RLVR1 for 7b_fkd.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage DPO for 13b_fkd.
这条结论尚无可执行实验方案。
Switch Distillation (SD, q=20%, tau=2) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher, demonstrating substantial gains on reasoning while maintaining factual recall.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage DPO for 13b_rkd.
这条结论尚无可执行实验方案。
Token-Routing Knowledge Distillation (TRKD) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.
这条结论尚无可执行实验方案。
Forward KL Distillation (FKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.
这条结论尚无可执行实验方案。
Standard Next-Token Prediction (NTP) baseline downstream evaluation performance after 60B mid-training tokens across 15 tasks covering Reasoning, Factual Recall, and Knowledge & Commonsense.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage SFT for ntp_shared.
这条结论尚无可执行实验方案。
Token-Routing Knowledge Distillation (TRKD) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.
这条结论尚无可执行实验方案。
Lower teacher predictive entropy corresponds to substantially higher teacher top-1 agreement with the ground-truth token across all data domains and teacher sizes.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage DPO for 7b_rkd.
这条结论尚无可执行实验方案。
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B Switch Distillation post-training performance.
这条结论尚无可执行实验方案。
Higher teacher entropy corresponds to strictly lower gold-token probability, attenuates the gold-token gradient relative to NTP (reaching ~0.5x NTP for Q5 facts), and produces larger downstream factual recall deficits.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage DPO for 7b_sd.
这条结论尚无可执行实验方案。
Across distillation strengths alpha in [0.0, 1.0], KL directions (forward and reverse), and teacher sizes (1B, 7B, 13B), mid-training distillation traces a reasoning-recall frontier that falls below NTP, which Switch Distillation mitigates.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage DPO for ntp_shared.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage SFT for 13b_trkd.
这条结论尚无可执行实验方案。
Knowledge distillation generally improves reasoning, but its effect on factual recall changes across training stages: during pre-training, distillation improves both reasoning and factual recall over NTP, whereas during mid-training it improves reasoning at the expense of factual recall.
这条结论尚无可执行实验方案。
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q30.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage RLVR1 for 7b_sd.
这条结论尚无可执行实验方案。
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q20.
这条结论尚无可执行实验方案。
Mid-training ablation results for Switch Distillation with FKL relative difference across Reasoning, Factual Recall, and Knowledge task groups.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage DPO for 13b_trkd.
这条结论尚无可执行实验方案。
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B Switch Distillation post-training performance.
这条结论尚无可执行实验方案。
Reverse KL Distillation (RKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage SFT for 7b_fkd.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage RLVR1 for 13b_sd.
这条结论尚无可执行实验方案。
Per-task downstream accuracy for mid-training ablation sd_fkl using OLMo-2 7B Instruct teacher.
这条结论尚无可执行实验方案。
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B RKD post-training performance.
这条结论尚无可执行实验方案。
Mid-training ablation results for Random Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage RLVR1 for 7b_trkd.
这条结论尚无可执行实验方案。
Per-task downstream accuracy for mid-training ablation teacher_correct using OLMo-2 7B Instruct teacher.
这条结论尚无可执行实验方案。
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q20.
这条结论尚无可执行实验方案。
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for NTP post-training baseline performance.
这条结论尚无可执行实验方案。
Mid-training ablation results for Teacher-Correct Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.
这条结论尚无可执行实验方案。
Per-task downstream accuracy for mid-training ablation always_ce using OLMo-2 7B Instruct teacher.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage RLVR1 for 13b_trkd.
这条结论尚无可执行实验方案。
Downstream task evaluation after post-training stage DPO for 7b_fkd.
这条结论尚无可执行实验方案。
Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).
Evaluate Teacher Supervision Asymmetry (ROC AUC) Across Training Stages
下一步
执行链路已完成;继续查看科学 Assessment。
结论与实验绑定已建立
证据已发布
证据发布 · 已完成
1 个不可变 Artifact
已有评估但无法判定
历史任务未记录目标级资源要求