Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
尚未评估计划受阻研究发现claim_fig7_smollm_tradeoff_generalizability

The stage-dependent distillation tradeoff generalizes to the SmolLM family (SmolLM2 360M student, 1.7B Instruct teacher, SmolLM3 Stage-3 data), where mid-training KD exhibits the reasoning-recall deficit and Switch Distillation mitigates it.

来源:source_paper:Figure 7, Page 26

报告指标与观测值

这条 Claim 暂无结构化指标。

Assessment(0)

这条 Claim 暂无已发布的不可变 Assessment。

关联运行(0)

这条 Claim 暂无关联执行。