Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
公开更多
研究结论
0 / 76 条结论已通过验证
其余仍在验证中
关键指标论文报告值 → 复现观测值 · 点击行展开详情
Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).证据不足4/12 项指标获得支持 · 12 项已有评估报告 0.815 score0.8079 score
报告
0.815 score
观测
0.8079 score
偏差
−0.0071 score
此次执行没有产生足够的可比较测量结果,因此现有证据不足以判断论文结论。 · 这不构成对论文结论的反驳,只表示目前的证据既不能确认,也不能否定它。
roc_auc · 获得支持
报告: 0.815 score
观测: 0.8079 score
Reported ROC AUC of 0.815 was observed at 0.8079 (relative diff 0.87%), showing strong predictive entropy separation on OLMo-2 1B Base.
roc_auc · 获得支持
报告: 0.77 score
观测: 0.7967 score
Reported ROC AUC of 0.770 was observed at 0.7967 (relative diff 3.47%), confirming high ROC AUC separation on OLMo-2 1B SFT.
roc_auc · 获得支持
报告: 0.744 score
观测: 0.7931 score
Reported ROC AUC of 0.744 was observed at 0.7931 (relative diff 6.60%), maintaining the substantive finding of strong discrimination on OLMo-2 1B DPO.
roc_auc · 获得支持
报告: 0.748 score
观测: 0.7973 score
Reported ROC AUC of 0.748 was observed at 0.7973 (relative diff 6.59%), confirming procedural entropy remains distinct on OLMo-2 1B Instruct.
roc_auc · 证据不足
报告: 0.816 score
观测: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json
roc_auc · 证据不足
报告: 0.777 score
观测: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json
roc_auc · 证据不足
报告: 0.76 score
观测: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json
roc_auc · 证据不足
报告: 0.761 score
观测: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json
roc_auc · 证据不足
报告: 0.826 score
观测: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json
roc_auc · 证据不足
报告: 0.814 score
观测: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json
roc_auc · 证据不足
报告: 0.809 score
观测: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json
roc_auc · 证据不足
报告: 0.811 score
观测: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json
Mid-training ablation results for Always CE relative difference across Reasoning, Factual Recall, and Knowledge task groups.实验受阻,详见具体原因报告 -0.3 个百分点
报告
-0.3 个百分点
观测
—
Ablation checkpoints not released and training exceeds compute limits.
delta_percentage_points · 尚未评估
报告: -0.3 个百分点
delta_percentage_points · 尚未评估
报告: 0.1 个百分点
delta_percentage_points · 尚未评估
报告: -0.6 个百分点
Downstream task evaluation after post-training stage RLVR1 for 13b_rkd.实验受阻,详见具体原因报告 75.6 个百分点
报告
75.6 个百分点
观测
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · 尚未评估
报告: 75.6 个百分点
downstream_accuracy · 尚未评估
报告: 60.1 个百分点
downstream_accuracy · 尚未评估
报告: 50.3 个百分点
downstream_accuracy · 尚未评估
报告: 33.3 个百分点
downstream_accuracy · 尚未评估
报告: 38.1 个百分点
downstream_accuracy · 尚未评估
报告: 15.8 个百分点
downstream_accuracy · 尚未评估
报告: 53.3 个百分点
downstream_accuracy · 尚未评估
报告: 22.3 个百分点
downstream_accuracy · 尚未评估
报告: 8 个百分点
downstream_accuracy · 尚未评估
报告: 36.2 个百分点
downstream_accuracy · 尚未评估
报告: 16.9 个百分点
downstream_accuracy · 尚未评估
报告: 55.9 个百分点
downstream_accuracy · 尚未评估
报告: 55.2 个百分点
downstream_accuracy · 尚未评估
报告: 51.3 个百分点
downstream_accuracy · 尚未评估
报告: 36.4 个百分点
downstream_accuracy · 尚未评估
报告: 63.6 个百分点
Downstream task evaluation after post-training stage RLVR1 for 13b_fkd.实验受阻,详见具体原因报告 73 个百分点
报告
73 个百分点
观测
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · 尚未评估
报告: 73 个百分点
downstream_accuracy · 尚未评估
报告: 55.5 个百分点
downstream_accuracy · 尚未评估
报告: 47.3 个百分点
downstream_accuracy · 尚未评估
报告: 33.7 个百分点
downstream_accuracy · 尚未评估
报告: 36.1 个百分点
downstream_accuracy · 尚未评估
报告: 14 个百分点
downstream_accuracy · 尚未评估
报告: 54.7 个百分点
downstream_accuracy · 尚未评估
报告: 23.3 个百分点
downstream_accuracy · 尚未评估
报告: 8.3 个百分点
downstream_accuracy · 尚未评估
报告: 35 个百分点
downstream_accuracy · 尚未评估
报告: 17.2 个百分点
downstream_accuracy · 尚未评估
报告: 54 个百分点
downstream_accuracy · 尚未评估
报告: 56.6 个百分点
downstream_accuracy · 尚未评估
报告: 51.7 个百分点
downstream_accuracy · 尚未评估
报告: 35.4 个百分点
downstream_accuracy · 尚未评估
报告: 63.2 个百分点
Downstream task evaluation after post-training stage SFT for 13b_rkd.实验受阻,详见具体原因报告 56.8 个百分点
报告
56.8 个百分点
观测
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · 尚未评估
报告: 56.8 个百分点
downstream_accuracy · 尚未评估
报告: 36.7 个百分点
downstream_accuracy · 尚未评估
报告: 31.3 个百分点
downstream_accuracy · 尚未评估
报告: 32.1 个百分点
downstream_accuracy · 尚未评估
报告: 38.9 个百分点
downstream_accuracy · 尚未评估
报告: 11.6 个百分点
downstream_accuracy · 尚未评估
报告: 55.4 个百分点
downstream_accuracy · 尚未评估
报告: 23.8 个百分点
downstream_accuracy · 尚未评估
报告: 8 个百分点
downstream_accuracy · 尚未评估
报告: 47.3 个百分点
downstream_accuracy · 尚未评估
报告: 17.3 个百分点
downstream_accuracy · 尚未评估
报告: 57.8 个百分点
downstream_accuracy · 尚未评估
报告: 57.4 个百分点
downstream_accuracy · 尚未评估
报告: 52.2 个百分点
downstream_accuracy · 尚未评估
报告: 38.1 个百分点
downstream_accuracy · 尚未评估
报告: 43.6 个百分点
Downstream task evaluation after post-training stage SFT for 13b_sd.实验受阻,详见具体原因报告 62.7 个百分点
报告
62.7 个百分点
观测
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · 尚未评估
报告: 62.7 个百分点
downstream_accuracy · 尚未评估
报告: 42.5 个百分点
downstream_accuracy · 尚未评估
报告: 34.8 个百分点
downstream_accuracy · 尚未评估
报告: 31.6 个百分点
downstream_accuracy · 尚未评估
报告: 43.5 个百分点
downstream_accuracy · 尚未评估
报告: 10.2 个百分点
downstream_accuracy · 尚未评估
报告: 54.9 个百分点
downstream_accuracy · 尚未评估
报告: 24.1 个百分点
downstream_accuracy · 尚未评估
报告: 8.5 个百分点
downstream_accuracy · 尚未评估
报告: 48 个百分点
downstream_accuracy · 尚未评估
报告: 17.7 个百分点
downstream_accuracy · 尚未评估
报告: 55.1 个百分点
downstream_accuracy · 尚未评估
报告: 59.4 个百分点
downstream_accuracy · 尚未评估
报告: 52.2 个百分点
downstream_accuracy · 尚未评估
报告: 38 个百分点
downstream_accuracy · 尚未评估
报告: 43.1 个百分点
Factual acquisition stratification by teacher entropy quintile persists when using OLMo-2 1B Instruct and 13B Instruct teachers.实验受阻,详见具体原因报告 76 个百分点
报告
76 个百分点
观测
—
Intermediate pre-training and mid-training NTP trajectory checkpoints are not public.
factual_examples_learned · 尚未评估
报告: 76 个百分点
factual_examples_learned · 尚未评估
报告: 48 个百分点
factual_examples_learned · 尚未评估
报告: 30 个百分点
factual_examples_learned · 尚未评估
报告: 12 个百分点
factual_examples_learned · 尚未评估
报告: 6 个百分点
factual_examples_learned · 尚未评估
报告: 81 个百分点
factual_examples_learned · 尚未评估
报告: 61 个百分点
factual_examples_learned · 尚未评估
报告: 40 个百分点
factual_examples_learned · 尚未评估
报告: 18 个百分点
factual_examples_learned · 尚未评估
报告: 8 个百分点
factual_examples_learned · 尚未评估
报告: 54 个百分点
factual_examples_learned · 尚未评估
报告: 48 个百分点
factual_examples_learned · 尚未评估
报告: 37 个百分点
factual_examples_learned · 尚未评估
报告: 28 个百分点
factual_examples_learned · 尚未评估
报告: 6 个百分点
factual_examples_learned · 尚未评估
报告: 66 个百分点
factual_examples_learned · 尚未评估
报告: 58 个百分点
factual_examples_learned · 尚未评估
报告: 43 个百分点
factual_examples_learned · 尚未评估
报告: 32 个百分点
factual_examples_learned · 尚未评估
报告: 8 个百分点
Teacher entropy under OLMo-2 7B Instruct strongly predicts factual acquisition under NTP: by the end of pre-training, the student learns 67% of Q1 facts vs 5% of Q5 facts; by mid-training initialization (4T tokens), 80% of Q1 facts are learned vs 7% of Q5 facts.实验受阻,详见具体原因报告 67 个百分点
报告
67 个百分点
观测
—
The intermediate NTP training checkpoints across pre-training and mid-training trajectories are not released.
factual_examples_learned · 尚未评估
报告: 67 个百分点
factual_examples_learned · 尚未评估
报告: 47 个百分点
factual_examples_learned · 尚未评估
报告: 34 个百分点
factual_examples_learned · 尚未评估
报告: 21 个百分点
factual_examples_learned · 尚未评估
报告: 5 个百分点
factual_examples_learned · 尚未评估
报告: 80 个百分点
factual_examples_learned · 尚未评估
报告: 55 个百分点
factual_examples_learned · 尚未评估
报告: 42 个百分点
factual_examples_learned · 尚未评估
报告: 24 个百分点
factual_examples_learned · 尚未评估
报告: 7 个百分点
Downstream task evaluation after post-training stage SFT for 7b_rkd.实验受阻,详见具体原因报告 57.7 个百分点
报告
57.7 个百分点
观测
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · 尚未评估
报告: 57.7 个百分点
downstream_accuracy · 尚未评估
报告: 35.5 个百分点
downstream_accuracy · 尚未评估
报告: 32 个百分点
downstream_accuracy · 尚未评估
报告: 32.1 个百分点
downstream_accuracy · 尚未评估
报告: 39.9 个百分点
downstream_accuracy · 尚未评估
报告: 11.2 个百分点
downstream_accuracy · 尚未评估
报告: 53.2 个百分点
downstream_accuracy · 尚未评估
报告: 22.4 个百分点
downstream_accuracy · 尚未评估
报告: 8.4 个百分点
downstream_accuracy · 尚未评估
报告: 48.9 个百分点
downstream_accuracy · 尚未评估
报告: 18.2 个百分点
downstream_accuracy · 尚未评估
报告: 60.9 个百分点
downstream_accuracy · 尚未评估
报告: 60.2 个百分点
downstream_accuracy · 尚未评估
报告: 52.2 个百分点
downstream_accuracy · 尚未评估
报告: 39.2 个百分点
downstream_accuracy · 尚未评估
报告: 45.8 个百分点
Switch Distillation (SD, q=20%, tau=2) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.实验受阻,详见具体原因报告 66 个百分点
报告
66 个百分点
观测
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · 尚未评估
报告: 66 个百分点
downstream_accuracy · 尚未评估
报告: 54.8 个百分点
downstream_accuracy · 尚未评估
报告: 42.2 个百分点
downstream_accuracy · 尚未评估
报告: 31.6 个百分点
downstream_accuracy · 尚未评估
报告: 45.8 个百分点
downstream_accuracy · 尚未评估
报告: 12 个百分点
downstream_accuracy · 尚未评估
报告: 54.1 个百分点
downstream_accuracy · 尚未评估
报告: 24.9 个百分点
downstream_accuracy · 尚未评估
报告: 9 个百分点
downstream_accuracy · 尚未评估
报告: 48.5 个百分点
downstream_accuracy · 尚未评估
报告: 17.7 个百分点
downstream_accuracy · 尚未评估
报告: 58.5 个百分点
downstream_accuracy · 尚未评估
报告: 63.8 个百分点
downstream_accuracy · 尚未评估
报告: 51.2 个百分点
downstream_accuracy · 尚未评估
报告: 39.2 个百分点
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q30.实验受阻,详见具体原因报告 69.8 个百分点
报告
69.8 个百分点
观测
—
Threshold sweep checkpoints not released.
downstream_accuracy · 尚未评估
报告: 69.8 个百分点
downstream_accuracy · 尚未评估
报告: 54.1 个百分点
downstream_accuracy · 尚未评估
报告: 45.6 个百分点
downstream_accuracy · 尚未评估
报告: 33.8 个百分点
downstream_accuracy · 尚未评估
报告: 47.2 个百分点
downstream_accuracy · 尚未评估
报告: 13 个百分点
downstream_accuracy · 尚未评估
报告: 43.9 个百分点
downstream_accuracy · 尚未评估
报告: 54.7 个百分点
downstream_accuracy · 尚未评估
报告: 25.3 个百分点
downstream_accuracy · 尚未评估
报告: 8.8 个百分点
downstream_accuracy · 尚未评估
报告: 29.6 个百分点
downstream_accuracy · 尚未评估
报告: 51 个百分点
downstream_accuracy · 尚未评估
报告: 19.5 个百分点
downstream_accuracy · 尚未评估
报告: 63.5 个百分点
downstream_accuracy · 尚未评估
报告: 62.4 个百分点
downstream_accuracy · 尚未评估
报告: 53 个百分点
downstream_accuracy · 尚未评估
报告: 40.7 个百分点
downstream_accuracy · 尚未评估
报告: 48.4 个百分点
Per-task downstream accuracy for mid-training ablation random_routing using OLMo-2 7B Instruct teacher.实验受阻,详见具体原因报告 61.5 个百分点
报告
61.5 个百分点
观测
—
Per-task ablation checkpoints not released.
downstream_accuracy · 尚未评估
报告: 61.5 个百分点
downstream_accuracy · 尚未评估
报告: 45.5 个百分点
downstream_accuracy · 尚未评估
报告: 38.6 个百分点
downstream_accuracy · 尚未评估
报告: 32 个百分点
downstream_accuracy · 尚未评估
报告: 40.6 个百分点
downstream_accuracy · 尚未评估
报告: 10.8 个百分点
downstream_accuracy · 尚未评估
报告: 53.2 个百分点
downstream_accuracy · 尚未评估
报告: 23.9 个百分点
downstream_accuracy · 尚未评估
报告: 8.5 个百分点
downstream_accuracy · 尚未评估
报告: 49.6 个百分点
downstream_accuracy · 尚未评估
报告: 17.8 个百分点
downstream_accuracy · 尚未评估
报告: 61.6 个百分点
downstream_accuracy · 尚未评估
报告: 62.6 个百分点
downstream_accuracy · 尚未评估
报告: 52.6 个百分点
downstream_accuracy · 尚未评估
报告: 39.5 个百分点
Per-task downstream accuracy for mid-training ablation teacher_top1 using OLMo-2 7B Instruct teacher.实验受阻,详见具体原因报告 62.2 个百分点
报告
62.2 个百分点
观测
—
Per-task ablation checkpoints not released.
downstream_accuracy · 尚未评估
报告: 62.2 个百分点
downstream_accuracy · 尚未评估
报告: 46 个百分点
downstream_accuracy · 尚未评估
报告: 40.1 个百分点
downstream_accuracy · 尚未评估
报告: 30.7 个百分点
downstream_accuracy · 尚未评估
报告: 42.1 个百分点
downstream_accuracy · 尚未评估
报告: 8.8 个百分点
downstream_accuracy · 尚未评估
报告: 57.9 个百分点
downstream_accuracy · 尚未评估
报告: 24.7 个百分点
downstream_accuracy · 尚未评估
报告: 9.3 个百分点
downstream_accuracy · 尚未评估
报告: 48.8 个百分点
downstream_accuracy · 尚未评估
报告: 18 个百分点
downstream_accuracy · 尚未评估
报告: 59.8 个百分点
downstream_accuracy · 尚未评估
报告: 62 个百分点
downstream_accuracy · 尚未评估
报告: 51 个百分点
downstream_accuracy · 尚未评估
报告: 39.2 个百分点
Mid-training ablation results for Switch Distillation absolute macro-averages across Reasoning, Factual Recall, and Knowledge task groups.实验受阻,详见具体原因报告 44.7 个百分点
报告
44.7 个百分点
观测
—
Ablation checkpoints not released and training exceeds compute limits.
macro_average · 尚未评估
报告: 44.7 个百分点
macro_average · 尚未评估
报告: 29.3 个百分点
macro_average · 尚未评估
报告: 49.3 个百分点
Downstream task evaluation after post-training stage SFT for 7b_sd.实验受阻,详见具体原因报告 63.7 个百分点
报告
63.7 个百分点
观测
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · 尚未评估
报告: 63.7 个百分点
downstream_accuracy · 尚未评估
报告: 42.7 个百分点
downstream_accuracy · 尚未评估
报告: 36.7 个百分点
downstream_accuracy · 尚未评估
报告: 33.5 个百分点
downstream_accuracy · 尚未评估
报告: 48.9 个百分点
downstream_accuracy · 尚未评估
报告: 12.8 个百分点
downstream_accuracy · 尚未评估
报告: 55.1 个百分点
downstream_accuracy · 尚未评估
报告: 24.5 个百分点
downstream_accuracy · 尚未评估
报告: 8.5 个百分点
downstream_accuracy · 尚未评估
报告: 50.3 个百分点
downstream_accuracy · 尚未评估
报告: 19.4 个百分点
downstream_accuracy · 尚未评估
报告: 61.3 个百分点
downstream_accuracy · 尚未评估
报告: 61.4 个百分点
downstream_accuracy · 尚未评估
报告: 56.1 个百分点
downstream_accuracy · 尚未评估
报告: 40.6 个百分点
downstream_accuracy · 尚未评估
报告: 45.5 个百分点
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q10.实验受阻,详见具体原因报告 62.1 个百分点
报告
62.1 个百分点
观测
—
Threshold sweep checkpoints not released.
downstream_accuracy · 尚未评估
报告: 62.1 个百分点
downstream_accuracy · 尚未评估
报告: 49.6 个百分点
downstream_accuracy · 尚未评估
报告: 38.3 个百分点
downstream_accuracy · 尚未评估
报告: 28.6 个百分点
downstream_accuracy · 尚未评估
报告: 44.2 个百分点
downstream_accuracy · 尚未评估
报告: 7.8 个百分点
downstream_accuracy · 尚未评估
报告: 38.5 个百分点
downstream_accuracy · 尚未评估
报告: 52.2 个百分点
downstream_accuracy · 尚未评估
报告: 23.1 个百分点
downstream_accuracy · 尚未评估
报告: 9.2 个百分点
downstream_accuracy · 尚未评估
报告: 28.1 个百分点
downstream_accuracy · 尚未评估
报告: 46.8 个百分点
downstream_accuracy · 尚未评估
报告: 17.1 个百分点
downstream_accuracy · 尚未评估
报告: 55.5 个百分点
downstream_accuracy · 尚未评估
报告: 59.6 个百分点
downstream_accuracy · 尚未评估
报告: 52.7 个百分点
downstream_accuracy · 尚未评估
报告: 37.9 个百分点
downstream_accuracy · 尚未评估
报告: 45 个百分点
Downstream task evaluation after post-training stage DPO for 7b_trkd.实验受阻,详见具体原因报告 57.2 个百分点
报告
57.2 个百分点
观测
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · 尚未评估
报告: 57.2 个百分点
downstream_accuracy · 尚未评估
报告: 34.9 个百分点
downstream_accuracy · 尚未评估
报告: 32.3 个百分点
downstream_accuracy · 尚未评估
报告: 32.4 个百分点
downstream_accuracy · 尚未评估
报告: 36.6 个百分点
downstream_accuracy · 尚未评估
报告: 7.2 个百分点
downstream_accuracy · 尚未评估
报告: 53.1 个百分点
downstream_accuracy · 尚未评估
报告: 23.2 个百分点
downstream_accuracy · 尚未评估
报告: 7.9 个百分点
downstream_accuracy · 尚未评估
报告: 47 个百分点
downstream_accuracy · 尚未评估
报告: 17.9 个百分点
downstream_accuracy · 尚未评估
报告: 56.7 个百分点
downstream_accuracy · 尚未评估
报告: 57.2 个百分点
downstream_accuracy · 尚未评估
报告: 51.9 个百分点
downstream_accuracy · 尚未评估
报告: 37.5 个百分点
downstream_accuracy · 尚未评估
报告: 62.5 个百分点
Forward KL Distillation (FKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.实验受阻,详见具体原因报告 57.8 个百分点
报告
57.8 个百分点
观测
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · 尚未评估
报告: 57.8 个百分点
downstream_accuracy · 尚未评估
报告: 42.6 个百分点
downstream_accuracy · 尚未评估
报告: 36.2 个百分点
downstream_accuracy · 尚未评估
报告: 30.6 个百分点
downstream_accuracy · 尚未评估
报告: 41.9 个百分点
downstream_accuracy · 尚未评估
报告: 7.6 个百分点
downstream_accuracy · 尚未评估
报告: 53.9 个百分点
downstream_accuracy · 尚未评估
报告: 24.6 个百分点
downstream_accuracy · 尚未评估
报告: 8.4 个百分点
downstream_accuracy · 尚未评估
报告: 49.5 个百分点
downstream_accuracy · 尚未评估
报告: 18.7 个百分点
downstream_accuracy · 尚未评估
报告: 61.9 个百分点
downstream_accuracy · 尚未评估
报告: 61.8 个百分点
downstream_accuracy · 尚未评估
报告: 51.5 个百分点
downstream_accuracy · 尚未评估
报告: 38.6 个百分点
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B TRKD post-training performance.实验受阻,详见具体原因报告 70.7 个百分点
报告
70.7 个百分点
观测
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · 尚未评估
报告: 70.7 个百分点
downstream_accuracy · 尚未评估
报告: 52.9 个百分点
downstream_accuracy · 尚未评估
报告: 46.3 个百分点
downstream_accuracy · 尚未评估
报告: 31.8 个百分点
downstream_accuracy · 尚未评估
报告: 36.8 个百分点
downstream_accuracy · 尚未评估
报告: 14.4 个百分点
downstream_accuracy · 尚未评估
报告: 51.7 个百分点
downstream_accuracy · 尚未评估
报告: 22.3 个百分点
downstream_accuracy · 尚未评估
报告: 8 个百分点
downstream_accuracy · 尚未评估
报告: 40.9 个百分点
downstream_accuracy · 尚未评估
报告: 17.4 个百分点
downstream_accuracy · 尚未评估
报告: 59 个百分点
downstream_accuracy · 尚未评估
报告: 53.6 个百分点
downstream_accuracy · 尚未评估
报告: 51.5 个百分点
downstream_accuracy · 尚未评估
报告: 35.6 个百分点
downstream_accuracy · 尚未评估
报告: 65.2 个百分点
Teacher predictive entropy distinguishes procedural from knowledge-intensive domains across diverse instruction-tuned open-weight model families (OLMo-3 7B Instruct: 0.771, Qwen 3 8B: 0.705, Gemma-3 12B it: 0.707, Granite 3.3 8B Instruct: 0.696).等待复现报告 0.771 score
报告
0.771 score
观测
—
实验方案已生成,还没有运行记录
roc_auc · 尚未评估
报告: 0.771 score
roc_auc · 尚未评估
报告: 0.705 score
roc_auc · 尚未评估
报告: 0.707 score
roc_auc · 尚未评估
报告: 0.696 score
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B FKD post-training performance.实验受阻,详见具体原因报告 76.2 个百分点
报告
76.2 个百分点
观测
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · 尚未评估
报告: 76.2 个百分点
downstream_accuracy · 尚未评估
报告: 56.4 个百分点
downstream_accuracy · 尚未评估
报告: 51.6 个百分点
downstream_accuracy · 尚未评估
报告: 33.8 个百分点
downstream_accuracy · 尚未评估
报告: 38.4 个百分点
downstream_accuracy · 尚未评估
报告: 18 个百分点
downstream_accuracy · 尚未评估
报告: 51.9 个百分点
downstream_accuracy · 尚未评估
报告: 22.2 个百分点
downstream_accuracy · 尚未评估
报告: 8.1 个百分点
downstream_accuracy · 尚未评估
报告: 42.2 个百分点
downstream_accuracy · 尚未评估
报告: 18.6 个百分点
downstream_accuracy · 尚未评估
报告: 60.2 个百分点
downstream_accuracy · 尚未评估
报告: 57.4 个百分点
downstream_accuracy · 尚未评估
报告: 51.5 个百分点
downstream_accuracy · 尚未评估
报告: 38 个百分点
downstream_accuracy · 尚未评估
报告: 64.7 个百分点
Downstream task evaluation after post-training stage DPO for 13b_sd.实验受阻,详见具体原因报告 69.9 个百分点
报告
69.9 个百分点
观测
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · 尚未评估
报告: 69.9 个百分点
downstream_accuracy · 尚未评估
报告: 45.7 个百分点
downstream_accuracy · 尚未评估
报告: 39.5 个百分点
downstream_accuracy · 尚未评估
报告: 32.2 个百分点
downstream_accuracy · 尚未评估
报告: 44.2 个百分点
downstream_accuracy · 尚未评估
报告: 10 个百分点
downstream_accuracy · 尚未评估
报告: 54.5 个百分点
downstream_accuracy · 尚未评估
报告: 23.9 个百分点
downstream_accuracy · 尚未评估
报告: 8.2 个百分点
downstream_accuracy · 尚未评估
报告: 47.3 个百分点
downstream_accuracy · 尚未评估
报告: 17.9 个百分点
downstream_accuracy · 尚未评估
报告: 56.1 个百分点
downstream_accuracy · 尚未评估
报告: 59.2 个百分点
downstream_accuracy · 尚未评估
报告: 52.5 个百分点
downstream_accuracy · 尚未评估
报告: 38.7 个百分点
downstream_accuracy · 尚未评估
报告: 61 个百分点
Mid-training ablation results for Oracle Domain Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.实验受阻,详见具体原因报告 -7.2 个百分点
报告
-7.2 个百分点
观测
—
Ablation checkpoints not released and training exceeds compute limits.
delta_percentage_points · 尚未评估
报告: -7.2 个百分点
delta_percentage_points · 尚未评估
报告: -1.3 个百分点
delta_percentage_points · 尚未评估
报告: -2.3 个百分点
Downstream task evaluation after post-training stage RLVR1 for 7b_rkd.实验受阻,详见具体原因报告 77.9 个百分点
报告
77.9 个百分点
观测
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · 尚未评估
报告: 77.9 个百分点
downstream_accuracy · 尚未评估
报告: 54.3 个百分点
downstream_accuracy · 尚未评估
报告: 52.5 个百分点
downstream_accuracy · 尚未评估
报告: 34.2 个百分点
downstream_accuracy · 尚未评估
报告: 40.1 个百分点
downstream_accuracy · 尚未评估
报告: 15.6 个百分点
downstream_accuracy · 尚未评估
报告: 52.1 个百分点
downstream_accuracy · 尚未评估
报告: 21.9 个百分点
downstream_accuracy · 尚未评估
报告: 7.8 个百分点
downstream_accuracy · 尚未评估
报告: 45.3 个百分点
downstream_accuracy · 尚未评估
报告: 16.8 个百分点
downstream_accuracy · 尚未评估
报告: 57.6 个百分点
downstream_accuracy · 尚未评估
报告: 50.4 个百分点
downstream_accuracy · 尚未评估
报告: 51.5 个百分点
downstream_accuracy · 尚未评估
报告: 36.1 个百分点
downstream_accuracy · 尚未评估
报告: 59.3 个百分点
Reverse KL Distillation (RKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.实验受阻,详见具体原因报告 59.2 个百分点
报告
59.2 个百分点
观测
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · 尚未评估
报告: 59.2 个百分点
downstream_accuracy · 尚未评估
报告: 48.7 个百分点
downstream_accuracy · 尚未评估
报告: 37.4 个百分点
downstream_accuracy · 尚未评估
报告: 31.4 个百分点
downstream_accuracy · 尚未评估
报告: 37.3 个百分点
downstream_accuracy · 尚未评估
报告: 8 个百分点
downstream_accuracy · 尚未评估
报告: 53.4 个百分点
downstream_accuracy · 尚未评估
报告: 24 个百分点
downstream_accuracy · 尚未评估
报告: 8.5 个百分点
downstream_accuracy · 尚未评估
报告: 48.4 个百分点
downstream_accuracy · 尚未评估
报告: 17.3 个百分点
downstream_accuracy · 尚未评估
报告: 57.8 个百分点
downstream_accuracy · 尚未评估
报告: 60 个百分点
downstream_accuracy · 尚未评估
报告: 51 个百分点
downstream_accuracy · 尚未评估
报告: 38.3 个百分点
Downstream task evaluation after post-training stage SFT for 7b_trkd.实验受阻,详见具体原因报告 49.9 个百分点
报告
49.9 个百分点
观测
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · 尚未评估
报告: 49.9 个百分点
downstream_accuracy · 尚未评估
报告: 31.1 个百分点
downstream_accuracy · 尚未评估
报告: 26.9 个百分点
downstream_accuracy · 尚未评估
报告: 31.5 个百分点
downstream_accuracy · 尚未评估
报告: 37.1 个百分点
downstream_accuracy · 尚未评估
报告: 7.6 个百分点
downstream_accuracy · 尚未评估
报告: 53.7 个百分点
downstream_accuracy · 尚未评估
报告: 23.1 个百分点
downstream_accuracy · 尚未评估
报告: 8.4 个百分点
downstream_accuracy · 尚未评估
报告: 46.5 个百分点
downstream_accuracy · 尚未评估
报告: 17.4 个百分点
downstream_accuracy · 尚未评估
报告: 57.8 个百分点
downstream_accuracy · 尚未评估
报告: 56 个百分点
downstream_accuracy · 尚未评估
报告: 51.4 个百分点
downstream_accuracy · 尚未评估
报告: 37.5 个百分点
downstream_accuracy · 尚未评估
报告: 45.3 个百分点
Mid-training ablation results for Teacher Top-1 Labels relative difference across Reasoning, Factual Recall, and Knowledge task groups.实验受阻,详见具体原因报告 -6.4 个百分点
报告
-6.4 个百分点
观测
—
Ablation checkpoints not released and training exceeds compute limits.
delta_percentage_points · 尚未评估
报告: -6.4 个百分点
delta_percentage_points · 尚未评估
报告: 1.3 个百分点
delta_percentage_points · 尚未评估
报告: -2.8 个百分点
Downstream task evaluation after post-training stage SFT for 13b_fkd.实验受阻,详见具体原因报告 54.1 个百分点
报告
54.1 个百分点
观测
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · 尚未评估
报告: 54.1 个百分点
downstream_accuracy · 尚未评估
报告: 32.5 个百分点
downstream_accuracy · 尚未评估
报告: 28 个百分点
downstream_accuracy · 尚未评估
报告: 31.8 个百分点
downstream_accuracy · 尚未评估
报告: 36.3 个百分点
downstream_accuracy · 尚未评估
报告: 7.6 个百分点
downstream_accuracy · 尚未评估
报告: 54.9 个百分点
downstream_accuracy · 尚未评估
报告: 23.5 个百分点
downstream_accuracy · 尚未评估
报告: 7.9 个百分点
downstream_accuracy · 尚未评估
报告: 45.9 个百分点
downstream_accuracy · 尚未评估
报告: 17.1 个百分点
downstream_accuracy · 尚未评估
报告: 55.6 个百分点
downstream_accuracy · 尚未评估
报告: 57.2 个百分点
downstream_accuracy · 尚未评估
报告: 51.2 个百分点
downstream_accuracy · 尚未评估
报告: 37.1 个百分点
downstream_accuracy · 尚未评估
报告: 44.2 个百分点
Per-task downstream accuracy for mid-training ablation oracle_domain using OLMo-2 7B Instruct teacher.实验受阻,详见具体原因报告 61 个百分点
报告
61 个百分点
观测
—
Per-task ablation checkpoints not released.
downstream_accuracy · 尚未评估
报告: 61 个百分点
downstream_accuracy · 尚未评估
报告: 45.6 个百分点
downstream_accuracy · 尚未评估
报告: 38.5 个百分点
downstream_accuracy · 尚未评估
报告: 31.6 个百分点
downstream_accuracy · 尚未评估
报告: 40.1 个百分点
downstream_accuracy · 尚未评估
报告: 8 个百分点
downstream_accuracy · 尚未评估
报告: 53 个百分点
downstream_accuracy · 尚未评估
报告: 22.6 个百分点
downstream_accuracy · 尚未评估
报告: 8.3 个百分点
downstream_accuracy · 尚未评估
报告: 49.4 个百分点
downstream_accuracy · 尚未评估
报告: 18.8 个百分点
downstream_accuracy · 尚未评估
报告: 61.1 个百分点
downstream_accuracy · 尚未评估
报告: 61.4 个百分点
downstream_accuracy · 尚未评估
报告: 52.6 个百分点
downstream_accuracy · 尚未评估
报告: 38.9 个百分点
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B FKD post-training performance.实验受阻,详见具体原因报告 72.1 个百分点
报告
72.1 个百分点
观测
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · 尚未评估
报告: 72.1 个百分点
downstream_accuracy · 尚未评估
报告: 54.3 个百分点
downstream_accuracy · 尚未评估
报告: 46.5 个百分点
downstream_accuracy · 尚未评估
报告: 31.4 个百分点
downstream_accuracy · 尚未评估
报告: 35.7 个百分点
downstream_accuracy · 尚未评估
报告: 14.6 个百分点
downstream_accuracy · 尚未评估
报告: 54.1 个百分点
downstream_accuracy · 尚未评估
报告: 23.2 个百分点
downstream_accuracy · 尚未评估
报告: 8 个百分点
downstream_accuracy · 尚未评估
报告: 36.1 个百分点
downstream_accuracy · 尚未评估
报告: 18.1 个百分点
downstream_accuracy · 尚未评估
报告: 55 个百分点
downstream_accuracy · 尚未评估
报告: 56.4 个百分点
downstream_accuracy · 尚未评估
报告: 51.2 个百分点
downstream_accuracy · 尚未评估
报告: 35.9 个百分点
downstream_accuracy · 尚未评估
报告: 64.9 个百分点
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q10.实验受阻,详见具体原因报告 61.6 个百分点
报告
61.6 个百分点
观测
—
Threshold sweep checkpoints not released.
downstream_accuracy · 尚未评估
报告: 61.6 个百分点
downstream_accuracy · 尚未评估
报告: 48.8 个百分点
downstream_accuracy · 尚未评估
报告: 39.4 个百分点
downstream_accuracy · 尚未评估
报告: 32.5 个百分点
downstream_accuracy · 尚未评估
报告: 50.2 个百分点
downstream_accuracy · 尚未评估
报告: 11.4 个百分点
downstream_accuracy · 尚未评估
报告: 40.6 个百分点
downstream_accuracy · 尚未评估
报告: 53.7 个百分点
downstream_accuracy · 尚未评估
报告: 23.9 个百分点
downstream_accuracy · 尚未评估
报告: 8.6 个百分点
downstream_accuracy · 尚未评估
报告: 28.7 个百分点
downstream_accuracy · 尚未评估
报告: 50.6 个百分点
downstream_accuracy · 尚未评估
报告: 19 个百分点
downstream_accuracy · 尚未评估
报告: 62 个百分点
downstream_accuracy · 尚未评估
报告: 63.8 个百分点
downstream_accuracy · 尚未评估
报告: 55.5 个百分点
downstream_accuracy · 尚未评估
报告: 40.4 个百分点
downstream_accuracy · 尚未评估
报告: 48.5 个百分点
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B TRKD post-training performance.实验受阻,详见具体原因报告 69.6 个百分点
报告
69.6 个百分点
观测
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · 尚未评估
报告: 69.6 个百分点
downstream_accuracy · 尚未评估
报告: 48.5 个百分点
downstream_accuracy · 尚未评估
报告: 42.6 个百分点
downstream_accuracy · 尚未评估
报告: 31.5 个百分点
downstream_accuracy · 尚未评估
报告: 34.6 个百分点
downstream_accuracy · 尚未评估
报告: 12.2 个百分点
downstream_accuracy · 尚未评估
报告: 53.9 个百分点
downstream_accuracy · 尚未评估
报告: 22.2 个百分点
downstream_accuracy · 尚未评估
报告: 8.2 个百分点
downstream_accuracy · 尚未评估
报告: 37 个百分点
downstream_accuracy · 尚未评估
报告: 17.3 个百分点
downstream_accuracy · 尚未评估
报告: 55.4 个百分点
downstream_accuracy · 尚未评估
报告: 54 个百分点
downstream_accuracy · 尚未评估
报告: 51.3 个百分点
downstream_accuracy · 尚未评估
报告: 33.6 个百分点
downstream_accuracy · 尚未评估
报告: 64.1 个百分点
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B RKD post-training performance.实验受阻,详见具体原因报告 73.8 个百分点
报告
73.8 个百分点
观测
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · 尚未评估
报告: 73.8 个百分点
downstream_accuracy · 尚未评估
报告: 53.4 个百分点
downstream_accuracy · 尚未评估
报告: 49.5 个百分点
downstream_accuracy · 尚未评估
报告: 31.6 个百分点
downstream_accuracy · 尚未评估
报告: 40.5 个百分点
downstream_accuracy · 尚未评估
报告: 17.6 个百分点
downstream_accuracy · 尚未评估
报告: 51.8 个百分点
downstream_accuracy · 尚未评估
报告: 21.9 个百分点
downstream_accuracy · 尚未评估
报告: 7.9 个百分点
downstream_accuracy · 尚未评估
报告: 46.2 个百分点
downstream_accuracy · 尚未评估
报告: 18 个百分点
downstream_accuracy · 尚未评估
报告: 61.1 个百分点
downstream_accuracy · 尚未评估
报告: 53 个百分点
downstream_accuracy · 尚未评估
报告: 51.2 个百分点
downstream_accuracy · 尚未评估
报告: 37.1 个百分点
downstream_accuracy · 尚未评估
报告: 61.6 个百分点
Downstream task evaluation after post-training stage RLVR1 for 7b_fkd.实验受阻,详见具体原因报告 76.7 个百分点
报告
76.7 个百分点
观测
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · 尚未评估
报告: 76.7 个百分点
downstream_accuracy · 尚未评估
报告: 59.4 个百分点
downstream_accuracy · 尚未评估
报告: 51.9 个百分点
downstream_accuracy · 尚未评估
报告: 34.4 个百分点
downstream_accuracy · 尚未评估
报告: 39.2 个百分点
downstream_accuracy · 尚未评估
报告: 16.8 个百分点
downstream_accuracy · 尚未评估
报告: 52.8 个百分点
downstream_accuracy · 尚未评估
报告: 23.1 个百分点
downstream_accuracy · 尚未评估
报告: 8.3 个百分点
downstream_accuracy · 尚未评估
报告: 41.4 个百分点
downstream_accuracy · 尚未评估
报告: 18.3 个百分点
downstream_accuracy · 尚未评估
报告: 59.4 个百分点
downstream_accuracy · 尚未评估
报告: 56.2 个百分点
downstream_accuracy · 尚未评估
报告: 51.2 个百分点
downstream_accuracy · 尚未评估
报告: 38 个百分点
downstream_accuracy · 尚未评估
报告: 67.1 个百分点
Downstream task evaluation after post-training stage DPO for 13b_fkd.实验受阻,详见具体原因报告 61 个百分点
报告
61 个百分点
观测
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · 尚未评估
报告: 61 个百分点
downstream_accuracy · 尚未评估
报告: 34.2 个百分点
downstream_accuracy · 尚未评估
报告: 31.7 个百分点
downstream_accuracy · 尚未评估
报告: 33.4 个百分点
downstream_accuracy · 尚未评估
报告: 36.9 个百分点
downstream_accuracy · 尚未评估
报告: 8.2 个百分点
downstream_accuracy · 尚未评估
报告: 54.6 个百分点
downstream_accuracy · 尚未评估
报告: 23.2 个百分点
downstream_accuracy · 尚未评估
报告: 7.7 个百分点
downstream_accuracy · 尚未评估
报告: 45.9 个百分点
downstream_accuracy · 尚未评估
报告: 17.4 个百分点
downstream_accuracy · 尚未评估
报告: 53.2 个百分点
downstream_accuracy · 尚未评估
报告: 57.2 个百分点
downstream_accuracy · 尚未评估
报告: 52.3 个百分点
downstream_accuracy · 尚未评估
报告: 37.4 个百分点
downstream_accuracy · 尚未评估
报告: 64.7 个百分点
Switch Distillation (SD, q=20%, tau=2) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher, demonstrating substantial gains on reasoning while maintaining factual recall.实验受阻,详见具体原因报告 69.7 个百分点
报告
69.7 个百分点
观测
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · 尚未评估
报告: 69.7 个百分点
downstream_accuracy · 尚未评估
报告: 55.3 个百分点
downstream_accuracy · 尚未评估
报告: 46.1 个百分点
downstream_accuracy · 尚未评估
报告: 32.8 个百分点
downstream_accuracy · 尚未评估
报告: 49.6 个百分点
downstream_accuracy · 尚未评估
报告: 14.8 个百分点
downstream_accuracy · 尚未评估
报告: 54.9 个百分点
downstream_accuracy · 尚未评估
报告: 24.6 个百分点
downstream_accuracy · 尚未评估
报告: 8.4 个百分点
downstream_accuracy · 尚未评估
报告: 51.6 个百分点
downstream_accuracy · 尚未评估
报告: 19.8 个百分点
downstream_accuracy · 尚未评估
报告: 64.7 个百分点
downstream_accuracy · 尚未评估
报告: 64.2 个百分点
downstream_accuracy · 尚未评估
报告: 53.8 个百分点
downstream_accuracy · 尚未评估
报告: 41.5 个百分点
Downstream task evaluation after post-training stage DPO for 13b_rkd.实验受阻,详见具体原因报告 65.6 个百分点
报告
65.6 个百分点
观测
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · 尚未评估
报告: 65.6 个百分点
downstream_accuracy · 尚未评估
报告: 41.2 个百分点
downstream_accuracy · 尚未评估
报告: 35 个百分点
downstream_accuracy · 尚未评估
报告: 33.2 个百分点
downstream_accuracy · 尚未评估
报告: 39.9 个百分点
downstream_accuracy · 尚未评估
报告: 7.4 个百分点
downstream_accuracy · 尚未评估
报告: 53.5 个百分点
downstream_accuracy · 尚未评估
报告: 23.7 个百分点
downstream_accuracy · 尚未评估
报告: 7.9 个百分点
downstream_accuracy · 尚未评估
报告: 47.8 个百分点
downstream_accuracy · 尚未评估
报告: 17.9 个百分点
downstream_accuracy · 尚未评估
报告: 57.1 个百分点
downstream_accuracy · 尚未评估
报告: 56.2 个百分点
downstream_accuracy · 尚未评估
报告: 51.9 个百分点
downstream_accuracy · 尚未评估
报告: 38.5 个百分点
downstream_accuracy · 尚未评估
报告: 64.5 个百分点
Token-Routing Knowledge Distillation (TRKD) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.实验受阻,详见具体原因报告 52.5 个百分点
报告
52.5 个百分点
观测
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · 尚未评估
报告: 52.5 个百分点
downstream_accuracy · 尚未评估
报告: 36.9 个百分点
downstream_accuracy · 尚未评估
报告: 31 个百分点
downstream_accuracy · 尚未评估
报告: 30.3 个百分点
downstream_accuracy · 尚未评估
报告: 36.8 个百分点
downstream_accuracy · 尚未评估
报告: 5.6 个百分点
downstream_accuracy · 尚未评估
报告: 54 个百分点
downstream_accuracy · 尚未评估
报告: 24.1 个百分点
downstream_accuracy · 尚未评估
报告: 7.8 个百分点
downstream_accuracy · 尚未评估
报告: 47.9 个百分点
downstream_accuracy · 尚未评估
报告: 17.4 个百分点
downstream_accuracy · 尚未评估
报告: 58.9 个百分点
downstream_accuracy · 尚未评估
报告: 60 个百分点
downstream_accuracy · 尚未评估
报告: 51.2 个百分点
downstream_accuracy · 尚未评估
报告: 37.5 个百分点
Forward KL Distillation (FKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.实验受阻,详见具体原因报告 52.7 个百分点
报告
52.7 个百分点
观测
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · 尚未评估
报告: 52.7 个百分点
downstream_accuracy · 尚未评估
报告: 37.9 个百分点
downstream_accuracy · 尚未评估
报告: 31.8 个百分点
downstream_accuracy · 尚未评估
报告: 29.4 个百分点
downstream_accuracy · 尚未评估
报告: 33.1 个百分点
downstream_accuracy · 尚未评估
报告: 6.4 个百分点
downstream_accuracy · 尚未评估
报告: 54.2 个百分点
downstream_accuracy · 尚未评估
报告: 24.7 个百分点
downstream_accuracy · 尚未评估
报告: 8.1 个百分点
downstream_accuracy · 尚未评估
报告: 47.5 个百分点
downstream_accuracy · 尚未评估
报告: 16.3 个百分点
downstream_accuracy · 尚未评估
报告: 55.3 个百分点
downstream_accuracy · 尚未评估
报告: 56.2 个百分点
downstream_accuracy · 尚未评估
报告: 51.2 个百分点
downstream_accuracy · 尚未评估
报告: 36.2 个百分点
Standard Next-Token Prediction (NTP) baseline downstream evaluation performance after 60B mid-training tokens across 15 tasks covering Reasoning, Factual Recall, and Knowledge & Commonsense.实验受阻,详见具体原因报告 40.4 个百分点
报告
40.4 个百分点
观测
—
Official model checkpoints from mid-training are not publicly released, and mid-training 1B students from 4T tokens requires compute far exceeding the resource budget.
downstream_accuracy · 尚未评估
报告: 40.4 个百分点
downstream_accuracy · 尚未评估
报告: 29.8 个百分点
downstream_accuracy · 尚未评估
报告: 23.1 个百分点
downstream_accuracy · 尚未评估
报告: 29.9 个百分点
downstream_accuracy · 尚未评估
报告: 29.8 个百分点
downstream_accuracy · 尚未评估
报告: 3.8 个百分点
downstream_accuracy · 尚未评估
报告: 56.7 个百分点
downstream_accuracy · 尚未评估
报告: 25.5 个百分点
downstream_accuracy · 尚未评估
报告: 8.7 个百分点
downstream_accuracy · 尚未评估
报告: 43.6 个百分点
downstream_accuracy · 尚未评估
报告: 15.5 个百分点
downstream_accuracy · 尚未评估
报告: 51.1 个百分点
downstream_accuracy · 尚未评估
报告: 51.8 个百分点
downstream_accuracy · 尚未评估
报告: 51.4 个百分点
downstream_accuracy · 尚未评估
报告: 34.1 个百分点
Token-Routing Knowledge Distillation (TRKD) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.实验受阻,详见具体原因报告 47.8 个百分点
报告
47.8 个百分点
观测
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · 尚未评估
报告: 47.8 个百分点
downstream_accuracy · 尚未评估
报告: 32.7 个百分点
downstream_accuracy · 尚未评估
报告: 28.4 个百分点
downstream_accuracy · 尚未评估
报告: 29.3 个百分点
downstream_accuracy · 尚未评估
报告: 34 个百分点
downstream_accuracy · 尚未评估
报告: 5.6 个百分点
downstream_accuracy · 尚未评估
报告: 53.8 个百分点
downstream_accuracy · 尚未评估
报告: 24.2 个百分点
downstream_accuracy · 尚未评估
报告: 8.3 个百分点
downstream_accuracy · 尚未评估
报告: 45.7 个百分点
downstream_accuracy · 尚未评估
报告: 15.4 个百分点
downstream_accuracy · 尚未评估
报告: 53 个百分点
downstream_accuracy · 尚未评估
报告: 56.6 个百分点
downstream_accuracy · 尚未评估
报告: 50.7 个百分点
downstream_accuracy · 尚未评估
报告: 35.2 个百分点
Downstream task evaluation after post-training stage DPO for 7b_rkd.实验受阻,详见具体原因报告 64.5 个百分点
报告
64.5 个百分点
观测
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · 尚未评估
报告: 64.5 个百分点
downstream_accuracy · 尚未评估
报告: 40 个百分点
downstream_accuracy · 尚未评估
报告: 36.4 个百分点
downstream_accuracy · 尚未评估
报告: 34.2 个百分点
downstream_accuracy · 尚未评估
报告: 40.5 个百分点
downstream_accuracy · 尚未评估
报告: 10 个百分点
downstream_accuracy · 尚未评估
报告: 52.8 个百分点
downstream_accuracy · 尚未评估
报告: 22.3 个百分点
downstream_accuracy · 尚未评估
报告: 7.8 个百分点
downstream_accuracy · 尚未评估
报告: 48.8 个百分点
downstream_accuracy · 尚未评估
报告: 17.3 个百分点
downstream_accuracy · 尚未评估
报告: 57.4 个百分点
downstream_accuracy · 尚未评估
报告: 55 个百分点
downstream_accuracy · 尚未评估
报告: 52.5 个百分点
downstream_accuracy · 尚未评估
报告: 38.2 个百分点
downstream_accuracy · 尚未评估
报告: 64.1 个百分点
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B Switch Distillation post-training performance.实验受阻,详见具体原因报告 77.8 个百分点
报告
77.8 个百分点
观测
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · 尚未评估
报告: 77.8 个百分点
downstream_accuracy · 尚未评估
报告: 62.8 个百分点
downstream_accuracy · 尚未评估
报告: 51.8 个百分点
downstream_accuracy · 尚未评估
报告: 33.3 个百分点
downstream_accuracy · 尚未评估
报告: 42.8 个百分点
downstream_accuracy · 尚未评估
报告: 19.6 个百分点
downstream_accuracy · 尚未评估
报告: 54.2 个百分点
downstream_accuracy · 尚未评估
报告: 24.1 个百分点
downstream_accuracy · 尚未评估
报告: 8.4 个百分点
downstream_accuracy · 尚未评估
报告: 43.1 个百分点
downstream_accuracy · 尚未评估
报告: 18.1 个百分点
downstream_accuracy · 尚未评估
报告: 56.5 个百分点
downstream_accuracy · 尚未评估
报告: 59 个百分点
downstream_accuracy · 尚未评估
报告: 52.5 个百分点
downstream_accuracy · 尚未评估
报告: 38.1 个百分点
downstream_accuracy · 尚未评估
报告: 67.1 个百分点
Downstream task evaluation after post-training stage DPO for 7b_sd.实验受阻,详见具体原因报告 70.9 个百分点
报告
70.9 个百分点
观测
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · 尚未评估
报告: 70.9 个百分点
downstream_accuracy · 尚未评估
报告: 50.8 个百分点
downstream_accuracy · 尚未评估
报告: 44.2 个百分点
downstream_accuracy · 尚未评估
报告: 35.5 个百分点
downstream_accuracy · 尚未评估
报告: 49 个百分点
downstream_accuracy · 尚未评估
报告: 15 个百分点
downstream_accuracy · 尚未评估
报告: 54.7 个百分点
downstream_accuracy · 尚未评估
报告: 23.6 个百分点
downstream_accuracy · 尚未评估
报告: 8.1 个百分点
downstream_accuracy · 尚未评估
报告: 50 个百分点
downstream_accuracy · 尚未评估
报告: 20.1 个百分点
downstream_accuracy · 尚未评估
报告: 61.2 个百分点
downstream_accuracy · 尚未评估
报告: 61 个百分点
downstream_accuracy · 尚未评估
报告: 55.5 个百分点
downstream_accuracy · 尚未评估
报告: 40.9 个百分点
downstream_accuracy · 尚未评估
报告: 62.1 个百分点
Downstream task evaluation after post-training stage SFT for 13b_trkd.实验受阻,详见具体原因报告 49 个百分点
报告
49 个百分点
观测
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · 尚未评估
报告: 49 个百分点
downstream_accuracy · 尚未评估
报告: 29.8 个百分点
downstream_accuracy · 尚未评估
报告: 25.4 个百分点
downstream_accuracy · 尚未评估
报告: 31.6 个百分点
downstream_accuracy · 尚未评估
报告: 35.2 个百分点
downstream_accuracy · 尚未评估
报告: 7.6 个百分点
downstream_accuracy · 尚未评估
报告: 55.4 个百分点
downstream_accuracy · 尚未评估
报告: 23.6 个百分点
downstream_accuracy · 尚未评估
报告: 8.8 个百分点
downstream_accuracy · 尚未评估
报告: 43.8 个百分点
downstream_accuracy · 尚未评估
报告: 16.7 个百分点
downstream_accuracy · 尚未评估
报告: 53.8 个百分点
downstream_accuracy · 尚未评估
报告: 55 个百分点
downstream_accuracy · 尚未评估
报告: 51.3 个百分点
downstream_accuracy · 尚未评估
报告: 36.2 个百分点
downstream_accuracy · 尚未评估
报告: 43.4 个百分点
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q30.实验受阻,详见具体原因报告 66 个百分点
报告
66 个百分点
观测
—
Threshold sweep checkpoints not released.
downstream_accuracy · 尚未评估
报告: 66 个百分点
downstream_accuracy · 尚未评估
报告: 53.1 个百分点
downstream_accuracy · 尚未评估
报告: 42.9 个百分点
downstream_accuracy · 尚未评估
报告: 32.5 个百分点
downstream_accuracy · 尚未评估
报告: 44.4 个百分点
downstream_accuracy · 尚未评估
报告: 11 个百分点
downstream_accuracy · 尚未评估
报告: 41.6 个百分点
downstream_accuracy · 尚未评估
报告: 54.2 个百分点
downstream_accuracy · 尚未评估
报告: 24.7 个百分点
downstream_accuracy · 尚未评估
报告: 8.6 个百分点
downstream_accuracy · 尚未评估
报告: 29.2 个百分点
downstream_accuracy · 尚未评估
报告: 49 个百分点
downstream_accuracy · 尚未评估
报告: 18.4 个百分点
downstream_accuracy · 尚未评估
报告: 60.9 个百分点
downstream_accuracy · 尚未评估
报告: 63.6 个百分点
downstream_accuracy · 尚未评估
报告: 52.9 个百分点
downstream_accuracy · 尚未评估
报告: 38.7 个百分点
downstream_accuracy · 尚未评估
报告: 47.2 个百分点
Downstream task evaluation after post-training stage RLVR1 for 7b_sd.实验受阻,详见具体原因报告 79.3 个百分点
报告
79.3 个百分点
观测
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · 尚未评估
报告: 79.3 个百分点
downstream_accuracy · 尚未评估
报告: 63.4 个百分点
downstream_accuracy · 尚未评估
报告: 52.8 个百分点
downstream_accuracy · 尚未评估
报告: 36.1 个百分点
downstream_accuracy · 尚未评估
报告: 48.8 个百分点
downstream_accuracy · 尚未评估
报告: 17.4 个百分点
downstream_accuracy · 尚未评估
报告: 54.1 个百分点
downstream_accuracy · 尚未评估
报告: 23 个百分点
downstream_accuracy · 尚未评估
报告: 8 个百分点
downstream_accuracy · 尚未评估
报告: 47.2 个百分点
downstream_accuracy · 尚未评估
报告: 19.7 个百分点
downstream_accuracy · 尚未评估
报告: 61.7 个百分点
downstream_accuracy · 尚未评估
报告: 59.8 个百分点
downstream_accuracy · 尚未评估
报告: 53 个百分点
downstream_accuracy · 尚未评估
报告: 39.7 个百分点
downstream_accuracy · 尚未评估
报告: 68.8 个百分点
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q20.实验受阻,详见具体原因报告 66 个百分点
报告
66 个百分点
观测
—
Threshold sweep checkpoints not released.
downstream_accuracy · 尚未评估
报告: 66 个百分点
downstream_accuracy · 尚未评估
报告: 54.8 个百分点
downstream_accuracy · 尚未评估
报告: 42.2 个百分点
downstream_accuracy · 尚未评估
报告: 31.6 个百分点
downstream_accuracy · 尚未评估
报告: 45.8 个百分点
downstream_accuracy · 尚未评估
报告: 12 个百分点
downstream_accuracy · 尚未评估
报告: 42.1 个百分点
downstream_accuracy · 尚未评估
报告: 54.1 个百分点
downstream_accuracy · 尚未评估
报告: 24.9 个百分点
downstream_accuracy · 尚未评估
报告: 9 个百分点
downstream_accuracy · 尚未评估
报告: 29.3 个百分点
downstream_accuracy · 尚未评估
报告: 48.5 个百分点
downstream_accuracy · 尚未评估
报告: 17.7 个百分点
downstream_accuracy · 尚未评估
报告: 58.5 个百分点
downstream_accuracy · 尚未评估
报告: 63.8 个百分点
downstream_accuracy · 尚未评估
报告: 51.2 个百分点
downstream_accuracy · 尚未评估
报告: 39.2 个百分点
downstream_accuracy · 尚未评估
报告: 46.5 个百分点
Mid-training ablation results for Switch Distillation with FKL relative difference across Reasoning, Factual Recall, and Knowledge task groups.实验受阻,详见具体原因报告 -2.9 个百分点
报告
-2.9 个百分点
观测
—
Ablation checkpoints not released and training exceeds compute limits.
delta_percentage_points · 尚未评估
报告: -2.9 个百分点
delta_percentage_points · 尚未评估
报告: -0.2 个百分点
delta_percentage_points · 尚未评估
报告: -1.4 个百分点
Downstream task evaluation after post-training stage DPO for 13b_trkd.实验受阻,详见具体原因报告 53.1 个百分点
报告
53.1 个百分点
观测
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · 尚未评估
报告: 53.1 个百分点
downstream_accuracy · 尚未评估
报告: 34.9 个百分点
downstream_accuracy · 尚未评估
报告: 30.7 个百分点
downstream_accuracy · 尚未评估
报告: 32.9 个百分点
downstream_accuracy · 尚未评估
报告: 35.3 个百分点
downstream_accuracy · 尚未评估
报告: 6.8 个百分点
downstream_accuracy · 尚未评估
报告: 54.9 个百分点
downstream_accuracy · 尚未评估
报告: 22.4 个百分点
downstream_accuracy · 尚未评估
报告: 8.5 个百分点
downstream_accuracy · 尚未评估
报告: 43.4 个百分点
downstream_accuracy · 尚未评估
报告: 16.5 个百分点
downstream_accuracy · 尚未评估
报告: 53.2 个百分点
downstream_accuracy · 尚未评估
报告: 56 个百分点
downstream_accuracy · 尚未评估
报告: 52 个百分点
downstream_accuracy · 尚未评估
报告: 35.3 个百分点
downstream_accuracy · 尚未评估
报告: 61.7 个百分点
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B Switch Distillation post-training performance.实验受阻,详见具体原因报告 79.8 个百分点
报告
79.8 个百分点
观测
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · 尚未评估
报告: 79.8 个百分点
downstream_accuracy · 尚未评估
报告: 65 个百分点
downstream_accuracy · 尚未评估
报告: 52.7 个百分点
downstream_accuracy · 尚未评估
报告: 35.6 个百分点
downstream_accuracy · 尚未评估
报告: 48.2 个百分点
downstream_accuracy · 尚未评估
报告: 22.4 个百分点
downstream_accuracy · 尚未评估
报告: 53.6 个百分点
downstream_accuracy · 尚未评估
报告: 22.7 个百分点
downstream_accuracy · 尚未评估
报告: 8.2 个百分点
downstream_accuracy · 尚未评估
报告: 48.1 个百分点
downstream_accuracy · 尚未评估
报告: 19.8 个百分点
downstream_accuracy · 尚未评估
报告: 62.6 个百分点
downstream_accuracy · 尚未评估
报告: 59.4 个百分点
downstream_accuracy · 尚未评估
报告: 54.6 个百分点
downstream_accuracy · 尚未评估
报告: 39.6 个百分点
downstream_accuracy · 尚未评估
报告: 69.5 个百分点
Reverse KL Distillation (RKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.实验受阻,详见具体原因报告 62.1 个百分点
报告
62.1 个百分点
观测
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · 尚未评估
报告: 62.1 个百分点
downstream_accuracy · 尚未评估
报告: 46.2 个百分点
downstream_accuracy · 尚未评估
报告: 39.2 个百分点
downstream_accuracy · 尚未评估
报告: 31.6 个百分点
downstream_accuracy · 尚未评估
报告: 43.1 个百分点
downstream_accuracy · 尚未评估
报告: 10.4 个百分点
downstream_accuracy · 尚未评估
报告: 53.1 个百分点
downstream_accuracy · 尚未评估
报告: 24.1 个百分点
downstream_accuracy · 尚未评估
报告: 8.4 个百分点
downstream_accuracy · 尚未评估
报告: 50.1 个百分点
downstream_accuracy · 尚未评估
报告: 18.9 个百分点
downstream_accuracy · 尚未评估
报告: 62.5 个百分点
downstream_accuracy · 尚未评估
报告: 62.8 个百分点
downstream_accuracy · 尚未评估
报告: 52.5 个百分点
downstream_accuracy · 尚未评估
报告: 39.3 个百分点
Downstream task evaluation after post-training stage SFT for 7b_fkd.实验受阻,详见具体原因报告 54.7 个百分点
报告
54.7 个百分点
观测
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · 尚未评估
报告: 54.7 个百分点
downstream_accuracy · 尚未评估
报告: 34.2 个百分点
downstream_accuracy · 尚未评估
报告: 30.7 个百分点
downstream_accuracy · 尚未评估
报告: 31.3 个百分点
downstream_accuracy · 尚未评估
报告: 39.1 个百分点
downstream_accuracy · 尚未评估
报告: 9.4 个百分点
downstream_accuracy · 尚未评估
报告: 54.1 个百分点
downstream_accuracy · 尚未评估
报告: 23.3 个百分点
downstream_accuracy · 尚未评估
报告: 8.2 个百分点
downstream_accuracy · 尚未评估
报告: 49.1 个百分点
downstream_accuracy · 尚未评估
报告: 18.5 个百分点
downstream_accuracy · 尚未评估
报告: 61.7 个百分点
downstream_accuracy · 尚未评估
报告: 59.4 个百分点
downstream_accuracy · 尚未评估
报告: 51.9 个百分点
downstream_accuracy · 尚未评估
报告: 39.7 个百分点
downstream_accuracy · 尚未评估
报告: 46 个百分点
Downstream task evaluation after post-training stage RLVR1 for 13b_sd.实验受阻,详见具体原因报告 78.8 个百分点
报告
78.8 个百分点
观测
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · 尚未评估
报告: 78.8 个百分点
downstream_accuracy · 尚未评估
报告: 65.6 个百分点
downstream_accuracy · 尚未评估
报告: 52.5 个百分点
downstream_accuracy · 尚未评估
报告: 33.5 个百分点
downstream_accuracy · 尚未评估
报告: 43.3 个百分点
downstream_accuracy · 尚未评估
报告: 17.8 个百分点
downstream_accuracy · 尚未评估
报告: 54.3 个百分点
downstream_accuracy · 尚未评估
报告: 24 个百分点
downstream_accuracy · 尚未评估
报告: 8.2 个百分点
downstream_accuracy · 尚未评估
报告: 38.7 个百分点
downstream_accuracy · 尚未评估
报告: 17.7 个百分点
downstream_accuracy · 尚未评估
报告: 56.2 个百分点
downstream_accuracy · 尚未评估
报告: 57.8 个百分点
downstream_accuracy · 尚未评估
报告: 51.7 个百分点
downstream_accuracy · 尚未评估
报告: 37.7 个百分点
downstream_accuracy · 尚未评估
报告: 68.9 个百分点
Per-task downstream accuracy for mid-training ablation sd_fkl using OLMo-2 7B Instruct teacher.实验受阻,详见具体原因报告 66.7 个百分点
报告
66.7 个百分点
观测
—
Per-task ablation checkpoints not released.
downstream_accuracy · 尚未评估
报告: 66.7 个百分点
downstream_accuracy · 尚未评估
报告: 52 个百分点
downstream_accuracy · 尚未评估
报告: 43.7 个百分点
downstream_accuracy · 尚未评估
报告: 31.2 个百分点
downstream_accuracy · 尚未评估
报告: 46.4 个百分点
downstream_accuracy · 尚未评估
报告: 11 个百分点
downstream_accuracy · 尚未评估
报告: 54.8 个百分点
downstream_accuracy · 尚未评估
报告: 24.3 个百分点
downstream_accuracy · 尚未评估
报告: 8.2 个百分点
downstream_accuracy · 尚未评估
报告: 50.5 个百分点
downstream_accuracy · 尚未评估
报告: 19.1 个百分点
downstream_accuracy · 尚未评估
报告: 63.7 个百分点
downstream_accuracy · 尚未评估
报告: 62.6 个百分点
downstream_accuracy · 尚未评估
报告: 52.2 个百分点
downstream_accuracy · 尚未评估
报告: 39.6 个百分点
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B RKD post-training performance.实验受阻,详见具体原因报告 75.1 个百分点
报告
75.1 个百分点
观测
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · 尚未评估
报告: 75.1 个百分点
downstream_accuracy · 尚未评估
报告: 58.1 个百分点
downstream_accuracy · 尚未评估
报告: 49.3 个百分点
downstream_accuracy · 尚未评估
报告: 33.3 个百分点
downstream_accuracy · 尚未评估
报告: 38.7 个百分点
downstream_accuracy · 尚未评估
报告: 16 个百分点
downstream_accuracy · 尚未评估
报告: 53.9 个百分点
downstream_accuracy · 尚未评估
报告: 22.3 个百分点
downstream_accuracy · 尚未评估
报告: 7.7 个百分点
downstream_accuracy · 尚未评估
报告: 38.2 个百分点
downstream_accuracy · 尚未评估
报告: 17 个百分点
downstream_accuracy · 尚未评估
报告: 56 个百分点
downstream_accuracy · 尚未评估
报告: 55.2 个百分点
downstream_accuracy · 尚未评估
报告: 51.1 个百分点
downstream_accuracy · 尚未评估
报告: 36.6 个百分点
downstream_accuracy · 尚未评估
报告: 64.1 个百分点
Mid-training ablation results for Random Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.实验受阻,详见具体原因报告 -6.5 个百分点
报告
-6.5 个百分点
观测
—
Ablation checkpoints not released and training exceeds compute limits.
delta_percentage_points · 尚未评估
报告: -6.5 个百分点
delta_percentage_points · 尚未评估
报告: -0.8 个百分点
delta_percentage_points · 尚未评估
报告: -2 个百分点
Downstream task evaluation after post-training stage RLVR1 for 7b_trkd.实验受阻,详见具体原因报告 71 个百分点
报告
71 个百分点
观测
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · 尚未评估
报告: 71 个百分点
downstream_accuracy · 尚未评估
报告: 54.6 个百分点
downstream_accuracy · 尚未评估
报告: 46.8 个百分点
downstream_accuracy · 尚未评估
报告: 33 个百分点
downstream_accuracy · 尚未评估
报告: 37.4 个百分点
downstream_accuracy · 尚未评估
报告: 11.4 个百分点
downstream_accuracy · 尚未评估
报告: 52.8 个百分点
downstream_accuracy · 尚未评估
报告: 22.7 个百分点
downstream_accuracy · 尚未评估
报告: 8 个百分点
downstream_accuracy · 尚未评估
报告: 40.2 个百分点
downstream_accuracy · 尚未评估
报告: 17.8 个百分点
downstream_accuracy · 尚未评估
报告: 58.9 个百分点
downstream_accuracy · 尚未评估
报告: 55.6 个百分点
downstream_accuracy · 尚未评估
报告: 52 个百分点
downstream_accuracy · 尚未评估
报告: 35.9 个百分点
downstream_accuracy · 尚未评估
报告: 65.2 个百分点
Per-task downstream accuracy for mid-training ablation teacher_correct using OLMo-2 7B Instruct teacher.实验受阻,详见具体原因报告 63.8 个百分点
报告
63.8 个百分点
观测
—
Per-task ablation checkpoints not released.
downstream_accuracy · 尚未评估
报告: 63.8 个百分点
downstream_accuracy · 尚未评估
报告: 49.5 个百分点
downstream_accuracy · 尚未评估
报告: 41.7 个百分点
downstream_accuracy · 尚未评估
报告: 31.8 个百分点
downstream_accuracy · 尚未评估
报告: 45.2 个百分点
downstream_accuracy · 尚未评估
报告: 9.8 个百分点
downstream_accuracy · 尚未评估
报告: 44.5 个百分点
downstream_accuracy · 尚未评估
报告: 19.4 个百分点
downstream_accuracy · 尚未评估
报告: 8.7 个百分点
downstream_accuracy · 尚未评估
报告: 50.6 个百分点
downstream_accuracy · 尚未评估
报告: 19.3 个百分点
downstream_accuracy · 尚未评估
报告: 63.2 个百分点
downstream_accuracy · 尚未评估
报告: 64.6 个百分点
downstream_accuracy · 尚未评估
报告: 52 个百分点
downstream_accuracy · 尚未评估
报告: 39.9 个百分点
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q20.实验受阻,详见具体原因报告 69.7 个百分点
报告
69.7 个百分点
观测
—
Threshold sweep checkpoints not released.
downstream_accuracy · 尚未评估
报告: 69.7 个百分点
downstream_accuracy · 尚未评估
报告: 55.3 个百分点
downstream_accuracy · 尚未评估
报告: 46.1 个百分点
downstream_accuracy · 尚未评估
报告: 32.8 个百分点
downstream_accuracy · 尚未评估
报告: 49.6 个百分点
downstream_accuracy · 尚未评估
报告: 14.8 个百分点
downstream_accuracy · 尚未评估
报告: 44.7 个百分点
downstream_accuracy · 尚未评估
报告: 54.9 个百分点
downstream_accuracy · 尚未评估
报告: 24.6 个百分点
downstream_accuracy · 尚未评估
报告: 8.4 个百分点
downstream_accuracy · 尚未评估
报告: 29.3 个百分点
downstream_accuracy · 尚未评估
报告: 51.6 个百分点
downstream_accuracy · 尚未评估
报告: 19.8 个百分点
downstream_accuracy · 尚未评估
报告: 64.7 个百分点
downstream_accuracy · 尚未评估
报告: 64.2 个百分点
downstream_accuracy · 尚未评估
报告: 53.8 个百分点
downstream_accuracy · 尚未评估
报告: 41.5 个百分点
downstream_accuracy · 尚未评估
报告: 49.3 个百分点
Mid-training ablation results for Teacher-Correct Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.实验受阻,详见具体原因报告 -4.4 个百分点
报告
-4.4 个百分点
观测
—
Ablation checkpoints not released and training exceeds compute limits.
delta_percentage_points · 尚未评估
报告: -4.4 个百分点
delta_percentage_points · 尚未评估
报告: -5.1 个百分点
delta_percentage_points · 尚未评估
报告: -1 个百分点
Per-task downstream accuracy for mid-training ablation always_ce using OLMo-2 7B Instruct teacher.实验受阻,详见具体原因报告 69.1 个百分点
报告
69.1 个百分点
观测
—
Per-task ablation checkpoints not released.
downstream_accuracy · 尚未评估
报告: 69.1 个百分点
downstream_accuracy · 尚未评估
报告: 55.2 个百分点
downstream_accuracy · 尚未评估
报告: 46.5 个百分点
downstream_accuracy · 尚未评估
报告: 33.3 个百分点
downstream_accuracy · 尚未评估
报告: 50.1 个百分点
downstream_accuracy · 尚未评估
报告: 12 个百分点
downstream_accuracy · 尚未评估
报告: 54.3 个百分点
downstream_accuracy · 尚未评估
报告: 24.5 个百分点
downstream_accuracy · 尚未评估
报告: 9.4 个百分点
downstream_accuracy · 尚未评估
报告: 51.1 个百分点
downstream_accuracy · 尚未评估
报告: 19.9 个百分点
downstream_accuracy · 尚未评估
报告: 63.7 个百分点
downstream_accuracy · 尚未评估
报告: 63.6 个百分点
downstream_accuracy · 尚未评估
报告: 52.7 个百分点
downstream_accuracy · 尚未评估
报告: 41.1 个百分点
Downstream task evaluation after post-training stage RLVR1 for 13b_trkd.实验受阻,详见具体原因报告 72.4 个百分点
报告
72.4 个百分点
观测
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · 尚未评估
报告: 72.4 个百分点
downstream_accuracy · 尚未评估
报告: 53.3 个百分点
downstream_accuracy · 尚未评估
报告: 47.6 个百分点
downstream_accuracy · 尚未评估
报告: 32.7 个百分点
downstream_accuracy · 尚未评估
报告: 35 个百分点
downstream_accuracy · 尚未评估
报告: 10.4 个百分点
downstream_accuracy · 尚未评估
报告: 54.7 个百分点
downstream_accuracy · 尚未评估
报告: 22.2 个百分点
downstream_accuracy · 尚未评估
报告: 8.6 个百分点
downstream_accuracy · 尚未评估
报告: 36.6 个百分点
downstream_accuracy · 尚未评估
报告: 16.6 个百分点
downstream_accuracy · 尚未评估
报告: 54.1 个百分点
downstream_accuracy · 尚未评估
报告: 54.6 个百分点
downstream_accuracy · 尚未评估
报告: 51.7 个百分点
downstream_accuracy · 尚未评估
报告: 34.5 个百分点
downstream_accuracy · 尚未评估
报告: 67.1 个百分点
Downstream task evaluation after post-training stage DPO for 7b_fkd.实验受阻,详见具体原因报告 60.4 个百分点
报告
60.4 个百分点
观测
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · 尚未评估
报告: 60.4 个百分点
downstream_accuracy · 尚未评估
报告: 35.8 个百分点
downstream_accuracy · 尚未评估
报告: 33.1 个百分点
downstream_accuracy · 尚未评估
报告: 33.1 个百分点
downstream_accuracy · 尚未评估
报告: 39.8 个百分点
downstream_accuracy · 尚未评估
报告: 7.8 个百分点
downstream_accuracy · 尚未评估
报告: 53.6 个百分点
downstream_accuracy · 尚未评估
报告: 23.1 个百分点
downstream_accuracy · 尚未评估
报告: 8.3 个百分点
downstream_accuracy · 尚未评估
报告: 49.2 个百分点
downstream_accuracy · 尚未评估
报告: 18.5 个百分点
downstream_accuracy · 尚未评估
报告: 59.2 个百分点
downstream_accuracy · 尚未评估
报告: 58 个百分点
downstream_accuracy · 尚未评估
报告: 51.6 个百分点
downstream_accuracy · 尚未评估
报告: 39.7 个百分点
downstream_accuracy · 尚未评估
报告: 64.1 个百分点
定性结论
The stage-dependent distillation tradeoff generalizes to the SmolLM family (SmolLM2 360M student, 1.7B Instruct teacher, SmolLM3 Stage-3 data), where mid-training KD exhibits the reasoning-recall deficit and Switch Distillation mitigates it.实验受阻,详见具体原因
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Forward KL consistently exhibits higher CE-KL gradient cosine alignment than Reverse KL across pre-training and mid-training, with the gap widening at higher alpha and later training steps.实验受阻,详见具体原因
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Gold-token probability, gradient attenuation relative to NTP, and factual recall deficits under KD hold consistently for OLMo-2 1B and 13B Instruct teachers.实验受阻,详见具体原因
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Switch Distillation substantially accelerates reasoning acquisition during mid-training, surpassing the final 60B-token NTP reasoning macro-average within 2.5B tokens (24x fewer tokens), while standard FKD requires 5.0B tokens (12x fewer).实验受阻,详见具体原因
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Lower teacher predictive entropy corresponds to substantially higher teacher top-1 agreement with the ground-truth token across all data domains and teacher sizes.实验受阻,详见具体原因
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Higher teacher entropy corresponds to strictly lower gold-token probability, attenuates the gold-token gradient relative to NTP (reaching ~0.5x NTP for Q5 facts), and produces larger downstream factual recall deficits.实验受阻,详见具体原因
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Across distillation strengths alpha in [0.0, 1.0], KL directions (forward and reverse), and teacher sizes (1B, 7B, 13B), mid-training distillation traces a reasoning-recall frontier that falls below NTP, which Switch Distillation mitigates.实验受阻,详见具体原因
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Knowledge distillation generally improves reasoning, but its effect on factual recall changes across training stages: during pre-training, distillation improves both reasoning and factual recall over NTP, whereas during mid-training it improves reasoning at the expense of factual recall.实验受阻,详见具体原因
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.