Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
PublicMore
Research claims
0 / 76 claims verified
the rest still being verified
Key metricsPaper-reported → reproduced · click a row to expand
Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).Inconclusive4/12 measurements supported · 12 assessedReported 0.815 score0.8079 score
Reported
0.815 score
Observed
0.8079 score
Delta
−0.0071 score
The execution did not produce enough comparable measurements to determine whether the paper's claim is supported. · This does not refute the paper's claim; it means the available evidence can neither confirm nor refute it yet.
roc_auc · Supported
Reported: 0.815 score
Observed: 0.8079 score
Reported ROC AUC of 0.815 was observed at 0.8079 (relative diff 0.87%), showing strong predictive entropy separation on OLMo-2 1B Base.
roc_auc · Supported
Reported: 0.77 score
Observed: 0.7967 score
Reported ROC AUC of 0.770 was observed at 0.7967 (relative diff 3.47%), confirming high ROC AUC separation on OLMo-2 1B SFT.
roc_auc · Supported
Reported: 0.744 score
Observed: 0.7931 score
Reported ROC AUC of 0.744 was observed at 0.7931 (relative diff 6.60%), maintaining the substantive finding of strong discrimination on OLMo-2 1B DPO.
roc_auc · Supported
Reported: 0.748 score
Observed: 0.7973 score
Reported ROC AUC of 0.748 was observed at 0.7973 (relative diff 6.59%), confirming procedural entropy remains distinct on OLMo-2 1B Instruct.
roc_auc · Inconclusive
Reported: 0.816 score
Observed: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json
roc_auc · Inconclusive
Reported: 0.777 score
Observed: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json
roc_auc · Inconclusive
Reported: 0.76 score
Observed: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json
roc_auc · Inconclusive
Reported: 0.761 score
Observed: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json
roc_auc · Inconclusive
Reported: 0.826 score
Observed: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json
roc_auc · Inconclusive
Reported: 0.814 score
Observed: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json
roc_auc · Inconclusive
Reported: 0.809 score
Observed: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json
roc_auc · Inconclusive
Reported: 0.811 score
Observed: —
The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json
Mid-training ablation results for Always CE relative difference across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported -0.3 pp
Reported
-0.3 pp
Observed
—
Ablation checkpoints not released and training exceeds compute limits.
delta_percentage_points · Not assessed
Reported: -0.3 pp
delta_percentage_points · Not assessed
Reported: 0.1 pp
delta_percentage_points · Not assessed
Reported: -0.6 pp
Downstream task evaluation after post-training stage RLVR1 for 13b_rkd.Experiment blocked; see the specific reasonReported 75.6 pp
Reported
75.6 pp
Observed
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · Not assessed
Reported: 75.6 pp
downstream_accuracy · Not assessed
Reported: 60.1 pp
downstream_accuracy · Not assessed
Reported: 50.3 pp
downstream_accuracy · Not assessed
Reported: 33.3 pp
downstream_accuracy · Not assessed
Reported: 38.1 pp
downstream_accuracy · Not assessed
Reported: 15.8 pp
downstream_accuracy · Not assessed
Reported: 53.3 pp
downstream_accuracy · Not assessed
Reported: 22.3 pp
downstream_accuracy · Not assessed
Reported: 8 pp
downstream_accuracy · Not assessed
Reported: 36.2 pp
downstream_accuracy · Not assessed
Reported: 16.9 pp
downstream_accuracy · Not assessed
Reported: 55.9 pp
downstream_accuracy · Not assessed
Reported: 55.2 pp
downstream_accuracy · Not assessed
Reported: 51.3 pp
downstream_accuracy · Not assessed
Reported: 36.4 pp
downstream_accuracy · Not assessed
Reported: 63.6 pp
Downstream task evaluation after post-training stage RLVR1 for 13b_fkd.Experiment blocked; see the specific reasonReported 73 pp
Reported
73 pp
Observed
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · Not assessed
Reported: 73 pp
downstream_accuracy · Not assessed
Reported: 55.5 pp
downstream_accuracy · Not assessed
Reported: 47.3 pp
downstream_accuracy · Not assessed
Reported: 33.7 pp
downstream_accuracy · Not assessed
Reported: 36.1 pp
downstream_accuracy · Not assessed
Reported: 14 pp
downstream_accuracy · Not assessed
Reported: 54.7 pp
downstream_accuracy · Not assessed
Reported: 23.3 pp
downstream_accuracy · Not assessed
Reported: 8.3 pp
downstream_accuracy · Not assessed
Reported: 35 pp
downstream_accuracy · Not assessed
Reported: 17.2 pp
downstream_accuracy · Not assessed
Reported: 54 pp
downstream_accuracy · Not assessed
Reported: 56.6 pp
downstream_accuracy · Not assessed
Reported: 51.7 pp
downstream_accuracy · Not assessed
Reported: 35.4 pp
downstream_accuracy · Not assessed
Reported: 63.2 pp
Downstream task evaluation after post-training stage SFT for 13b_rkd.Experiment blocked; see the specific reasonReported 56.8 pp
Reported
56.8 pp
Observed
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · Not assessed
Reported: 56.8 pp
downstream_accuracy · Not assessed
Reported: 36.7 pp
downstream_accuracy · Not assessed
Reported: 31.3 pp
downstream_accuracy · Not assessed
Reported: 32.1 pp
downstream_accuracy · Not assessed
Reported: 38.9 pp
downstream_accuracy · Not assessed
Reported: 11.6 pp
downstream_accuracy · Not assessed
Reported: 55.4 pp
downstream_accuracy · Not assessed
Reported: 23.8 pp
downstream_accuracy · Not assessed
Reported: 8 pp
downstream_accuracy · Not assessed
Reported: 47.3 pp
downstream_accuracy · Not assessed
Reported: 17.3 pp
downstream_accuracy · Not assessed
Reported: 57.8 pp
downstream_accuracy · Not assessed
Reported: 57.4 pp
downstream_accuracy · Not assessed
Reported: 52.2 pp
downstream_accuracy · Not assessed
Reported: 38.1 pp
downstream_accuracy · Not assessed
Reported: 43.6 pp
Downstream task evaluation after post-training stage SFT for 13b_sd.Experiment blocked; see the specific reasonReported 62.7 pp
Reported
62.7 pp
Observed
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · Not assessed
Reported: 62.7 pp
downstream_accuracy · Not assessed
Reported: 42.5 pp
downstream_accuracy · Not assessed
Reported: 34.8 pp
downstream_accuracy · Not assessed
Reported: 31.6 pp
downstream_accuracy · Not assessed
Reported: 43.5 pp
downstream_accuracy · Not assessed
Reported: 10.2 pp
downstream_accuracy · Not assessed
Reported: 54.9 pp
downstream_accuracy · Not assessed
Reported: 24.1 pp
downstream_accuracy · Not assessed
Reported: 8.5 pp
downstream_accuracy · Not assessed
Reported: 48 pp
downstream_accuracy · Not assessed
Reported: 17.7 pp
downstream_accuracy · Not assessed
Reported: 55.1 pp
downstream_accuracy · Not assessed
Reported: 59.4 pp
downstream_accuracy · Not assessed
Reported: 52.2 pp
downstream_accuracy · Not assessed
Reported: 38 pp
downstream_accuracy · Not assessed
Reported: 43.1 pp
Factual acquisition stratification by teacher entropy quintile persists when using OLMo-2 1B Instruct and 13B Instruct teachers.Experiment blocked; see the specific reasonReported 76 pp
Reported
76 pp
Observed
—
Intermediate pre-training and mid-training NTP trajectory checkpoints are not public.
factual_examples_learned · Not assessed
Reported: 76 pp
factual_examples_learned · Not assessed
Reported: 48 pp
factual_examples_learned · Not assessed
Reported: 30 pp
factual_examples_learned · Not assessed
Reported: 12 pp
factual_examples_learned · Not assessed
Reported: 6 pp
factual_examples_learned · Not assessed
Reported: 81 pp
factual_examples_learned · Not assessed
Reported: 61 pp
factual_examples_learned · Not assessed
Reported: 40 pp
factual_examples_learned · Not assessed
Reported: 18 pp
factual_examples_learned · Not assessed
Reported: 8 pp
factual_examples_learned · Not assessed
Reported: 54 pp
factual_examples_learned · Not assessed
Reported: 48 pp
factual_examples_learned · Not assessed
Reported: 37 pp
factual_examples_learned · Not assessed
Reported: 28 pp
factual_examples_learned · Not assessed
Reported: 6 pp
factual_examples_learned · Not assessed
Reported: 66 pp
factual_examples_learned · Not assessed
Reported: 58 pp
factual_examples_learned · Not assessed
Reported: 43 pp
factual_examples_learned · Not assessed
Reported: 32 pp
factual_examples_learned · Not assessed
Reported: 8 pp
Teacher entropy under OLMo-2 7B Instruct strongly predicts factual acquisition under NTP: by the end of pre-training, the student learns 67% of Q1 facts vs 5% of Q5 facts; by mid-training initialization (4T tokens), 80% of Q1 facts are learned vs 7% of Q5 facts.Experiment blocked; see the specific reasonReported 67 pp
Reported
67 pp
Observed
—
The intermediate NTP training checkpoints across pre-training and mid-training trajectories are not released.
factual_examples_learned · Not assessed
Reported: 67 pp
factual_examples_learned · Not assessed
Reported: 47 pp
factual_examples_learned · Not assessed
Reported: 34 pp
factual_examples_learned · Not assessed
Reported: 21 pp
factual_examples_learned · Not assessed
Reported: 5 pp
factual_examples_learned · Not assessed
Reported: 80 pp
factual_examples_learned · Not assessed
Reported: 55 pp
factual_examples_learned · Not assessed
Reported: 42 pp
factual_examples_learned · Not assessed
Reported: 24 pp
factual_examples_learned · Not assessed
Reported: 7 pp
Downstream task evaluation after post-training stage SFT for 7b_rkd.Experiment blocked; see the specific reasonReported 57.7 pp
Reported
57.7 pp
Observed
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · Not assessed
Reported: 57.7 pp
downstream_accuracy · Not assessed
Reported: 35.5 pp
downstream_accuracy · Not assessed
Reported: 32 pp
downstream_accuracy · Not assessed
Reported: 32.1 pp
downstream_accuracy · Not assessed
Reported: 39.9 pp
downstream_accuracy · Not assessed
Reported: 11.2 pp
downstream_accuracy · Not assessed
Reported: 53.2 pp
downstream_accuracy · Not assessed
Reported: 22.4 pp
downstream_accuracy · Not assessed
Reported: 8.4 pp
downstream_accuracy · Not assessed
Reported: 48.9 pp
downstream_accuracy · Not assessed
Reported: 18.2 pp
downstream_accuracy · Not assessed
Reported: 60.9 pp
downstream_accuracy · Not assessed
Reported: 60.2 pp
downstream_accuracy · Not assessed
Reported: 52.2 pp
downstream_accuracy · Not assessed
Reported: 39.2 pp
downstream_accuracy · Not assessed
Reported: 45.8 pp
Switch Distillation (SD, q=20%, tau=2) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.Experiment blocked; see the specific reasonReported 66 pp
Reported
66 pp
Observed
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · Not assessed
Reported: 66 pp
downstream_accuracy · Not assessed
Reported: 54.8 pp
downstream_accuracy · Not assessed
Reported: 42.2 pp
downstream_accuracy · Not assessed
Reported: 31.6 pp
downstream_accuracy · Not assessed
Reported: 45.8 pp
downstream_accuracy · Not assessed
Reported: 12 pp
downstream_accuracy · Not assessed
Reported: 54.1 pp
downstream_accuracy · Not assessed
Reported: 24.9 pp
downstream_accuracy · Not assessed
Reported: 9 pp
downstream_accuracy · Not assessed
Reported: 48.5 pp
downstream_accuracy · Not assessed
Reported: 17.7 pp
downstream_accuracy · Not assessed
Reported: 58.5 pp
downstream_accuracy · Not assessed
Reported: 63.8 pp
downstream_accuracy · Not assessed
Reported: 51.2 pp
downstream_accuracy · Not assessed
Reported: 39.2 pp
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q30.Experiment blocked; see the specific reasonReported 69.8 pp
Reported
69.8 pp
Observed
—
Threshold sweep checkpoints not released.
downstream_accuracy · Not assessed
Reported: 69.8 pp
downstream_accuracy · Not assessed
Reported: 54.1 pp
downstream_accuracy · Not assessed
Reported: 45.6 pp
downstream_accuracy · Not assessed
Reported: 33.8 pp
downstream_accuracy · Not assessed
Reported: 47.2 pp
downstream_accuracy · Not assessed
Reported: 13 pp
downstream_accuracy · Not assessed
Reported: 43.9 pp
downstream_accuracy · Not assessed
Reported: 54.7 pp
downstream_accuracy · Not assessed
Reported: 25.3 pp
downstream_accuracy · Not assessed
Reported: 8.8 pp
downstream_accuracy · Not assessed
Reported: 29.6 pp
downstream_accuracy · Not assessed
Reported: 51 pp
downstream_accuracy · Not assessed
Reported: 19.5 pp
downstream_accuracy · Not assessed
Reported: 63.5 pp
downstream_accuracy · Not assessed
Reported: 62.4 pp
downstream_accuracy · Not assessed
Reported: 53 pp
downstream_accuracy · Not assessed
Reported: 40.7 pp
downstream_accuracy · Not assessed
Reported: 48.4 pp
Per-task downstream accuracy for mid-training ablation random_routing using OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 61.5 pp
Reported
61.5 pp
Observed
—
Per-task ablation checkpoints not released.
downstream_accuracy · Not assessed
Reported: 61.5 pp
downstream_accuracy · Not assessed
Reported: 45.5 pp
downstream_accuracy · Not assessed
Reported: 38.6 pp
downstream_accuracy · Not assessed
Reported: 32 pp
downstream_accuracy · Not assessed
Reported: 40.6 pp
downstream_accuracy · Not assessed
Reported: 10.8 pp
downstream_accuracy · Not assessed
Reported: 53.2 pp
downstream_accuracy · Not assessed
Reported: 23.9 pp
downstream_accuracy · Not assessed
Reported: 8.5 pp
downstream_accuracy · Not assessed
Reported: 49.6 pp
downstream_accuracy · Not assessed
Reported: 17.8 pp
downstream_accuracy · Not assessed
Reported: 61.6 pp
downstream_accuracy · Not assessed
Reported: 62.6 pp
downstream_accuracy · Not assessed
Reported: 52.6 pp
downstream_accuracy · Not assessed
Reported: 39.5 pp
Per-task downstream accuracy for mid-training ablation teacher_top1 using OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 62.2 pp
Reported
62.2 pp
Observed
—
Per-task ablation checkpoints not released.
downstream_accuracy · Not assessed
Reported: 62.2 pp
downstream_accuracy · Not assessed
Reported: 46 pp
downstream_accuracy · Not assessed
Reported: 40.1 pp
downstream_accuracy · Not assessed
Reported: 30.7 pp
downstream_accuracy · Not assessed
Reported: 42.1 pp
downstream_accuracy · Not assessed
Reported: 8.8 pp
downstream_accuracy · Not assessed
Reported: 57.9 pp
downstream_accuracy · Not assessed
Reported: 24.7 pp
downstream_accuracy · Not assessed
Reported: 9.3 pp
downstream_accuracy · Not assessed
Reported: 48.8 pp
downstream_accuracy · Not assessed
Reported: 18 pp
downstream_accuracy · Not assessed
Reported: 59.8 pp
downstream_accuracy · Not assessed
Reported: 62 pp
downstream_accuracy · Not assessed
Reported: 51 pp
downstream_accuracy · Not assessed
Reported: 39.2 pp
Mid-training ablation results for Switch Distillation absolute macro-averages across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported 44.7 pp
Reported
44.7 pp
Observed
—
Ablation checkpoints not released and training exceeds compute limits.
macro_average · Not assessed
Reported: 44.7 pp
macro_average · Not assessed
Reported: 29.3 pp
macro_average · Not assessed
Reported: 49.3 pp
Downstream task evaluation after post-training stage SFT for 7b_sd.Experiment blocked; see the specific reasonReported 63.7 pp
Reported
63.7 pp
Observed
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · Not assessed
Reported: 63.7 pp
downstream_accuracy · Not assessed
Reported: 42.7 pp
downstream_accuracy · Not assessed
Reported: 36.7 pp
downstream_accuracy · Not assessed
Reported: 33.5 pp
downstream_accuracy · Not assessed
Reported: 48.9 pp
downstream_accuracy · Not assessed
Reported: 12.8 pp
downstream_accuracy · Not assessed
Reported: 55.1 pp
downstream_accuracy · Not assessed
Reported: 24.5 pp
downstream_accuracy · Not assessed
Reported: 8.5 pp
downstream_accuracy · Not assessed
Reported: 50.3 pp
downstream_accuracy · Not assessed
Reported: 19.4 pp
downstream_accuracy · Not assessed
Reported: 61.3 pp
downstream_accuracy · Not assessed
Reported: 61.4 pp
downstream_accuracy · Not assessed
Reported: 56.1 pp
downstream_accuracy · Not assessed
Reported: 40.6 pp
downstream_accuracy · Not assessed
Reported: 45.5 pp
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q10.Experiment blocked; see the specific reasonReported 62.1 pp
Reported
62.1 pp
Observed
—
Threshold sweep checkpoints not released.
downstream_accuracy · Not assessed
Reported: 62.1 pp
downstream_accuracy · Not assessed
Reported: 49.6 pp
downstream_accuracy · Not assessed
Reported: 38.3 pp
downstream_accuracy · Not assessed
Reported: 28.6 pp
downstream_accuracy · Not assessed
Reported: 44.2 pp
downstream_accuracy · Not assessed
Reported: 7.8 pp
downstream_accuracy · Not assessed
Reported: 38.5 pp
downstream_accuracy · Not assessed
Reported: 52.2 pp
downstream_accuracy · Not assessed
Reported: 23.1 pp
downstream_accuracy · Not assessed
Reported: 9.2 pp
downstream_accuracy · Not assessed
Reported: 28.1 pp
downstream_accuracy · Not assessed
Reported: 46.8 pp
downstream_accuracy · Not assessed
Reported: 17.1 pp
downstream_accuracy · Not assessed
Reported: 55.5 pp
downstream_accuracy · Not assessed
Reported: 59.6 pp
downstream_accuracy · Not assessed
Reported: 52.7 pp
downstream_accuracy · Not assessed
Reported: 37.9 pp
downstream_accuracy · Not assessed
Reported: 45 pp
Downstream task evaluation after post-training stage DPO for 7b_trkd.Experiment blocked; see the specific reasonReported 57.2 pp
Reported
57.2 pp
Observed
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · Not assessed
Reported: 57.2 pp
downstream_accuracy · Not assessed
Reported: 34.9 pp
downstream_accuracy · Not assessed
Reported: 32.3 pp
downstream_accuracy · Not assessed
Reported: 32.4 pp
downstream_accuracy · Not assessed
Reported: 36.6 pp
downstream_accuracy · Not assessed
Reported: 7.2 pp
downstream_accuracy · Not assessed
Reported: 53.1 pp
downstream_accuracy · Not assessed
Reported: 23.2 pp
downstream_accuracy · Not assessed
Reported: 7.9 pp
downstream_accuracy · Not assessed
Reported: 47 pp
downstream_accuracy · Not assessed
Reported: 17.9 pp
downstream_accuracy · Not assessed
Reported: 56.7 pp
downstream_accuracy · Not assessed
Reported: 57.2 pp
downstream_accuracy · Not assessed
Reported: 51.9 pp
downstream_accuracy · Not assessed
Reported: 37.5 pp
downstream_accuracy · Not assessed
Reported: 62.5 pp
Forward KL Distillation (FKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 57.8 pp
Reported
57.8 pp
Observed
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · Not assessed
Reported: 57.8 pp
downstream_accuracy · Not assessed
Reported: 42.6 pp
downstream_accuracy · Not assessed
Reported: 36.2 pp
downstream_accuracy · Not assessed
Reported: 30.6 pp
downstream_accuracy · Not assessed
Reported: 41.9 pp
downstream_accuracy · Not assessed
Reported: 7.6 pp
downstream_accuracy · Not assessed
Reported: 53.9 pp
downstream_accuracy · Not assessed
Reported: 24.6 pp
downstream_accuracy · Not assessed
Reported: 8.4 pp
downstream_accuracy · Not assessed
Reported: 49.5 pp
downstream_accuracy · Not assessed
Reported: 18.7 pp
downstream_accuracy · Not assessed
Reported: 61.9 pp
downstream_accuracy · Not assessed
Reported: 61.8 pp
downstream_accuracy · Not assessed
Reported: 51.5 pp
downstream_accuracy · Not assessed
Reported: 38.6 pp
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B TRKD post-training performance.Experiment blocked; see the specific reasonReported 70.7 pp
Reported
70.7 pp
Observed
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · Not assessed
Reported: 70.7 pp
downstream_accuracy · Not assessed
Reported: 52.9 pp
downstream_accuracy · Not assessed
Reported: 46.3 pp
downstream_accuracy · Not assessed
Reported: 31.8 pp
downstream_accuracy · Not assessed
Reported: 36.8 pp
downstream_accuracy · Not assessed
Reported: 14.4 pp
downstream_accuracy · Not assessed
Reported: 51.7 pp
downstream_accuracy · Not assessed
Reported: 22.3 pp
downstream_accuracy · Not assessed
Reported: 8 pp
downstream_accuracy · Not assessed
Reported: 40.9 pp
downstream_accuracy · Not assessed
Reported: 17.4 pp
downstream_accuracy · Not assessed
Reported: 59 pp
downstream_accuracy · Not assessed
Reported: 53.6 pp
downstream_accuracy · Not assessed
Reported: 51.5 pp
downstream_accuracy · Not assessed
Reported: 35.6 pp
downstream_accuracy · Not assessed
Reported: 65.2 pp
Teacher predictive entropy distinguishes procedural from knowledge-intensive domains across diverse instruction-tuned open-weight model families (OLMo-3 7B Instruct: 0.771, Qwen 3 8B: 0.705, Gemma-3 12B it: 0.707, Granite 3.3 8B Instruct: 0.696).Awaiting reproductionReported 0.771 score
Reported
0.771 score
Observed
—
Experiment plan ready; no runs yet.
roc_auc · Not assessed
Reported: 0.771 score
roc_auc · Not assessed
Reported: 0.705 score
roc_auc · Not assessed
Reported: 0.707 score
roc_auc · Not assessed
Reported: 0.696 score
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B FKD post-training performance.Experiment blocked; see the specific reasonReported 76.2 pp
Reported
76.2 pp
Observed
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · Not assessed
Reported: 76.2 pp
downstream_accuracy · Not assessed
Reported: 56.4 pp
downstream_accuracy · Not assessed
Reported: 51.6 pp
downstream_accuracy · Not assessed
Reported: 33.8 pp
downstream_accuracy · Not assessed
Reported: 38.4 pp
downstream_accuracy · Not assessed
Reported: 18 pp
downstream_accuracy · Not assessed
Reported: 51.9 pp
downstream_accuracy · Not assessed
Reported: 22.2 pp
downstream_accuracy · Not assessed
Reported: 8.1 pp
downstream_accuracy · Not assessed
Reported: 42.2 pp
downstream_accuracy · Not assessed
Reported: 18.6 pp
downstream_accuracy · Not assessed
Reported: 60.2 pp
downstream_accuracy · Not assessed
Reported: 57.4 pp
downstream_accuracy · Not assessed
Reported: 51.5 pp
downstream_accuracy · Not assessed
Reported: 38 pp
downstream_accuracy · Not assessed
Reported: 64.7 pp
Downstream task evaluation after post-training stage DPO for 13b_sd.Experiment blocked; see the specific reasonReported 69.9 pp
Reported
69.9 pp
Observed
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · Not assessed
Reported: 69.9 pp
downstream_accuracy · Not assessed
Reported: 45.7 pp
downstream_accuracy · Not assessed
Reported: 39.5 pp
downstream_accuracy · Not assessed
Reported: 32.2 pp
downstream_accuracy · Not assessed
Reported: 44.2 pp
downstream_accuracy · Not assessed
Reported: 10 pp
downstream_accuracy · Not assessed
Reported: 54.5 pp
downstream_accuracy · Not assessed
Reported: 23.9 pp
downstream_accuracy · Not assessed
Reported: 8.2 pp
downstream_accuracy · Not assessed
Reported: 47.3 pp
downstream_accuracy · Not assessed
Reported: 17.9 pp
downstream_accuracy · Not assessed
Reported: 56.1 pp
downstream_accuracy · Not assessed
Reported: 59.2 pp
downstream_accuracy · Not assessed
Reported: 52.5 pp
downstream_accuracy · Not assessed
Reported: 38.7 pp
downstream_accuracy · Not assessed
Reported: 61 pp
Mid-training ablation results for Oracle Domain Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported -7.2 pp
Reported
-7.2 pp
Observed
—
Ablation checkpoints not released and training exceeds compute limits.
delta_percentage_points · Not assessed
Reported: -7.2 pp
delta_percentage_points · Not assessed
Reported: -1.3 pp
delta_percentage_points · Not assessed
Reported: -2.3 pp
Downstream task evaluation after post-training stage RLVR1 for 7b_rkd.Experiment blocked; see the specific reasonReported 77.9 pp
Reported
77.9 pp
Observed
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · Not assessed
Reported: 77.9 pp
downstream_accuracy · Not assessed
Reported: 54.3 pp
downstream_accuracy · Not assessed
Reported: 52.5 pp
downstream_accuracy · Not assessed
Reported: 34.2 pp
downstream_accuracy · Not assessed
Reported: 40.1 pp
downstream_accuracy · Not assessed
Reported: 15.6 pp
downstream_accuracy · Not assessed
Reported: 52.1 pp
downstream_accuracy · Not assessed
Reported: 21.9 pp
downstream_accuracy · Not assessed
Reported: 7.8 pp
downstream_accuracy · Not assessed
Reported: 45.3 pp
downstream_accuracy · Not assessed
Reported: 16.8 pp
downstream_accuracy · Not assessed
Reported: 57.6 pp
downstream_accuracy · Not assessed
Reported: 50.4 pp
downstream_accuracy · Not assessed
Reported: 51.5 pp
downstream_accuracy · Not assessed
Reported: 36.1 pp
downstream_accuracy · Not assessed
Reported: 59.3 pp
Reverse KL Distillation (RKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.Experiment blocked; see the specific reasonReported 59.2 pp
Reported
59.2 pp
Observed
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · Not assessed
Reported: 59.2 pp
downstream_accuracy · Not assessed
Reported: 48.7 pp
downstream_accuracy · Not assessed
Reported: 37.4 pp
downstream_accuracy · Not assessed
Reported: 31.4 pp
downstream_accuracy · Not assessed
Reported: 37.3 pp
downstream_accuracy · Not assessed
Reported: 8 pp
downstream_accuracy · Not assessed
Reported: 53.4 pp
downstream_accuracy · Not assessed
Reported: 24 pp
downstream_accuracy · Not assessed
Reported: 8.5 pp
downstream_accuracy · Not assessed
Reported: 48.4 pp
downstream_accuracy · Not assessed
Reported: 17.3 pp
downstream_accuracy · Not assessed
Reported: 57.8 pp
downstream_accuracy · Not assessed
Reported: 60 pp
downstream_accuracy · Not assessed
Reported: 51 pp
downstream_accuracy · Not assessed
Reported: 38.3 pp
Downstream task evaluation after post-training stage SFT for 7b_trkd.Experiment blocked; see the specific reasonReported 49.9 pp
Reported
49.9 pp
Observed
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · Not assessed
Reported: 49.9 pp
downstream_accuracy · Not assessed
Reported: 31.1 pp
downstream_accuracy · Not assessed
Reported: 26.9 pp
downstream_accuracy · Not assessed
Reported: 31.5 pp
downstream_accuracy · Not assessed
Reported: 37.1 pp
downstream_accuracy · Not assessed
Reported: 7.6 pp
downstream_accuracy · Not assessed
Reported: 53.7 pp
downstream_accuracy · Not assessed
Reported: 23.1 pp
downstream_accuracy · Not assessed
Reported: 8.4 pp
downstream_accuracy · Not assessed
Reported: 46.5 pp
downstream_accuracy · Not assessed
Reported: 17.4 pp
downstream_accuracy · Not assessed
Reported: 57.8 pp
downstream_accuracy · Not assessed
Reported: 56 pp
downstream_accuracy · Not assessed
Reported: 51.4 pp
downstream_accuracy · Not assessed
Reported: 37.5 pp
downstream_accuracy · Not assessed
Reported: 45.3 pp
Mid-training ablation results for Teacher Top-1 Labels relative difference across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported -6.4 pp
Reported
-6.4 pp
Observed
—
Ablation checkpoints not released and training exceeds compute limits.
delta_percentage_points · Not assessed
Reported: -6.4 pp
delta_percentage_points · Not assessed
Reported: 1.3 pp
delta_percentage_points · Not assessed
Reported: -2.8 pp
Downstream task evaluation after post-training stage SFT for 13b_fkd.Experiment blocked; see the specific reasonReported 54.1 pp
Reported
54.1 pp
Observed
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · Not assessed
Reported: 54.1 pp
downstream_accuracy · Not assessed
Reported: 32.5 pp
downstream_accuracy · Not assessed
Reported: 28 pp
downstream_accuracy · Not assessed
Reported: 31.8 pp
downstream_accuracy · Not assessed
Reported: 36.3 pp
downstream_accuracy · Not assessed
Reported: 7.6 pp
downstream_accuracy · Not assessed
Reported: 54.9 pp
downstream_accuracy · Not assessed
Reported: 23.5 pp
downstream_accuracy · Not assessed
Reported: 7.9 pp
downstream_accuracy · Not assessed
Reported: 45.9 pp
downstream_accuracy · Not assessed
Reported: 17.1 pp
downstream_accuracy · Not assessed
Reported: 55.6 pp
downstream_accuracy · Not assessed
Reported: 57.2 pp
downstream_accuracy · Not assessed
Reported: 51.2 pp
downstream_accuracy · Not assessed
Reported: 37.1 pp
downstream_accuracy · Not assessed
Reported: 44.2 pp
Per-task downstream accuracy for mid-training ablation oracle_domain using OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 61 pp
Reported
61 pp
Observed
—
Per-task ablation checkpoints not released.
downstream_accuracy · Not assessed
Reported: 61 pp
downstream_accuracy · Not assessed
Reported: 45.6 pp
downstream_accuracy · Not assessed
Reported: 38.5 pp
downstream_accuracy · Not assessed
Reported: 31.6 pp
downstream_accuracy · Not assessed
Reported: 40.1 pp
downstream_accuracy · Not assessed
Reported: 8 pp
downstream_accuracy · Not assessed
Reported: 53 pp
downstream_accuracy · Not assessed
Reported: 22.6 pp
downstream_accuracy · Not assessed
Reported: 8.3 pp
downstream_accuracy · Not assessed
Reported: 49.4 pp
downstream_accuracy · Not assessed
Reported: 18.8 pp
downstream_accuracy · Not assessed
Reported: 61.1 pp
downstream_accuracy · Not assessed
Reported: 61.4 pp
downstream_accuracy · Not assessed
Reported: 52.6 pp
downstream_accuracy · Not assessed
Reported: 38.9 pp
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B FKD post-training performance.Experiment blocked; see the specific reasonReported 72.1 pp
Reported
72.1 pp
Observed
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · Not assessed
Reported: 72.1 pp
downstream_accuracy · Not assessed
Reported: 54.3 pp
downstream_accuracy · Not assessed
Reported: 46.5 pp
downstream_accuracy · Not assessed
Reported: 31.4 pp
downstream_accuracy · Not assessed
Reported: 35.7 pp
downstream_accuracy · Not assessed
Reported: 14.6 pp
downstream_accuracy · Not assessed
Reported: 54.1 pp
downstream_accuracy · Not assessed
Reported: 23.2 pp
downstream_accuracy · Not assessed
Reported: 8 pp
downstream_accuracy · Not assessed
Reported: 36.1 pp
downstream_accuracy · Not assessed
Reported: 18.1 pp
downstream_accuracy · Not assessed
Reported: 55 pp
downstream_accuracy · Not assessed
Reported: 56.4 pp
downstream_accuracy · Not assessed
Reported: 51.2 pp
downstream_accuracy · Not assessed
Reported: 35.9 pp
downstream_accuracy · Not assessed
Reported: 64.9 pp
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q10.Experiment blocked; see the specific reasonReported 61.6 pp
Reported
61.6 pp
Observed
—
Threshold sweep checkpoints not released.
downstream_accuracy · Not assessed
Reported: 61.6 pp
downstream_accuracy · Not assessed
Reported: 48.8 pp
downstream_accuracy · Not assessed
Reported: 39.4 pp
downstream_accuracy · Not assessed
Reported: 32.5 pp
downstream_accuracy · Not assessed
Reported: 50.2 pp
downstream_accuracy · Not assessed
Reported: 11.4 pp
downstream_accuracy · Not assessed
Reported: 40.6 pp
downstream_accuracy · Not assessed
Reported: 53.7 pp
downstream_accuracy · Not assessed
Reported: 23.9 pp
downstream_accuracy · Not assessed
Reported: 8.6 pp
downstream_accuracy · Not assessed
Reported: 28.7 pp
downstream_accuracy · Not assessed
Reported: 50.6 pp
downstream_accuracy · Not assessed
Reported: 19 pp
downstream_accuracy · Not assessed
Reported: 62 pp
downstream_accuracy · Not assessed
Reported: 63.8 pp
downstream_accuracy · Not assessed
Reported: 55.5 pp
downstream_accuracy · Not assessed
Reported: 40.4 pp
downstream_accuracy · Not assessed
Reported: 48.5 pp
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B TRKD post-training performance.Experiment blocked; see the specific reasonReported 69.6 pp
Reported
69.6 pp
Observed
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · Not assessed
Reported: 69.6 pp
downstream_accuracy · Not assessed
Reported: 48.5 pp
downstream_accuracy · Not assessed
Reported: 42.6 pp
downstream_accuracy · Not assessed
Reported: 31.5 pp
downstream_accuracy · Not assessed
Reported: 34.6 pp
downstream_accuracy · Not assessed
Reported: 12.2 pp
downstream_accuracy · Not assessed
Reported: 53.9 pp
downstream_accuracy · Not assessed
Reported: 22.2 pp
downstream_accuracy · Not assessed
Reported: 8.2 pp
downstream_accuracy · Not assessed
Reported: 37 pp
downstream_accuracy · Not assessed
Reported: 17.3 pp
downstream_accuracy · Not assessed
Reported: 55.4 pp
downstream_accuracy · Not assessed
Reported: 54 pp
downstream_accuracy · Not assessed
Reported: 51.3 pp
downstream_accuracy · Not assessed
Reported: 33.6 pp
downstream_accuracy · Not assessed
Reported: 64.1 pp
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B RKD post-training performance.Experiment blocked; see the specific reasonReported 73.8 pp
Reported
73.8 pp
Observed
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · Not assessed
Reported: 73.8 pp
downstream_accuracy · Not assessed
Reported: 53.4 pp
downstream_accuracy · Not assessed
Reported: 49.5 pp
downstream_accuracy · Not assessed
Reported: 31.6 pp
downstream_accuracy · Not assessed
Reported: 40.5 pp
downstream_accuracy · Not assessed
Reported: 17.6 pp
downstream_accuracy · Not assessed
Reported: 51.8 pp
downstream_accuracy · Not assessed
Reported: 21.9 pp
downstream_accuracy · Not assessed
Reported: 7.9 pp
downstream_accuracy · Not assessed
Reported: 46.2 pp
downstream_accuracy · Not assessed
Reported: 18 pp
downstream_accuracy · Not assessed
Reported: 61.1 pp
downstream_accuracy · Not assessed
Reported: 53 pp
downstream_accuracy · Not assessed
Reported: 51.2 pp
downstream_accuracy · Not assessed
Reported: 37.1 pp
downstream_accuracy · Not assessed
Reported: 61.6 pp
Downstream task evaluation after post-training stage RLVR1 for 7b_fkd.Experiment blocked; see the specific reasonReported 76.7 pp
Reported
76.7 pp
Observed
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · Not assessed
Reported: 76.7 pp
downstream_accuracy · Not assessed
Reported: 59.4 pp
downstream_accuracy · Not assessed
Reported: 51.9 pp
downstream_accuracy · Not assessed
Reported: 34.4 pp
downstream_accuracy · Not assessed
Reported: 39.2 pp
downstream_accuracy · Not assessed
Reported: 16.8 pp
downstream_accuracy · Not assessed
Reported: 52.8 pp
downstream_accuracy · Not assessed
Reported: 23.1 pp
downstream_accuracy · Not assessed
Reported: 8.3 pp
downstream_accuracy · Not assessed
Reported: 41.4 pp
downstream_accuracy · Not assessed
Reported: 18.3 pp
downstream_accuracy · Not assessed
Reported: 59.4 pp
downstream_accuracy · Not assessed
Reported: 56.2 pp
downstream_accuracy · Not assessed
Reported: 51.2 pp
downstream_accuracy · Not assessed
Reported: 38 pp
downstream_accuracy · Not assessed
Reported: 67.1 pp
Downstream task evaluation after post-training stage DPO for 13b_fkd.Experiment blocked; see the specific reasonReported 61 pp
Reported
61 pp
Observed
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · Not assessed
Reported: 61 pp
downstream_accuracy · Not assessed
Reported: 34.2 pp
downstream_accuracy · Not assessed
Reported: 31.7 pp
downstream_accuracy · Not assessed
Reported: 33.4 pp
downstream_accuracy · Not assessed
Reported: 36.9 pp
downstream_accuracy · Not assessed
Reported: 8.2 pp
downstream_accuracy · Not assessed
Reported: 54.6 pp
downstream_accuracy · Not assessed
Reported: 23.2 pp
downstream_accuracy · Not assessed
Reported: 7.7 pp
downstream_accuracy · Not assessed
Reported: 45.9 pp
downstream_accuracy · Not assessed
Reported: 17.4 pp
downstream_accuracy · Not assessed
Reported: 53.2 pp
downstream_accuracy · Not assessed
Reported: 57.2 pp
downstream_accuracy · Not assessed
Reported: 52.3 pp
downstream_accuracy · Not assessed
Reported: 37.4 pp
downstream_accuracy · Not assessed
Reported: 64.7 pp
Switch Distillation (SD, q=20%, tau=2) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher, demonstrating substantial gains on reasoning while maintaining factual recall.Experiment blocked; see the specific reasonReported 69.7 pp
Reported
69.7 pp
Observed
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · Not assessed
Reported: 69.7 pp
downstream_accuracy · Not assessed
Reported: 55.3 pp
downstream_accuracy · Not assessed
Reported: 46.1 pp
downstream_accuracy · Not assessed
Reported: 32.8 pp
downstream_accuracy · Not assessed
Reported: 49.6 pp
downstream_accuracy · Not assessed
Reported: 14.8 pp
downstream_accuracy · Not assessed
Reported: 54.9 pp
downstream_accuracy · Not assessed
Reported: 24.6 pp
downstream_accuracy · Not assessed
Reported: 8.4 pp
downstream_accuracy · Not assessed
Reported: 51.6 pp
downstream_accuracy · Not assessed
Reported: 19.8 pp
downstream_accuracy · Not assessed
Reported: 64.7 pp
downstream_accuracy · Not assessed
Reported: 64.2 pp
downstream_accuracy · Not assessed
Reported: 53.8 pp
downstream_accuracy · Not assessed
Reported: 41.5 pp
Downstream task evaluation after post-training stage DPO for 13b_rkd.Experiment blocked; see the specific reasonReported 65.6 pp
Reported
65.6 pp
Observed
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · Not assessed
Reported: 65.6 pp
downstream_accuracy · Not assessed
Reported: 41.2 pp
downstream_accuracy · Not assessed
Reported: 35 pp
downstream_accuracy · Not assessed
Reported: 33.2 pp
downstream_accuracy · Not assessed
Reported: 39.9 pp
downstream_accuracy · Not assessed
Reported: 7.4 pp
downstream_accuracy · Not assessed
Reported: 53.5 pp
downstream_accuracy · Not assessed
Reported: 23.7 pp
downstream_accuracy · Not assessed
Reported: 7.9 pp
downstream_accuracy · Not assessed
Reported: 47.8 pp
downstream_accuracy · Not assessed
Reported: 17.9 pp
downstream_accuracy · Not assessed
Reported: 57.1 pp
downstream_accuracy · Not assessed
Reported: 56.2 pp
downstream_accuracy · Not assessed
Reported: 51.9 pp
downstream_accuracy · Not assessed
Reported: 38.5 pp
downstream_accuracy · Not assessed
Reported: 64.5 pp
Token-Routing Knowledge Distillation (TRKD) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 52.5 pp
Reported
52.5 pp
Observed
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · Not assessed
Reported: 52.5 pp
downstream_accuracy · Not assessed
Reported: 36.9 pp
downstream_accuracy · Not assessed
Reported: 31 pp
downstream_accuracy · Not assessed
Reported: 30.3 pp
downstream_accuracy · Not assessed
Reported: 36.8 pp
downstream_accuracy · Not assessed
Reported: 5.6 pp
downstream_accuracy · Not assessed
Reported: 54 pp
downstream_accuracy · Not assessed
Reported: 24.1 pp
downstream_accuracy · Not assessed
Reported: 7.8 pp
downstream_accuracy · Not assessed
Reported: 47.9 pp
downstream_accuracy · Not assessed
Reported: 17.4 pp
downstream_accuracy · Not assessed
Reported: 58.9 pp
downstream_accuracy · Not assessed
Reported: 60 pp
downstream_accuracy · Not assessed
Reported: 51.2 pp
downstream_accuracy · Not assessed
Reported: 37.5 pp
Forward KL Distillation (FKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.Experiment blocked; see the specific reasonReported 52.7 pp
Reported
52.7 pp
Observed
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · Not assessed
Reported: 52.7 pp
downstream_accuracy · Not assessed
Reported: 37.9 pp
downstream_accuracy · Not assessed
Reported: 31.8 pp
downstream_accuracy · Not assessed
Reported: 29.4 pp
downstream_accuracy · Not assessed
Reported: 33.1 pp
downstream_accuracy · Not assessed
Reported: 6.4 pp
downstream_accuracy · Not assessed
Reported: 54.2 pp
downstream_accuracy · Not assessed
Reported: 24.7 pp
downstream_accuracy · Not assessed
Reported: 8.1 pp
downstream_accuracy · Not assessed
Reported: 47.5 pp
downstream_accuracy · Not assessed
Reported: 16.3 pp
downstream_accuracy · Not assessed
Reported: 55.3 pp
downstream_accuracy · Not assessed
Reported: 56.2 pp
downstream_accuracy · Not assessed
Reported: 51.2 pp
downstream_accuracy · Not assessed
Reported: 36.2 pp
Standard Next-Token Prediction (NTP) baseline downstream evaluation performance after 60B mid-training tokens across 15 tasks covering Reasoning, Factual Recall, and Knowledge & Commonsense.Experiment blocked; see the specific reasonReported 40.4 pp
Reported
40.4 pp
Observed
—
Official model checkpoints from mid-training are not publicly released, and mid-training 1B students from 4T tokens requires compute far exceeding the resource budget.
downstream_accuracy · Not assessed
Reported: 40.4 pp
downstream_accuracy · Not assessed
Reported: 29.8 pp
downstream_accuracy · Not assessed
Reported: 23.1 pp
downstream_accuracy · Not assessed
Reported: 29.9 pp
downstream_accuracy · Not assessed
Reported: 29.8 pp
downstream_accuracy · Not assessed
Reported: 3.8 pp
downstream_accuracy · Not assessed
Reported: 56.7 pp
downstream_accuracy · Not assessed
Reported: 25.5 pp
downstream_accuracy · Not assessed
Reported: 8.7 pp
downstream_accuracy · Not assessed
Reported: 43.6 pp
downstream_accuracy · Not assessed
Reported: 15.5 pp
downstream_accuracy · Not assessed
Reported: 51.1 pp
downstream_accuracy · Not assessed
Reported: 51.8 pp
downstream_accuracy · Not assessed
Reported: 51.4 pp
downstream_accuracy · Not assessed
Reported: 34.1 pp
Token-Routing Knowledge Distillation (TRKD) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.Experiment blocked; see the specific reasonReported 47.8 pp
Reported
47.8 pp
Observed
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · Not assessed
Reported: 47.8 pp
downstream_accuracy · Not assessed
Reported: 32.7 pp
downstream_accuracy · Not assessed
Reported: 28.4 pp
downstream_accuracy · Not assessed
Reported: 29.3 pp
downstream_accuracy · Not assessed
Reported: 34 pp
downstream_accuracy · Not assessed
Reported: 5.6 pp
downstream_accuracy · Not assessed
Reported: 53.8 pp
downstream_accuracy · Not assessed
Reported: 24.2 pp
downstream_accuracy · Not assessed
Reported: 8.3 pp
downstream_accuracy · Not assessed
Reported: 45.7 pp
downstream_accuracy · Not assessed
Reported: 15.4 pp
downstream_accuracy · Not assessed
Reported: 53 pp
downstream_accuracy · Not assessed
Reported: 56.6 pp
downstream_accuracy · Not assessed
Reported: 50.7 pp
downstream_accuracy · Not assessed
Reported: 35.2 pp
Downstream task evaluation after post-training stage DPO for 7b_rkd.Experiment blocked; see the specific reasonReported 64.5 pp
Reported
64.5 pp
Observed
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · Not assessed
Reported: 64.5 pp
downstream_accuracy · Not assessed
Reported: 40 pp
downstream_accuracy · Not assessed
Reported: 36.4 pp
downstream_accuracy · Not assessed
Reported: 34.2 pp
downstream_accuracy · Not assessed
Reported: 40.5 pp
downstream_accuracy · Not assessed
Reported: 10 pp
downstream_accuracy · Not assessed
Reported: 52.8 pp
downstream_accuracy · Not assessed
Reported: 22.3 pp
downstream_accuracy · Not assessed
Reported: 7.8 pp
downstream_accuracy · Not assessed
Reported: 48.8 pp
downstream_accuracy · Not assessed
Reported: 17.3 pp
downstream_accuracy · Not assessed
Reported: 57.4 pp
downstream_accuracy · Not assessed
Reported: 55 pp
downstream_accuracy · Not assessed
Reported: 52.5 pp
downstream_accuracy · Not assessed
Reported: 38.2 pp
downstream_accuracy · Not assessed
Reported: 64.1 pp
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B Switch Distillation post-training performance.Experiment blocked; see the specific reasonReported 77.8 pp
Reported
77.8 pp
Observed
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · Not assessed
Reported: 77.8 pp
downstream_accuracy · Not assessed
Reported: 62.8 pp
downstream_accuracy · Not assessed
Reported: 51.8 pp
downstream_accuracy · Not assessed
Reported: 33.3 pp
downstream_accuracy · Not assessed
Reported: 42.8 pp
downstream_accuracy · Not assessed
Reported: 19.6 pp
downstream_accuracy · Not assessed
Reported: 54.2 pp
downstream_accuracy · Not assessed
Reported: 24.1 pp
downstream_accuracy · Not assessed
Reported: 8.4 pp
downstream_accuracy · Not assessed
Reported: 43.1 pp
downstream_accuracy · Not assessed
Reported: 18.1 pp
downstream_accuracy · Not assessed
Reported: 56.5 pp
downstream_accuracy · Not assessed
Reported: 59 pp
downstream_accuracy · Not assessed
Reported: 52.5 pp
downstream_accuracy · Not assessed
Reported: 38.1 pp
downstream_accuracy · Not assessed
Reported: 67.1 pp
Downstream task evaluation after post-training stage DPO for 7b_sd.Experiment blocked; see the specific reasonReported 70.9 pp
Reported
70.9 pp
Observed
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · Not assessed
Reported: 70.9 pp
downstream_accuracy · Not assessed
Reported: 50.8 pp
downstream_accuracy · Not assessed
Reported: 44.2 pp
downstream_accuracy · Not assessed
Reported: 35.5 pp
downstream_accuracy · Not assessed
Reported: 49 pp
downstream_accuracy · Not assessed
Reported: 15 pp
downstream_accuracy · Not assessed
Reported: 54.7 pp
downstream_accuracy · Not assessed
Reported: 23.6 pp
downstream_accuracy · Not assessed
Reported: 8.1 pp
downstream_accuracy · Not assessed
Reported: 50 pp
downstream_accuracy · Not assessed
Reported: 20.1 pp
downstream_accuracy · Not assessed
Reported: 61.2 pp
downstream_accuracy · Not assessed
Reported: 61 pp
downstream_accuracy · Not assessed
Reported: 55.5 pp
downstream_accuracy · Not assessed
Reported: 40.9 pp
downstream_accuracy · Not assessed
Reported: 62.1 pp
Downstream task evaluation after post-training stage SFT for 13b_trkd.Experiment blocked; see the specific reasonReported 49 pp
Reported
49 pp
Observed
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · Not assessed
Reported: 49 pp
downstream_accuracy · Not assessed
Reported: 29.8 pp
downstream_accuracy · Not assessed
Reported: 25.4 pp
downstream_accuracy · Not assessed
Reported: 31.6 pp
downstream_accuracy · Not assessed
Reported: 35.2 pp
downstream_accuracy · Not assessed
Reported: 7.6 pp
downstream_accuracy · Not assessed
Reported: 55.4 pp
downstream_accuracy · Not assessed
Reported: 23.6 pp
downstream_accuracy · Not assessed
Reported: 8.8 pp
downstream_accuracy · Not assessed
Reported: 43.8 pp
downstream_accuracy · Not assessed
Reported: 16.7 pp
downstream_accuracy · Not assessed
Reported: 53.8 pp
downstream_accuracy · Not assessed
Reported: 55 pp
downstream_accuracy · Not assessed
Reported: 51.3 pp
downstream_accuracy · Not assessed
Reported: 36.2 pp
downstream_accuracy · Not assessed
Reported: 43.4 pp
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q30.Experiment blocked; see the specific reasonReported 66 pp
Reported
66 pp
Observed
—
Threshold sweep checkpoints not released.
downstream_accuracy · Not assessed
Reported: 66 pp
downstream_accuracy · Not assessed
Reported: 53.1 pp
downstream_accuracy · Not assessed
Reported: 42.9 pp
downstream_accuracy · Not assessed
Reported: 32.5 pp
downstream_accuracy · Not assessed
Reported: 44.4 pp
downstream_accuracy · Not assessed
Reported: 11 pp
downstream_accuracy · Not assessed
Reported: 41.6 pp
downstream_accuracy · Not assessed
Reported: 54.2 pp
downstream_accuracy · Not assessed
Reported: 24.7 pp
downstream_accuracy · Not assessed
Reported: 8.6 pp
downstream_accuracy · Not assessed
Reported: 29.2 pp
downstream_accuracy · Not assessed
Reported: 49 pp
downstream_accuracy · Not assessed
Reported: 18.4 pp
downstream_accuracy · Not assessed
Reported: 60.9 pp
downstream_accuracy · Not assessed
Reported: 63.6 pp
downstream_accuracy · Not assessed
Reported: 52.9 pp
downstream_accuracy · Not assessed
Reported: 38.7 pp
downstream_accuracy · Not assessed
Reported: 47.2 pp
Downstream task evaluation after post-training stage RLVR1 for 7b_sd.Experiment blocked; see the specific reasonReported 79.3 pp
Reported
79.3 pp
Observed
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · Not assessed
Reported: 79.3 pp
downstream_accuracy · Not assessed
Reported: 63.4 pp
downstream_accuracy · Not assessed
Reported: 52.8 pp
downstream_accuracy · Not assessed
Reported: 36.1 pp
downstream_accuracy · Not assessed
Reported: 48.8 pp
downstream_accuracy · Not assessed
Reported: 17.4 pp
downstream_accuracy · Not assessed
Reported: 54.1 pp
downstream_accuracy · Not assessed
Reported: 23 pp
downstream_accuracy · Not assessed
Reported: 8 pp
downstream_accuracy · Not assessed
Reported: 47.2 pp
downstream_accuracy · Not assessed
Reported: 19.7 pp
downstream_accuracy · Not assessed
Reported: 61.7 pp
downstream_accuracy · Not assessed
Reported: 59.8 pp
downstream_accuracy · Not assessed
Reported: 53 pp
downstream_accuracy · Not assessed
Reported: 39.7 pp
downstream_accuracy · Not assessed
Reported: 68.8 pp
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q20.Experiment blocked; see the specific reasonReported 66 pp
Reported
66 pp
Observed
—
Threshold sweep checkpoints not released.
downstream_accuracy · Not assessed
Reported: 66 pp
downstream_accuracy · Not assessed
Reported: 54.8 pp
downstream_accuracy · Not assessed
Reported: 42.2 pp
downstream_accuracy · Not assessed
Reported: 31.6 pp
downstream_accuracy · Not assessed
Reported: 45.8 pp
downstream_accuracy · Not assessed
Reported: 12 pp
downstream_accuracy · Not assessed
Reported: 42.1 pp
downstream_accuracy · Not assessed
Reported: 54.1 pp
downstream_accuracy · Not assessed
Reported: 24.9 pp
downstream_accuracy · Not assessed
Reported: 9 pp
downstream_accuracy · Not assessed
Reported: 29.3 pp
downstream_accuracy · Not assessed
Reported: 48.5 pp
downstream_accuracy · Not assessed
Reported: 17.7 pp
downstream_accuracy · Not assessed
Reported: 58.5 pp
downstream_accuracy · Not assessed
Reported: 63.8 pp
downstream_accuracy · Not assessed
Reported: 51.2 pp
downstream_accuracy · Not assessed
Reported: 39.2 pp
downstream_accuracy · Not assessed
Reported: 46.5 pp
Mid-training ablation results for Switch Distillation with FKL relative difference across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported -2.9 pp
Reported
-2.9 pp
Observed
—
Ablation checkpoints not released and training exceeds compute limits.
delta_percentage_points · Not assessed
Reported: -2.9 pp
delta_percentage_points · Not assessed
Reported: -0.2 pp
delta_percentage_points · Not assessed
Reported: -1.4 pp
Downstream task evaluation after post-training stage DPO for 13b_trkd.Experiment blocked; see the specific reasonReported 53.1 pp
Reported
53.1 pp
Observed
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · Not assessed
Reported: 53.1 pp
downstream_accuracy · Not assessed
Reported: 34.9 pp
downstream_accuracy · Not assessed
Reported: 30.7 pp
downstream_accuracy · Not assessed
Reported: 32.9 pp
downstream_accuracy · Not assessed
Reported: 35.3 pp
downstream_accuracy · Not assessed
Reported: 6.8 pp
downstream_accuracy · Not assessed
Reported: 54.9 pp
downstream_accuracy · Not assessed
Reported: 22.4 pp
downstream_accuracy · Not assessed
Reported: 8.5 pp
downstream_accuracy · Not assessed
Reported: 43.4 pp
downstream_accuracy · Not assessed
Reported: 16.5 pp
downstream_accuracy · Not assessed
Reported: 53.2 pp
downstream_accuracy · Not assessed
Reported: 56 pp
downstream_accuracy · Not assessed
Reported: 52 pp
downstream_accuracy · Not assessed
Reported: 35.3 pp
downstream_accuracy · Not assessed
Reported: 61.7 pp
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B Switch Distillation post-training performance.Experiment blocked; see the specific reasonReported 79.8 pp
Reported
79.8 pp
Observed
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · Not assessed
Reported: 79.8 pp
downstream_accuracy · Not assessed
Reported: 65 pp
downstream_accuracy · Not assessed
Reported: 52.7 pp
downstream_accuracy · Not assessed
Reported: 35.6 pp
downstream_accuracy · Not assessed
Reported: 48.2 pp
downstream_accuracy · Not assessed
Reported: 22.4 pp
downstream_accuracy · Not assessed
Reported: 53.6 pp
downstream_accuracy · Not assessed
Reported: 22.7 pp
downstream_accuracy · Not assessed
Reported: 8.2 pp
downstream_accuracy · Not assessed
Reported: 48.1 pp
downstream_accuracy · Not assessed
Reported: 19.8 pp
downstream_accuracy · Not assessed
Reported: 62.6 pp
downstream_accuracy · Not assessed
Reported: 59.4 pp
downstream_accuracy · Not assessed
Reported: 54.6 pp
downstream_accuracy · Not assessed
Reported: 39.6 pp
downstream_accuracy · Not assessed
Reported: 69.5 pp
Reverse KL Distillation (RKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 62.1 pp
Reported
62.1 pp
Observed
—
Mid-trained checkpoint weights not released and mid-training exceeds compute limits.
downstream_accuracy · Not assessed
Reported: 62.1 pp
downstream_accuracy · Not assessed
Reported: 46.2 pp
downstream_accuracy · Not assessed
Reported: 39.2 pp
downstream_accuracy · Not assessed
Reported: 31.6 pp
downstream_accuracy · Not assessed
Reported: 43.1 pp
downstream_accuracy · Not assessed
Reported: 10.4 pp
downstream_accuracy · Not assessed
Reported: 53.1 pp
downstream_accuracy · Not assessed
Reported: 24.1 pp
downstream_accuracy · Not assessed
Reported: 8.4 pp
downstream_accuracy · Not assessed
Reported: 50.1 pp
downstream_accuracy · Not assessed
Reported: 18.9 pp
downstream_accuracy · Not assessed
Reported: 62.5 pp
downstream_accuracy · Not assessed
Reported: 62.8 pp
downstream_accuracy · Not assessed
Reported: 52.5 pp
downstream_accuracy · Not assessed
Reported: 39.3 pp
Downstream task evaluation after post-training stage SFT for 7b_fkd.Experiment blocked; see the specific reasonReported 54.7 pp
Reported
54.7 pp
Observed
—
Intermediate checkpoint weights after post-training stage SFT not released.
downstream_accuracy · Not assessed
Reported: 54.7 pp
downstream_accuracy · Not assessed
Reported: 34.2 pp
downstream_accuracy · Not assessed
Reported: 30.7 pp
downstream_accuracy · Not assessed
Reported: 31.3 pp
downstream_accuracy · Not assessed
Reported: 39.1 pp
downstream_accuracy · Not assessed
Reported: 9.4 pp
downstream_accuracy · Not assessed
Reported: 54.1 pp
downstream_accuracy · Not assessed
Reported: 23.3 pp
downstream_accuracy · Not assessed
Reported: 8.2 pp
downstream_accuracy · Not assessed
Reported: 49.1 pp
downstream_accuracy · Not assessed
Reported: 18.5 pp
downstream_accuracy · Not assessed
Reported: 61.7 pp
downstream_accuracy · Not assessed
Reported: 59.4 pp
downstream_accuracy · Not assessed
Reported: 51.9 pp
downstream_accuracy · Not assessed
Reported: 39.7 pp
downstream_accuracy · Not assessed
Reported: 46 pp
Downstream task evaluation after post-training stage RLVR1 for 13b_sd.Experiment blocked; see the specific reasonReported 78.8 pp
Reported
78.8 pp
Observed
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · Not assessed
Reported: 78.8 pp
downstream_accuracy · Not assessed
Reported: 65.6 pp
downstream_accuracy · Not assessed
Reported: 52.5 pp
downstream_accuracy · Not assessed
Reported: 33.5 pp
downstream_accuracy · Not assessed
Reported: 43.3 pp
downstream_accuracy · Not assessed
Reported: 17.8 pp
downstream_accuracy · Not assessed
Reported: 54.3 pp
downstream_accuracy · Not assessed
Reported: 24 pp
downstream_accuracy · Not assessed
Reported: 8.2 pp
downstream_accuracy · Not assessed
Reported: 38.7 pp
downstream_accuracy · Not assessed
Reported: 17.7 pp
downstream_accuracy · Not assessed
Reported: 56.2 pp
downstream_accuracy · Not assessed
Reported: 57.8 pp
downstream_accuracy · Not assessed
Reported: 51.7 pp
downstream_accuracy · Not assessed
Reported: 37.7 pp
downstream_accuracy · Not assessed
Reported: 68.9 pp
Per-task downstream accuracy for mid-training ablation sd_fkl using OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 66.7 pp
Reported
66.7 pp
Observed
—
Per-task ablation checkpoints not released.
downstream_accuracy · Not assessed
Reported: 66.7 pp
downstream_accuracy · Not assessed
Reported: 52 pp
downstream_accuracy · Not assessed
Reported: 43.7 pp
downstream_accuracy · Not assessed
Reported: 31.2 pp
downstream_accuracy · Not assessed
Reported: 46.4 pp
downstream_accuracy · Not assessed
Reported: 11 pp
downstream_accuracy · Not assessed
Reported: 54.8 pp
downstream_accuracy · Not assessed
Reported: 24.3 pp
downstream_accuracy · Not assessed
Reported: 8.2 pp
downstream_accuracy · Not assessed
Reported: 50.5 pp
downstream_accuracy · Not assessed
Reported: 19.1 pp
downstream_accuracy · Not assessed
Reported: 63.7 pp
downstream_accuracy · Not assessed
Reported: 62.6 pp
downstream_accuracy · Not assessed
Reported: 52.2 pp
downstream_accuracy · Not assessed
Reported: 39.6 pp
Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B RKD post-training performance.Experiment blocked; see the specific reasonReported 75.1 pp
Reported
75.1 pp
Observed
—
Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.
downstream_accuracy · Not assessed
Reported: 75.1 pp
downstream_accuracy · Not assessed
Reported: 58.1 pp
downstream_accuracy · Not assessed
Reported: 49.3 pp
downstream_accuracy · Not assessed
Reported: 33.3 pp
downstream_accuracy · Not assessed
Reported: 38.7 pp
downstream_accuracy · Not assessed
Reported: 16 pp
downstream_accuracy · Not assessed
Reported: 53.9 pp
downstream_accuracy · Not assessed
Reported: 22.3 pp
downstream_accuracy · Not assessed
Reported: 7.7 pp
downstream_accuracy · Not assessed
Reported: 38.2 pp
downstream_accuracy · Not assessed
Reported: 17 pp
downstream_accuracy · Not assessed
Reported: 56 pp
downstream_accuracy · Not assessed
Reported: 55.2 pp
downstream_accuracy · Not assessed
Reported: 51.1 pp
downstream_accuracy · Not assessed
Reported: 36.6 pp
downstream_accuracy · Not assessed
Reported: 64.1 pp
Mid-training ablation results for Random Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported -6.5 pp
Reported
-6.5 pp
Observed
—
Ablation checkpoints not released and training exceeds compute limits.
delta_percentage_points · Not assessed
Reported: -6.5 pp
delta_percentage_points · Not assessed
Reported: -0.8 pp
delta_percentage_points · Not assessed
Reported: -2 pp
Downstream task evaluation after post-training stage RLVR1 for 7b_trkd.Experiment blocked; see the specific reasonReported 71 pp
Reported
71 pp
Observed
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · Not assessed
Reported: 71 pp
downstream_accuracy · Not assessed
Reported: 54.6 pp
downstream_accuracy · Not assessed
Reported: 46.8 pp
downstream_accuracy · Not assessed
Reported: 33 pp
downstream_accuracy · Not assessed
Reported: 37.4 pp
downstream_accuracy · Not assessed
Reported: 11.4 pp
downstream_accuracy · Not assessed
Reported: 52.8 pp
downstream_accuracy · Not assessed
Reported: 22.7 pp
downstream_accuracy · Not assessed
Reported: 8 pp
downstream_accuracy · Not assessed
Reported: 40.2 pp
downstream_accuracy · Not assessed
Reported: 17.8 pp
downstream_accuracy · Not assessed
Reported: 58.9 pp
downstream_accuracy · Not assessed
Reported: 55.6 pp
downstream_accuracy · Not assessed
Reported: 52 pp
downstream_accuracy · Not assessed
Reported: 35.9 pp
downstream_accuracy · Not assessed
Reported: 65.2 pp
Per-task downstream accuracy for mid-training ablation teacher_correct using OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 63.8 pp
Reported
63.8 pp
Observed
—
Per-task ablation checkpoints not released.
downstream_accuracy · Not assessed
Reported: 63.8 pp
downstream_accuracy · Not assessed
Reported: 49.5 pp
downstream_accuracy · Not assessed
Reported: 41.7 pp
downstream_accuracy · Not assessed
Reported: 31.8 pp
downstream_accuracy · Not assessed
Reported: 45.2 pp
downstream_accuracy · Not assessed
Reported: 9.8 pp
downstream_accuracy · Not assessed
Reported: 44.5 pp
downstream_accuracy · Not assessed
Reported: 19.4 pp
downstream_accuracy · Not assessed
Reported: 8.7 pp
downstream_accuracy · Not assessed
Reported: 50.6 pp
downstream_accuracy · Not assessed
Reported: 19.3 pp
downstream_accuracy · Not assessed
Reported: 63.2 pp
downstream_accuracy · Not assessed
Reported: 64.6 pp
downstream_accuracy · Not assessed
Reported: 52 pp
downstream_accuracy · Not assessed
Reported: 39.9 pp
Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q20.Experiment blocked; see the specific reasonReported 69.7 pp
Reported
69.7 pp
Observed
—
Threshold sweep checkpoints not released.
downstream_accuracy · Not assessed
Reported: 69.7 pp
downstream_accuracy · Not assessed
Reported: 55.3 pp
downstream_accuracy · Not assessed
Reported: 46.1 pp
downstream_accuracy · Not assessed
Reported: 32.8 pp
downstream_accuracy · Not assessed
Reported: 49.6 pp
downstream_accuracy · Not assessed
Reported: 14.8 pp
downstream_accuracy · Not assessed
Reported: 44.7 pp
downstream_accuracy · Not assessed
Reported: 54.9 pp
downstream_accuracy · Not assessed
Reported: 24.6 pp
downstream_accuracy · Not assessed
Reported: 8.4 pp
downstream_accuracy · Not assessed
Reported: 29.3 pp
downstream_accuracy · Not assessed
Reported: 51.6 pp
downstream_accuracy · Not assessed
Reported: 19.8 pp
downstream_accuracy · Not assessed
Reported: 64.7 pp
downstream_accuracy · Not assessed
Reported: 64.2 pp
downstream_accuracy · Not assessed
Reported: 53.8 pp
downstream_accuracy · Not assessed
Reported: 41.5 pp
downstream_accuracy · Not assessed
Reported: 49.3 pp
Mid-training ablation results for Teacher-Correct Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported -4.4 pp
Reported
-4.4 pp
Observed
—
Ablation checkpoints not released and training exceeds compute limits.
delta_percentage_points · Not assessed
Reported: -4.4 pp
delta_percentage_points · Not assessed
Reported: -5.1 pp
delta_percentage_points · Not assessed
Reported: -1 pp
Per-task downstream accuracy for mid-training ablation always_ce using OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 69.1 pp
Reported
69.1 pp
Observed
—
Per-task ablation checkpoints not released.
downstream_accuracy · Not assessed
Reported: 69.1 pp
downstream_accuracy · Not assessed
Reported: 55.2 pp
downstream_accuracy · Not assessed
Reported: 46.5 pp
downstream_accuracy · Not assessed
Reported: 33.3 pp
downstream_accuracy · Not assessed
Reported: 50.1 pp
downstream_accuracy · Not assessed
Reported: 12 pp
downstream_accuracy · Not assessed
Reported: 54.3 pp
downstream_accuracy · Not assessed
Reported: 24.5 pp
downstream_accuracy · Not assessed
Reported: 9.4 pp
downstream_accuracy · Not assessed
Reported: 51.1 pp
downstream_accuracy · Not assessed
Reported: 19.9 pp
downstream_accuracy · Not assessed
Reported: 63.7 pp
downstream_accuracy · Not assessed
Reported: 63.6 pp
downstream_accuracy · Not assessed
Reported: 52.7 pp
downstream_accuracy · Not assessed
Reported: 41.1 pp
Downstream task evaluation after post-training stage RLVR1 for 13b_trkd.Experiment blocked; see the specific reasonReported 72.4 pp
Reported
72.4 pp
Observed
—
Intermediate checkpoint weights after post-training stage RLVR1 not released.
downstream_accuracy · Not assessed
Reported: 72.4 pp
downstream_accuracy · Not assessed
Reported: 53.3 pp
downstream_accuracy · Not assessed
Reported: 47.6 pp
downstream_accuracy · Not assessed
Reported: 32.7 pp
downstream_accuracy · Not assessed
Reported: 35 pp
downstream_accuracy · Not assessed
Reported: 10.4 pp
downstream_accuracy · Not assessed
Reported: 54.7 pp
downstream_accuracy · Not assessed
Reported: 22.2 pp
downstream_accuracy · Not assessed
Reported: 8.6 pp
downstream_accuracy · Not assessed
Reported: 36.6 pp
downstream_accuracy · Not assessed
Reported: 16.6 pp
downstream_accuracy · Not assessed
Reported: 54.1 pp
downstream_accuracy · Not assessed
Reported: 54.6 pp
downstream_accuracy · Not assessed
Reported: 51.7 pp
downstream_accuracy · Not assessed
Reported: 34.5 pp
downstream_accuracy · Not assessed
Reported: 67.1 pp
Downstream task evaluation after post-training stage DPO for 7b_fkd.Experiment blocked; see the specific reasonReported 60.4 pp
Reported
60.4 pp
Observed
—
Intermediate checkpoint weights after post-training stage DPO not released.
downstream_accuracy · Not assessed
Reported: 60.4 pp
downstream_accuracy · Not assessed
Reported: 35.8 pp
downstream_accuracy · Not assessed
Reported: 33.1 pp
downstream_accuracy · Not assessed
Reported: 33.1 pp
downstream_accuracy · Not assessed
Reported: 39.8 pp
downstream_accuracy · Not assessed
Reported: 7.8 pp
downstream_accuracy · Not assessed
Reported: 53.6 pp
downstream_accuracy · Not assessed
Reported: 23.1 pp
downstream_accuracy · Not assessed
Reported: 8.3 pp
downstream_accuracy · Not assessed
Reported: 49.2 pp
downstream_accuracy · Not assessed
Reported: 18.5 pp
downstream_accuracy · Not assessed
Reported: 59.2 pp
downstream_accuracy · Not assessed
Reported: 58 pp
downstream_accuracy · Not assessed
Reported: 51.6 pp
downstream_accuracy · Not assessed
Reported: 39.7 pp
downstream_accuracy · Not assessed
Reported: 64.1 pp
Qualitative findings
The stage-dependent distillation tradeoff generalizes to the SmolLM family (SmolLM2 360M student, 1.7B Instruct teacher, SmolLM3 Stage-3 data), where mid-training KD exhibits the reasoning-recall deficit and Switch Distillation mitigates it.Experiment blocked; see the specific reason
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Forward KL consistently exhibits higher CE-KL gradient cosine alignment than Reverse KL across pre-training and mid-training, with the gap widening at higher alpha and later training steps.Experiment blocked; see the specific reason
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Gold-token probability, gradient attenuation relative to NTP, and factual recall deficits under KD hold consistently for OLMo-2 1B and 13B Instruct teachers.Experiment blocked; see the specific reason
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Switch Distillation substantially accelerates reasoning acquisition during mid-training, surpassing the final 60B-token NTP reasoning macro-average within 2.5B tokens (24x fewer tokens), while standard FKD requires 5.0B tokens (12x fewer).Experiment blocked; see the specific reason
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Lower teacher predictive entropy corresponds to substantially higher teacher top-1 agreement with the ground-truth token across all data domains and teacher sizes.Experiment blocked; see the specific reason
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Higher teacher entropy corresponds to strictly lower gold-token probability, attenuates the gold-token gradient relative to NTP (reaching ~0.5x NTP for Q5 facts), and produces larger downstream factual recall deficits.Experiment blocked; see the specific reason
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Across distillation strengths alpha in [0.0, 1.0], KL directions (forward and reverse), and teacher sizes (1B, 7B, 13B), mid-training distillation traces a reasoning-recall frontier that falls below NTP, which Switch Distillation mitigates.Experiment blocked; see the specific reason
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.
Knowledge distillation generally improves reasoning, but its effect on factual recall changes across training stages: during pre-training, distillation improves both reasoning and factual recall over NTP, whereas during mid-training it improves reasoning at the expense of factual recall.Experiment blocked; see the specific reason
Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.