Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Public
Authors:Jacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li, Hanqing Zeng, Yinglong Xia, Benyu Zhang, Zhuokai Zhao, Qiang Zhang, Pang Wei Koh, Luke Zettlemoyer, Wen-tau Yih
More
Copy repository link

Research claims

0 / 76 claims verified

the rest still being verified

Key metricsPaper-reported → reproduced · click a row to expand

Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).Inconclusive4/12 measurements supported · 12 assessedReported 0.815 score0.8079 score

Reported

0.815 score

Observed

0.8079 score

Delta

−0.0071 score

The execution did not produce enough comparable measurements to determine whether the paper's claim is supported. · This does not refute the paper's claim; it means the available evidence can neither confirm nor refute it yet.

roc_auc · Supported

Reported: 0.815 score

Observed: 0.8079 score

Reported ROC AUC of 0.815 was observed at 0.8079 (relative diff 0.87%), showing strong predictive entropy separation on OLMo-2 1B Base.

roc_auc · Supported

Reported: 0.77 score

Observed: 0.7967 score

Reported ROC AUC of 0.770 was observed at 0.7967 (relative diff 3.47%), confirming high ROC AUC separation on OLMo-2 1B SFT.

roc_auc · Supported

Reported: 0.744 score

Observed: 0.7931 score

Reported ROC AUC of 0.744 was observed at 0.7931 (relative diff 6.60%), maintaining the substantive finding of strong discrimination on OLMo-2 1B DPO.

roc_auc · Supported

Reported: 0.748 score

Observed: 0.7973 score

Reported ROC AUC of 0.748 was observed at 0.7973 (relative diff 6.59%), confirming procedural entropy remains distinct on OLMo-2 1B Instruct.

roc_auc · Inconclusive

Reported: 0.816 score

Observed:

The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json

roc_auc · Inconclusive

Reported: 0.777 score

Observed:

The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json

roc_auc · Inconclusive

Reported: 0.76 score

Observed:

The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json

roc_auc · Inconclusive

Reported: 0.761 score

Observed:

The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json

roc_auc · Inconclusive

Reported: 0.826 score

Observed:

The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json

roc_auc · Inconclusive

Reported: 0.814 score

Observed:

The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json

roc_auc · Inconclusive

Reported: 0.809 score

Observed:

The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json

roc_auc · Inconclusive

Reported: 0.811 score

Observed:

The execution was partial during measurement because of metric_unavailable: Evaluated token-level predictive entropy and computed ROC AUC separating procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) on 240 Dolmino documents across 107,555 tokens for the OLMo-2 1B model suite across four training stages (Base, SFT, DPO, Instruct). The observed ROC AUC values were 0.8079 (Base), 0.7967 (SFT), 0.7931 (DPO), and 0.7973 (Instruct), confirming the paper's key finding that procedural tokens exhibit substantially lower predictive entropy than knowledge-intensive tokens across all training stages (ROC AUC ~0.80 vs 0.50 random chance). Evaluated the complete 1B model family across all four post-training stages (Base, SFT, DPO, Instruct) on the paper's 240 Dolmino document benchmark (107,555 tokens); the 7B and 13B model suites were omitted due to cumulative network bandwidth and runtime budget limits (>220 GB total checkpoint download size across remaining 8 models). Unresolved metric interpretation: fig8_roc_auc_7b_base: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_sft: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_dpo: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_7b_instruct: 7B checkpoint evaluation is constrained by total runtime budget and bandwidth limits; 12 checkpoints across 7B/13B total 221GB which exceeds execution budget.; fig8_roc_auc_13b_base: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_sft: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_dpo: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits.; fig8_roc_auc_13b_instruct: 13B checkpoint evaluation is constrained by total runtime budget, network bandwidth (153GB for 13B suite), and single GPU VRAM limits. Unresolved measurements: fig8_roc_auc_7b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_base.json; fig8_roc_auc_7b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_sft.json; fig8_roc_auc_7b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_dpo.json; fig8_roc_auc_7b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_7b_instruct.json; fig8_roc_auc_13b_base: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_base.json; fig8_roc_auc_13b_sft: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_sft.json; fig8_roc_auc_13b_dpo: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_dpo.json; fig8_roc_auc_13b_instruct: The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at results/fig8_fig8_roc_auc_13b_instruct.json

Open claim details
Mid-training ablation results for Always CE relative difference across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported -0.3 pp

Reported

-0.3 pp

Observed

Ablation checkpoints not released and training exceeds compute limits.

delta_percentage_points · Not assessed

Reported: -0.3 pp

delta_percentage_points · Not assessed

Reported: 0.1 pp

delta_percentage_points · Not assessed

Reported: -0.6 pp

Downstream task evaluation after post-training stage RLVR1 for 13b_rkd.Experiment blocked; see the specific reasonReported 75.6 pp

Reported

75.6 pp

Observed

Intermediate checkpoint weights after post-training stage RLVR1 not released.

downstream_accuracy · Not assessed

Reported: 75.6 pp

downstream_accuracy · Not assessed

Reported: 60.1 pp

downstream_accuracy · Not assessed

Reported: 50.3 pp

downstream_accuracy · Not assessed

Reported: 33.3 pp

downstream_accuracy · Not assessed

Reported: 38.1 pp

downstream_accuracy · Not assessed

Reported: 15.8 pp

downstream_accuracy · Not assessed

Reported: 53.3 pp

downstream_accuracy · Not assessed

Reported: 22.3 pp

downstream_accuracy · Not assessed

Reported: 8 pp

downstream_accuracy · Not assessed

Reported: 36.2 pp

downstream_accuracy · Not assessed

Reported: 16.9 pp

downstream_accuracy · Not assessed

Reported: 55.9 pp

downstream_accuracy · Not assessed

Reported: 55.2 pp

downstream_accuracy · Not assessed

Reported: 51.3 pp

downstream_accuracy · Not assessed

Reported: 36.4 pp

downstream_accuracy · Not assessed

Reported: 63.6 pp

Downstream task evaluation after post-training stage RLVR1 for 13b_fkd.Experiment blocked; see the specific reasonReported 73 pp

Reported

73 pp

Observed

Intermediate checkpoint weights after post-training stage RLVR1 not released.

downstream_accuracy · Not assessed

Reported: 73 pp

downstream_accuracy · Not assessed

Reported: 55.5 pp

downstream_accuracy · Not assessed

Reported: 47.3 pp

downstream_accuracy · Not assessed

Reported: 33.7 pp

downstream_accuracy · Not assessed

Reported: 36.1 pp

downstream_accuracy · Not assessed

Reported: 14 pp

downstream_accuracy · Not assessed

Reported: 54.7 pp

downstream_accuracy · Not assessed

Reported: 23.3 pp

downstream_accuracy · Not assessed

Reported: 8.3 pp

downstream_accuracy · Not assessed

Reported: 35 pp

downstream_accuracy · Not assessed

Reported: 17.2 pp

downstream_accuracy · Not assessed

Reported: 54 pp

downstream_accuracy · Not assessed

Reported: 56.6 pp

downstream_accuracy · Not assessed

Reported: 51.7 pp

downstream_accuracy · Not assessed

Reported: 35.4 pp

downstream_accuracy · Not assessed

Reported: 63.2 pp

Downstream task evaluation after post-training stage SFT for 13b_rkd.Experiment blocked; see the specific reasonReported 56.8 pp

Reported

56.8 pp

Observed

Intermediate checkpoint weights after post-training stage SFT not released.

downstream_accuracy · Not assessed

Reported: 56.8 pp

downstream_accuracy · Not assessed

Reported: 36.7 pp

downstream_accuracy · Not assessed

Reported: 31.3 pp

downstream_accuracy · Not assessed

Reported: 32.1 pp

downstream_accuracy · Not assessed

Reported: 38.9 pp

downstream_accuracy · Not assessed

Reported: 11.6 pp

downstream_accuracy · Not assessed

Reported: 55.4 pp

downstream_accuracy · Not assessed

Reported: 23.8 pp

downstream_accuracy · Not assessed

Reported: 8 pp

downstream_accuracy · Not assessed

Reported: 47.3 pp

downstream_accuracy · Not assessed

Reported: 17.3 pp

downstream_accuracy · Not assessed

Reported: 57.8 pp

downstream_accuracy · Not assessed

Reported: 57.4 pp

downstream_accuracy · Not assessed

Reported: 52.2 pp

downstream_accuracy · Not assessed

Reported: 38.1 pp

downstream_accuracy · Not assessed

Reported: 43.6 pp

Downstream task evaluation after post-training stage SFT for 13b_sd.Experiment blocked; see the specific reasonReported 62.7 pp

Reported

62.7 pp

Observed

Intermediate checkpoint weights after post-training stage SFT not released.

downstream_accuracy · Not assessed

Reported: 62.7 pp

downstream_accuracy · Not assessed

Reported: 42.5 pp

downstream_accuracy · Not assessed

Reported: 34.8 pp

downstream_accuracy · Not assessed

Reported: 31.6 pp

downstream_accuracy · Not assessed

Reported: 43.5 pp

downstream_accuracy · Not assessed

Reported: 10.2 pp

downstream_accuracy · Not assessed

Reported: 54.9 pp

downstream_accuracy · Not assessed

Reported: 24.1 pp

downstream_accuracy · Not assessed

Reported: 8.5 pp

downstream_accuracy · Not assessed

Reported: 48 pp

downstream_accuracy · Not assessed

Reported: 17.7 pp

downstream_accuracy · Not assessed

Reported: 55.1 pp

downstream_accuracy · Not assessed

Reported: 59.4 pp

downstream_accuracy · Not assessed

Reported: 52.2 pp

downstream_accuracy · Not assessed

Reported: 38 pp

downstream_accuracy · Not assessed

Reported: 43.1 pp

Factual acquisition stratification by teacher entropy quintile persists when using OLMo-2 1B Instruct and 13B Instruct teachers.Experiment blocked; see the specific reasonReported 76 pp

Reported

76 pp

Observed

Intermediate pre-training and mid-training NTP trajectory checkpoints are not public.

factual_examples_learned · Not assessed

Reported: 76 pp

factual_examples_learned · Not assessed

Reported: 48 pp

factual_examples_learned · Not assessed

Reported: 30 pp

factual_examples_learned · Not assessed

Reported: 12 pp

factual_examples_learned · Not assessed

Reported: 6 pp

factual_examples_learned · Not assessed

Reported: 81 pp

factual_examples_learned · Not assessed

Reported: 61 pp

factual_examples_learned · Not assessed

Reported: 40 pp

factual_examples_learned · Not assessed

Reported: 18 pp

factual_examples_learned · Not assessed

Reported: 8 pp

factual_examples_learned · Not assessed

Reported: 54 pp

factual_examples_learned · Not assessed

Reported: 48 pp

factual_examples_learned · Not assessed

Reported: 37 pp

factual_examples_learned · Not assessed

Reported: 28 pp

factual_examples_learned · Not assessed

Reported: 6 pp

factual_examples_learned · Not assessed

Reported: 66 pp

factual_examples_learned · Not assessed

Reported: 58 pp

factual_examples_learned · Not assessed

Reported: 43 pp

factual_examples_learned · Not assessed

Reported: 32 pp

factual_examples_learned · Not assessed

Reported: 8 pp

Teacher entropy under OLMo-2 7B Instruct strongly predicts factual acquisition under NTP: by the end of pre-training, the student learns 67% of Q1 facts vs 5% of Q5 facts; by mid-training initialization (4T tokens), 80% of Q1 facts are learned vs 7% of Q5 facts.Experiment blocked; see the specific reasonReported 67 pp

Reported

67 pp

Observed

The intermediate NTP training checkpoints across pre-training and mid-training trajectories are not released.

factual_examples_learned · Not assessed

Reported: 67 pp

factual_examples_learned · Not assessed

Reported: 47 pp

factual_examples_learned · Not assessed

Reported: 34 pp

factual_examples_learned · Not assessed

Reported: 21 pp

factual_examples_learned · Not assessed

Reported: 5 pp

factual_examples_learned · Not assessed

Reported: 80 pp

factual_examples_learned · Not assessed

Reported: 55 pp

factual_examples_learned · Not assessed

Reported: 42 pp

factual_examples_learned · Not assessed

Reported: 24 pp

factual_examples_learned · Not assessed

Reported: 7 pp

Downstream task evaluation after post-training stage RLVR1 for ntp_shared.Experiment blocked; see the specific reasonReported 69.1 pp

Reported

69.1 pp

Observed

Intermediate checkpoint weights after post-training stage RLVR1 not released.

downstream_accuracy · Not assessed

Reported: 69.1 pp

downstream_accuracy · Not assessed

Reported: 45.6 pp

downstream_accuracy · Not assessed

Reported: 40.5 pp

downstream_accuracy · Not assessed

Reported: 32.4 pp

downstream_accuracy · Not assessed

Reported: 33.4 pp

downstream_accuracy · Not assessed

Reported: 11.2 pp

downstream_accuracy · Not assessed

Reported: 54.5 pp

downstream_accuracy · Not assessed

Reported: 21.7 pp

downstream_accuracy · Not assessed

Reported: 8 pp

downstream_accuracy · Not assessed

Reported: 31.1 pp

downstream_accuracy · Not assessed

Reported: 16.9 pp

downstream_accuracy · Not assessed

Reported: 52.3 pp

downstream_accuracy · Not assessed

Reported: 54.2 pp

downstream_accuracy · Not assessed

Reported: 51.1 pp

downstream_accuracy · Not assessed

Reported: 33.2 pp

downstream_accuracy · Not assessed

Reported: 63.6 pp

Downstream task evaluation after post-training stage SFT for 7b_rkd.Experiment blocked; see the specific reasonReported 57.7 pp

Reported

57.7 pp

Observed

Intermediate checkpoint weights after post-training stage SFT not released.

downstream_accuracy · Not assessed

Reported: 57.7 pp

downstream_accuracy · Not assessed

Reported: 35.5 pp

downstream_accuracy · Not assessed

Reported: 32 pp

downstream_accuracy · Not assessed

Reported: 32.1 pp

downstream_accuracy · Not assessed

Reported: 39.9 pp

downstream_accuracy · Not assessed

Reported: 11.2 pp

downstream_accuracy · Not assessed

Reported: 53.2 pp

downstream_accuracy · Not assessed

Reported: 22.4 pp

downstream_accuracy · Not assessed

Reported: 8.4 pp

downstream_accuracy · Not assessed

Reported: 48.9 pp

downstream_accuracy · Not assessed

Reported: 18.2 pp

downstream_accuracy · Not assessed

Reported: 60.9 pp

downstream_accuracy · Not assessed

Reported: 60.2 pp

downstream_accuracy · Not assessed

Reported: 52.2 pp

downstream_accuracy · Not assessed

Reported: 39.2 pp

downstream_accuracy · Not assessed

Reported: 45.8 pp

Switch Distillation (SD, q=20%, tau=2) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.Experiment blocked; see the specific reasonReported 66 pp

Reported

66 pp

Observed

Mid-trained checkpoint weights not released and mid-training exceeds compute limits.

downstream_accuracy · Not assessed

Reported: 66 pp

downstream_accuracy · Not assessed

Reported: 54.8 pp

downstream_accuracy · Not assessed

Reported: 42.2 pp

downstream_accuracy · Not assessed

Reported: 31.6 pp

downstream_accuracy · Not assessed

Reported: 45.8 pp

downstream_accuracy · Not assessed

Reported: 12 pp

downstream_accuracy · Not assessed

Reported: 54.1 pp

downstream_accuracy · Not assessed

Reported: 24.9 pp

downstream_accuracy · Not assessed

Reported: 9 pp

downstream_accuracy · Not assessed

Reported: 48.5 pp

downstream_accuracy · Not assessed

Reported: 17.7 pp

downstream_accuracy · Not assessed

Reported: 58.5 pp

downstream_accuracy · Not assessed

Reported: 63.8 pp

downstream_accuracy · Not assessed

Reported: 51.2 pp

downstream_accuracy · Not assessed

Reported: 39.2 pp

Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q30.Experiment blocked; see the specific reasonReported 69.8 pp

Reported

69.8 pp

Observed

Threshold sweep checkpoints not released.

downstream_accuracy · Not assessed

Reported: 69.8 pp

downstream_accuracy · Not assessed

Reported: 54.1 pp

downstream_accuracy · Not assessed

Reported: 45.6 pp

downstream_accuracy · Not assessed

Reported: 33.8 pp

downstream_accuracy · Not assessed

Reported: 47.2 pp

downstream_accuracy · Not assessed

Reported: 13 pp

downstream_accuracy · Not assessed

Reported: 43.9 pp

downstream_accuracy · Not assessed

Reported: 54.7 pp

downstream_accuracy · Not assessed

Reported: 25.3 pp

downstream_accuracy · Not assessed

Reported: 8.8 pp

downstream_accuracy · Not assessed

Reported: 29.6 pp

downstream_accuracy · Not assessed

Reported: 51 pp

downstream_accuracy · Not assessed

Reported: 19.5 pp

downstream_accuracy · Not assessed

Reported: 63.5 pp

downstream_accuracy · Not assessed

Reported: 62.4 pp

downstream_accuracy · Not assessed

Reported: 53 pp

downstream_accuracy · Not assessed

Reported: 40.7 pp

downstream_accuracy · Not assessed

Reported: 48.4 pp

Per-task downstream accuracy for mid-training ablation random_routing using OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 61.5 pp

Reported

61.5 pp

Observed

Per-task ablation checkpoints not released.

downstream_accuracy · Not assessed

Reported: 61.5 pp

downstream_accuracy · Not assessed

Reported: 45.5 pp

downstream_accuracy · Not assessed

Reported: 38.6 pp

downstream_accuracy · Not assessed

Reported: 32 pp

downstream_accuracy · Not assessed

Reported: 40.6 pp

downstream_accuracy · Not assessed

Reported: 10.8 pp

downstream_accuracy · Not assessed

Reported: 53.2 pp

downstream_accuracy · Not assessed

Reported: 23.9 pp

downstream_accuracy · Not assessed

Reported: 8.5 pp

downstream_accuracy · Not assessed

Reported: 49.6 pp

downstream_accuracy · Not assessed

Reported: 17.8 pp

downstream_accuracy · Not assessed

Reported: 61.6 pp

downstream_accuracy · Not assessed

Reported: 62.6 pp

downstream_accuracy · Not assessed

Reported: 52.6 pp

downstream_accuracy · Not assessed

Reported: 39.5 pp

Per-task downstream accuracy for mid-training ablation teacher_top1 using OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 62.2 pp

Reported

62.2 pp

Observed

Per-task ablation checkpoints not released.

downstream_accuracy · Not assessed

Reported: 62.2 pp

downstream_accuracy · Not assessed

Reported: 46 pp

downstream_accuracy · Not assessed

Reported: 40.1 pp

downstream_accuracy · Not assessed

Reported: 30.7 pp

downstream_accuracy · Not assessed

Reported: 42.1 pp

downstream_accuracy · Not assessed

Reported: 8.8 pp

downstream_accuracy · Not assessed

Reported: 57.9 pp

downstream_accuracy · Not assessed

Reported: 24.7 pp

downstream_accuracy · Not assessed

Reported: 9.3 pp

downstream_accuracy · Not assessed

Reported: 48.8 pp

downstream_accuracy · Not assessed

Reported: 18 pp

downstream_accuracy · Not assessed

Reported: 59.8 pp

downstream_accuracy · Not assessed

Reported: 62 pp

downstream_accuracy · Not assessed

Reported: 51 pp

downstream_accuracy · Not assessed

Reported: 39.2 pp

Mid-training ablation results for Switch Distillation absolute macro-averages across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported 44.7 pp

Reported

44.7 pp

Observed

Ablation checkpoints not released and training exceeds compute limits.

macro_average · Not assessed

Reported: 44.7 pp

macro_average · Not assessed

Reported: 29.3 pp

macro_average · Not assessed

Reported: 49.3 pp

Downstream task evaluation after post-training stage SFT for 7b_sd.Experiment blocked; see the specific reasonReported 63.7 pp

Reported

63.7 pp

Observed

Intermediate checkpoint weights after post-training stage SFT not released.

downstream_accuracy · Not assessed

Reported: 63.7 pp

downstream_accuracy · Not assessed

Reported: 42.7 pp

downstream_accuracy · Not assessed

Reported: 36.7 pp

downstream_accuracy · Not assessed

Reported: 33.5 pp

downstream_accuracy · Not assessed

Reported: 48.9 pp

downstream_accuracy · Not assessed

Reported: 12.8 pp

downstream_accuracy · Not assessed

Reported: 55.1 pp

downstream_accuracy · Not assessed

Reported: 24.5 pp

downstream_accuracy · Not assessed

Reported: 8.5 pp

downstream_accuracy · Not assessed

Reported: 50.3 pp

downstream_accuracy · Not assessed

Reported: 19.4 pp

downstream_accuracy · Not assessed

Reported: 61.3 pp

downstream_accuracy · Not assessed

Reported: 61.4 pp

downstream_accuracy · Not assessed

Reported: 56.1 pp

downstream_accuracy · Not assessed

Reported: 40.6 pp

downstream_accuracy · Not assessed

Reported: 45.5 pp

Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q10.Experiment blocked; see the specific reasonReported 62.1 pp

Reported

62.1 pp

Observed

Threshold sweep checkpoints not released.

downstream_accuracy · Not assessed

Reported: 62.1 pp

downstream_accuracy · Not assessed

Reported: 49.6 pp

downstream_accuracy · Not assessed

Reported: 38.3 pp

downstream_accuracy · Not assessed

Reported: 28.6 pp

downstream_accuracy · Not assessed

Reported: 44.2 pp

downstream_accuracy · Not assessed

Reported: 7.8 pp

downstream_accuracy · Not assessed

Reported: 38.5 pp

downstream_accuracy · Not assessed

Reported: 52.2 pp

downstream_accuracy · Not assessed

Reported: 23.1 pp

downstream_accuracy · Not assessed

Reported: 9.2 pp

downstream_accuracy · Not assessed

Reported: 28.1 pp

downstream_accuracy · Not assessed

Reported: 46.8 pp

downstream_accuracy · Not assessed

Reported: 17.1 pp

downstream_accuracy · Not assessed

Reported: 55.5 pp

downstream_accuracy · Not assessed

Reported: 59.6 pp

downstream_accuracy · Not assessed

Reported: 52.7 pp

downstream_accuracy · Not assessed

Reported: 37.9 pp

downstream_accuracy · Not assessed

Reported: 45 pp

Downstream task evaluation after post-training stage DPO for 7b_trkd.Experiment blocked; see the specific reasonReported 57.2 pp

Reported

57.2 pp

Observed

Intermediate checkpoint weights after post-training stage DPO not released.

downstream_accuracy · Not assessed

Reported: 57.2 pp

downstream_accuracy · Not assessed

Reported: 34.9 pp

downstream_accuracy · Not assessed

Reported: 32.3 pp

downstream_accuracy · Not assessed

Reported: 32.4 pp

downstream_accuracy · Not assessed

Reported: 36.6 pp

downstream_accuracy · Not assessed

Reported: 7.2 pp

downstream_accuracy · Not assessed

Reported: 53.1 pp

downstream_accuracy · Not assessed

Reported: 23.2 pp

downstream_accuracy · Not assessed

Reported: 7.9 pp

downstream_accuracy · Not assessed

Reported: 47 pp

downstream_accuracy · Not assessed

Reported: 17.9 pp

downstream_accuracy · Not assessed

Reported: 56.7 pp

downstream_accuracy · Not assessed

Reported: 57.2 pp

downstream_accuracy · Not assessed

Reported: 51.9 pp

downstream_accuracy · Not assessed

Reported: 37.5 pp

downstream_accuracy · Not assessed

Reported: 62.5 pp

Forward KL Distillation (FKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 57.8 pp

Reported

57.8 pp

Observed

Mid-trained checkpoint weights not released and mid-training exceeds compute limits.

downstream_accuracy · Not assessed

Reported: 57.8 pp

downstream_accuracy · Not assessed

Reported: 42.6 pp

downstream_accuracy · Not assessed

Reported: 36.2 pp

downstream_accuracy · Not assessed

Reported: 30.6 pp

downstream_accuracy · Not assessed

Reported: 41.9 pp

downstream_accuracy · Not assessed

Reported: 7.6 pp

downstream_accuracy · Not assessed

Reported: 53.9 pp

downstream_accuracy · Not assessed

Reported: 24.6 pp

downstream_accuracy · Not assessed

Reported: 8.4 pp

downstream_accuracy · Not assessed

Reported: 49.5 pp

downstream_accuracy · Not assessed

Reported: 18.7 pp

downstream_accuracy · Not assessed

Reported: 61.9 pp

downstream_accuracy · Not assessed

Reported: 61.8 pp

downstream_accuracy · Not assessed

Reported: 51.5 pp

downstream_accuracy · Not assessed

Reported: 38.6 pp

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B TRKD post-training performance.Experiment blocked; see the specific reasonReported 70.7 pp

Reported

70.7 pp

Observed

Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.

downstream_accuracy · Not assessed

Reported: 70.7 pp

downstream_accuracy · Not assessed

Reported: 52.9 pp

downstream_accuracy · Not assessed

Reported: 46.3 pp

downstream_accuracy · Not assessed

Reported: 31.8 pp

downstream_accuracy · Not assessed

Reported: 36.8 pp

downstream_accuracy · Not assessed

Reported: 14.4 pp

downstream_accuracy · Not assessed

Reported: 51.7 pp

downstream_accuracy · Not assessed

Reported: 22.3 pp

downstream_accuracy · Not assessed

Reported: 8 pp

downstream_accuracy · Not assessed

Reported: 40.9 pp

downstream_accuracy · Not assessed

Reported: 17.4 pp

downstream_accuracy · Not assessed

Reported: 59 pp

downstream_accuracy · Not assessed

Reported: 53.6 pp

downstream_accuracy · Not assessed

Reported: 51.5 pp

downstream_accuracy · Not assessed

Reported: 35.6 pp

downstream_accuracy · Not assessed

Reported: 65.2 pp

Teacher predictive entropy distinguishes procedural from knowledge-intensive domains across diverse instruction-tuned open-weight model families (OLMo-3 7B Instruct: 0.771, Qwen 3 8B: 0.705, Gemma-3 12B it: 0.707, Granite 3.3 8B Instruct: 0.696).Awaiting reproductionReported 0.771 score

Reported

0.771 score

Observed

Experiment plan ready; no runs yet.

roc_auc · Not assessed

Reported: 0.771 score

roc_auc · Not assessed

Reported: 0.705 score

roc_auc · Not assessed

Reported: 0.707 score

roc_auc · Not assessed

Reported: 0.696 score

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B FKD post-training performance.Experiment blocked; see the specific reasonReported 76.2 pp

Reported

76.2 pp

Observed

Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.

downstream_accuracy · Not assessed

Reported: 76.2 pp

downstream_accuracy · Not assessed

Reported: 56.4 pp

downstream_accuracy · Not assessed

Reported: 51.6 pp

downstream_accuracy · Not assessed

Reported: 33.8 pp

downstream_accuracy · Not assessed

Reported: 38.4 pp

downstream_accuracy · Not assessed

Reported: 18 pp

downstream_accuracy · Not assessed

Reported: 51.9 pp

downstream_accuracy · Not assessed

Reported: 22.2 pp

downstream_accuracy · Not assessed

Reported: 8.1 pp

downstream_accuracy · Not assessed

Reported: 42.2 pp

downstream_accuracy · Not assessed

Reported: 18.6 pp

downstream_accuracy · Not assessed

Reported: 60.2 pp

downstream_accuracy · Not assessed

Reported: 57.4 pp

downstream_accuracy · Not assessed

Reported: 51.5 pp

downstream_accuracy · Not assessed

Reported: 38 pp

downstream_accuracy · Not assessed

Reported: 64.7 pp

Downstream task evaluation after post-training stage DPO for 13b_sd.Experiment blocked; see the specific reasonReported 69.9 pp

Reported

69.9 pp

Observed

Intermediate checkpoint weights after post-training stage DPO not released.

downstream_accuracy · Not assessed

Reported: 69.9 pp

downstream_accuracy · Not assessed

Reported: 45.7 pp

downstream_accuracy · Not assessed

Reported: 39.5 pp

downstream_accuracy · Not assessed

Reported: 32.2 pp

downstream_accuracy · Not assessed

Reported: 44.2 pp

downstream_accuracy · Not assessed

Reported: 10 pp

downstream_accuracy · Not assessed

Reported: 54.5 pp

downstream_accuracy · Not assessed

Reported: 23.9 pp

downstream_accuracy · Not assessed

Reported: 8.2 pp

downstream_accuracy · Not assessed

Reported: 47.3 pp

downstream_accuracy · Not assessed

Reported: 17.9 pp

downstream_accuracy · Not assessed

Reported: 56.1 pp

downstream_accuracy · Not assessed

Reported: 59.2 pp

downstream_accuracy · Not assessed

Reported: 52.5 pp

downstream_accuracy · Not assessed

Reported: 38.7 pp

downstream_accuracy · Not assessed

Reported: 61 pp

Mid-training ablation results for Oracle Domain Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported -7.2 pp

Reported

-7.2 pp

Observed

Ablation checkpoints not released and training exceeds compute limits.

delta_percentage_points · Not assessed

Reported: -7.2 pp

delta_percentage_points · Not assessed

Reported: -1.3 pp

delta_percentage_points · Not assessed

Reported: -2.3 pp

Downstream task evaluation after post-training stage RLVR1 for 7b_rkd.Experiment blocked; see the specific reasonReported 77.9 pp

Reported

77.9 pp

Observed

Intermediate checkpoint weights after post-training stage RLVR1 not released.

downstream_accuracy · Not assessed

Reported: 77.9 pp

downstream_accuracy · Not assessed

Reported: 54.3 pp

downstream_accuracy · Not assessed

Reported: 52.5 pp

downstream_accuracy · Not assessed

Reported: 34.2 pp

downstream_accuracy · Not assessed

Reported: 40.1 pp

downstream_accuracy · Not assessed

Reported: 15.6 pp

downstream_accuracy · Not assessed

Reported: 52.1 pp

downstream_accuracy · Not assessed

Reported: 21.9 pp

downstream_accuracy · Not assessed

Reported: 7.8 pp

downstream_accuracy · Not assessed

Reported: 45.3 pp

downstream_accuracy · Not assessed

Reported: 16.8 pp

downstream_accuracy · Not assessed

Reported: 57.6 pp

downstream_accuracy · Not assessed

Reported: 50.4 pp

downstream_accuracy · Not assessed

Reported: 51.5 pp

downstream_accuracy · Not assessed

Reported: 36.1 pp

downstream_accuracy · Not assessed

Reported: 59.3 pp

Reverse KL Distillation (RKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.Experiment blocked; see the specific reasonReported 59.2 pp

Reported

59.2 pp

Observed

Mid-trained checkpoint weights not released and mid-training exceeds compute limits.

downstream_accuracy · Not assessed

Reported: 59.2 pp

downstream_accuracy · Not assessed

Reported: 48.7 pp

downstream_accuracy · Not assessed

Reported: 37.4 pp

downstream_accuracy · Not assessed

Reported: 31.4 pp

downstream_accuracy · Not assessed

Reported: 37.3 pp

downstream_accuracy · Not assessed

Reported: 8 pp

downstream_accuracy · Not assessed

Reported: 53.4 pp

downstream_accuracy · Not assessed

Reported: 24 pp

downstream_accuracy · Not assessed

Reported: 8.5 pp

downstream_accuracy · Not assessed

Reported: 48.4 pp

downstream_accuracy · Not assessed

Reported: 17.3 pp

downstream_accuracy · Not assessed

Reported: 57.8 pp

downstream_accuracy · Not assessed

Reported: 60 pp

downstream_accuracy · Not assessed

Reported: 51 pp

downstream_accuracy · Not assessed

Reported: 38.3 pp

Downstream task evaluation after post-training stage SFT for 7b_trkd.Experiment blocked; see the specific reasonReported 49.9 pp

Reported

49.9 pp

Observed

Intermediate checkpoint weights after post-training stage SFT not released.

downstream_accuracy · Not assessed

Reported: 49.9 pp

downstream_accuracy · Not assessed

Reported: 31.1 pp

downstream_accuracy · Not assessed

Reported: 26.9 pp

downstream_accuracy · Not assessed

Reported: 31.5 pp

downstream_accuracy · Not assessed

Reported: 37.1 pp

downstream_accuracy · Not assessed

Reported: 7.6 pp

downstream_accuracy · Not assessed

Reported: 53.7 pp

downstream_accuracy · Not assessed

Reported: 23.1 pp

downstream_accuracy · Not assessed

Reported: 8.4 pp

downstream_accuracy · Not assessed

Reported: 46.5 pp

downstream_accuracy · Not assessed

Reported: 17.4 pp

downstream_accuracy · Not assessed

Reported: 57.8 pp

downstream_accuracy · Not assessed

Reported: 56 pp

downstream_accuracy · Not assessed

Reported: 51.4 pp

downstream_accuracy · Not assessed

Reported: 37.5 pp

downstream_accuracy · Not assessed

Reported: 45.3 pp

Mid-training ablation results for Teacher Top-1 Labels relative difference across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported -6.4 pp

Reported

-6.4 pp

Observed

Ablation checkpoints not released and training exceeds compute limits.

delta_percentage_points · Not assessed

Reported: -6.4 pp

delta_percentage_points · Not assessed

Reported: 1.3 pp

delta_percentage_points · Not assessed

Reported: -2.8 pp

Downstream task evaluation after post-training stage SFT for 13b_fkd.Experiment blocked; see the specific reasonReported 54.1 pp

Reported

54.1 pp

Observed

Intermediate checkpoint weights after post-training stage SFT not released.

downstream_accuracy · Not assessed

Reported: 54.1 pp

downstream_accuracy · Not assessed

Reported: 32.5 pp

downstream_accuracy · Not assessed

Reported: 28 pp

downstream_accuracy · Not assessed

Reported: 31.8 pp

downstream_accuracy · Not assessed

Reported: 36.3 pp

downstream_accuracy · Not assessed

Reported: 7.6 pp

downstream_accuracy · Not assessed

Reported: 54.9 pp

downstream_accuracy · Not assessed

Reported: 23.5 pp

downstream_accuracy · Not assessed

Reported: 7.9 pp

downstream_accuracy · Not assessed

Reported: 45.9 pp

downstream_accuracy · Not assessed

Reported: 17.1 pp

downstream_accuracy · Not assessed

Reported: 55.6 pp

downstream_accuracy · Not assessed

Reported: 57.2 pp

downstream_accuracy · Not assessed

Reported: 51.2 pp

downstream_accuracy · Not assessed

Reported: 37.1 pp

downstream_accuracy · Not assessed

Reported: 44.2 pp

Per-task downstream accuracy for mid-training ablation oracle_domain using OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 61 pp

Reported

61 pp

Observed

Per-task ablation checkpoints not released.

downstream_accuracy · Not assessed

Reported: 61 pp

downstream_accuracy · Not assessed

Reported: 45.6 pp

downstream_accuracy · Not assessed

Reported: 38.5 pp

downstream_accuracy · Not assessed

Reported: 31.6 pp

downstream_accuracy · Not assessed

Reported: 40.1 pp

downstream_accuracy · Not assessed

Reported: 8 pp

downstream_accuracy · Not assessed

Reported: 53 pp

downstream_accuracy · Not assessed

Reported: 22.6 pp

downstream_accuracy · Not assessed

Reported: 8.3 pp

downstream_accuracy · Not assessed

Reported: 49.4 pp

downstream_accuracy · Not assessed

Reported: 18.8 pp

downstream_accuracy · Not assessed

Reported: 61.1 pp

downstream_accuracy · Not assessed

Reported: 61.4 pp

downstream_accuracy · Not assessed

Reported: 52.6 pp

downstream_accuracy · Not assessed

Reported: 38.9 pp

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B FKD post-training performance.Experiment blocked; see the specific reasonReported 72.1 pp

Reported

72.1 pp

Observed

Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.

downstream_accuracy · Not assessed

Reported: 72.1 pp

downstream_accuracy · Not assessed

Reported: 54.3 pp

downstream_accuracy · Not assessed

Reported: 46.5 pp

downstream_accuracy · Not assessed

Reported: 31.4 pp

downstream_accuracy · Not assessed

Reported: 35.7 pp

downstream_accuracy · Not assessed

Reported: 14.6 pp

downstream_accuracy · Not assessed

Reported: 54.1 pp

downstream_accuracy · Not assessed

Reported: 23.2 pp

downstream_accuracy · Not assessed

Reported: 8 pp

downstream_accuracy · Not assessed

Reported: 36.1 pp

downstream_accuracy · Not assessed

Reported: 18.1 pp

downstream_accuracy · Not assessed

Reported: 55 pp

downstream_accuracy · Not assessed

Reported: 56.4 pp

downstream_accuracy · Not assessed

Reported: 51.2 pp

downstream_accuracy · Not assessed

Reported: 35.9 pp

downstream_accuracy · Not assessed

Reported: 64.9 pp

Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q10.Experiment blocked; see the specific reasonReported 61.6 pp

Reported

61.6 pp

Observed

Threshold sweep checkpoints not released.

downstream_accuracy · Not assessed

Reported: 61.6 pp

downstream_accuracy · Not assessed

Reported: 48.8 pp

downstream_accuracy · Not assessed

Reported: 39.4 pp

downstream_accuracy · Not assessed

Reported: 32.5 pp

downstream_accuracy · Not assessed

Reported: 50.2 pp

downstream_accuracy · Not assessed

Reported: 11.4 pp

downstream_accuracy · Not assessed

Reported: 40.6 pp

downstream_accuracy · Not assessed

Reported: 53.7 pp

downstream_accuracy · Not assessed

Reported: 23.9 pp

downstream_accuracy · Not assessed

Reported: 8.6 pp

downstream_accuracy · Not assessed

Reported: 28.7 pp

downstream_accuracy · Not assessed

Reported: 50.6 pp

downstream_accuracy · Not assessed

Reported: 19 pp

downstream_accuracy · Not assessed

Reported: 62 pp

downstream_accuracy · Not assessed

Reported: 63.8 pp

downstream_accuracy · Not assessed

Reported: 55.5 pp

downstream_accuracy · Not assessed

Reported: 40.4 pp

downstream_accuracy · Not assessed

Reported: 48.5 pp

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B TRKD post-training performance.Experiment blocked; see the specific reasonReported 69.6 pp

Reported

69.6 pp

Observed

Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.

downstream_accuracy · Not assessed

Reported: 69.6 pp

downstream_accuracy · Not assessed

Reported: 48.5 pp

downstream_accuracy · Not assessed

Reported: 42.6 pp

downstream_accuracy · Not assessed

Reported: 31.5 pp

downstream_accuracy · Not assessed

Reported: 34.6 pp

downstream_accuracy · Not assessed

Reported: 12.2 pp

downstream_accuracy · Not assessed

Reported: 53.9 pp

downstream_accuracy · Not assessed

Reported: 22.2 pp

downstream_accuracy · Not assessed

Reported: 8.2 pp

downstream_accuracy · Not assessed

Reported: 37 pp

downstream_accuracy · Not assessed

Reported: 17.3 pp

downstream_accuracy · Not assessed

Reported: 55.4 pp

downstream_accuracy · Not assessed

Reported: 54 pp

downstream_accuracy · Not assessed

Reported: 51.3 pp

downstream_accuracy · Not assessed

Reported: 33.6 pp

downstream_accuracy · Not assessed

Reported: 64.1 pp

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B RKD post-training performance.Experiment blocked; see the specific reasonReported 73.8 pp

Reported

73.8 pp

Observed

Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.

downstream_accuracy · Not assessed

Reported: 73.8 pp

downstream_accuracy · Not assessed

Reported: 53.4 pp

downstream_accuracy · Not assessed

Reported: 49.5 pp

downstream_accuracy · Not assessed

Reported: 31.6 pp

downstream_accuracy · Not assessed

Reported: 40.5 pp

downstream_accuracy · Not assessed

Reported: 17.6 pp

downstream_accuracy · Not assessed

Reported: 51.8 pp

downstream_accuracy · Not assessed

Reported: 21.9 pp

downstream_accuracy · Not assessed

Reported: 7.9 pp

downstream_accuracy · Not assessed

Reported: 46.2 pp

downstream_accuracy · Not assessed

Reported: 18 pp

downstream_accuracy · Not assessed

Reported: 61.1 pp

downstream_accuracy · Not assessed

Reported: 53 pp

downstream_accuracy · Not assessed

Reported: 51.2 pp

downstream_accuracy · Not assessed

Reported: 37.1 pp

downstream_accuracy · Not assessed

Reported: 61.6 pp

Downstream task evaluation after post-training stage RLVR1 for 7b_fkd.Experiment blocked; see the specific reasonReported 76.7 pp

Reported

76.7 pp

Observed

Intermediate checkpoint weights after post-training stage RLVR1 not released.

downstream_accuracy · Not assessed

Reported: 76.7 pp

downstream_accuracy · Not assessed

Reported: 59.4 pp

downstream_accuracy · Not assessed

Reported: 51.9 pp

downstream_accuracy · Not assessed

Reported: 34.4 pp

downstream_accuracy · Not assessed

Reported: 39.2 pp

downstream_accuracy · Not assessed

Reported: 16.8 pp

downstream_accuracy · Not assessed

Reported: 52.8 pp

downstream_accuracy · Not assessed

Reported: 23.1 pp

downstream_accuracy · Not assessed

Reported: 8.3 pp

downstream_accuracy · Not assessed

Reported: 41.4 pp

downstream_accuracy · Not assessed

Reported: 18.3 pp

downstream_accuracy · Not assessed

Reported: 59.4 pp

downstream_accuracy · Not assessed

Reported: 56.2 pp

downstream_accuracy · Not assessed

Reported: 51.2 pp

downstream_accuracy · Not assessed

Reported: 38 pp

downstream_accuracy · Not assessed

Reported: 67.1 pp

Downstream task evaluation after post-training stage DPO for 13b_fkd.Experiment blocked; see the specific reasonReported 61 pp

Reported

61 pp

Observed

Intermediate checkpoint weights after post-training stage DPO not released.

downstream_accuracy · Not assessed

Reported: 61 pp

downstream_accuracy · Not assessed

Reported: 34.2 pp

downstream_accuracy · Not assessed

Reported: 31.7 pp

downstream_accuracy · Not assessed

Reported: 33.4 pp

downstream_accuracy · Not assessed

Reported: 36.9 pp

downstream_accuracy · Not assessed

Reported: 8.2 pp

downstream_accuracy · Not assessed

Reported: 54.6 pp

downstream_accuracy · Not assessed

Reported: 23.2 pp

downstream_accuracy · Not assessed

Reported: 7.7 pp

downstream_accuracy · Not assessed

Reported: 45.9 pp

downstream_accuracy · Not assessed

Reported: 17.4 pp

downstream_accuracy · Not assessed

Reported: 53.2 pp

downstream_accuracy · Not assessed

Reported: 57.2 pp

downstream_accuracy · Not assessed

Reported: 52.3 pp

downstream_accuracy · Not assessed

Reported: 37.4 pp

downstream_accuracy · Not assessed

Reported: 64.7 pp

Switch Distillation (SD, q=20%, tau=2) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher, demonstrating substantial gains on reasoning while maintaining factual recall.Experiment blocked; see the specific reasonReported 69.7 pp

Reported

69.7 pp

Observed

Mid-trained checkpoint weights not released and mid-training exceeds compute limits.

downstream_accuracy · Not assessed

Reported: 69.7 pp

downstream_accuracy · Not assessed

Reported: 55.3 pp

downstream_accuracy · Not assessed

Reported: 46.1 pp

downstream_accuracy · Not assessed

Reported: 32.8 pp

downstream_accuracy · Not assessed

Reported: 49.6 pp

downstream_accuracy · Not assessed

Reported: 14.8 pp

downstream_accuracy · Not assessed

Reported: 54.9 pp

downstream_accuracy · Not assessed

Reported: 24.6 pp

downstream_accuracy · Not assessed

Reported: 8.4 pp

downstream_accuracy · Not assessed

Reported: 51.6 pp

downstream_accuracy · Not assessed

Reported: 19.8 pp

downstream_accuracy · Not assessed

Reported: 64.7 pp

downstream_accuracy · Not assessed

Reported: 64.2 pp

downstream_accuracy · Not assessed

Reported: 53.8 pp

downstream_accuracy · Not assessed

Reported: 41.5 pp

Downstream task evaluation after post-training stage DPO for 13b_rkd.Experiment blocked; see the specific reasonReported 65.6 pp

Reported

65.6 pp

Observed

Intermediate checkpoint weights after post-training stage DPO not released.

downstream_accuracy · Not assessed

Reported: 65.6 pp

downstream_accuracy · Not assessed

Reported: 41.2 pp

downstream_accuracy · Not assessed

Reported: 35 pp

downstream_accuracy · Not assessed

Reported: 33.2 pp

downstream_accuracy · Not assessed

Reported: 39.9 pp

downstream_accuracy · Not assessed

Reported: 7.4 pp

downstream_accuracy · Not assessed

Reported: 53.5 pp

downstream_accuracy · Not assessed

Reported: 23.7 pp

downstream_accuracy · Not assessed

Reported: 7.9 pp

downstream_accuracy · Not assessed

Reported: 47.8 pp

downstream_accuracy · Not assessed

Reported: 17.9 pp

downstream_accuracy · Not assessed

Reported: 57.1 pp

downstream_accuracy · Not assessed

Reported: 56.2 pp

downstream_accuracy · Not assessed

Reported: 51.9 pp

downstream_accuracy · Not assessed

Reported: 38.5 pp

downstream_accuracy · Not assessed

Reported: 64.5 pp

Token-Routing Knowledge Distillation (TRKD) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 52.5 pp

Reported

52.5 pp

Observed

Mid-trained checkpoint weights not released and mid-training exceeds compute limits.

downstream_accuracy · Not assessed

Reported: 52.5 pp

downstream_accuracy · Not assessed

Reported: 36.9 pp

downstream_accuracy · Not assessed

Reported: 31 pp

downstream_accuracy · Not assessed

Reported: 30.3 pp

downstream_accuracy · Not assessed

Reported: 36.8 pp

downstream_accuracy · Not assessed

Reported: 5.6 pp

downstream_accuracy · Not assessed

Reported: 54 pp

downstream_accuracy · Not assessed

Reported: 24.1 pp

downstream_accuracy · Not assessed

Reported: 7.8 pp

downstream_accuracy · Not assessed

Reported: 47.9 pp

downstream_accuracy · Not assessed

Reported: 17.4 pp

downstream_accuracy · Not assessed

Reported: 58.9 pp

downstream_accuracy · Not assessed

Reported: 60 pp

downstream_accuracy · Not assessed

Reported: 51.2 pp

downstream_accuracy · Not assessed

Reported: 37.5 pp

Forward KL Distillation (FKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.Experiment blocked; see the specific reasonReported 52.7 pp

Reported

52.7 pp

Observed

Mid-trained checkpoint weights not released and mid-training exceeds compute limits.

downstream_accuracy · Not assessed

Reported: 52.7 pp

downstream_accuracy · Not assessed

Reported: 37.9 pp

downstream_accuracy · Not assessed

Reported: 31.8 pp

downstream_accuracy · Not assessed

Reported: 29.4 pp

downstream_accuracy · Not assessed

Reported: 33.1 pp

downstream_accuracy · Not assessed

Reported: 6.4 pp

downstream_accuracy · Not assessed

Reported: 54.2 pp

downstream_accuracy · Not assessed

Reported: 24.7 pp

downstream_accuracy · Not assessed

Reported: 8.1 pp

downstream_accuracy · Not assessed

Reported: 47.5 pp

downstream_accuracy · Not assessed

Reported: 16.3 pp

downstream_accuracy · Not assessed

Reported: 55.3 pp

downstream_accuracy · Not assessed

Reported: 56.2 pp

downstream_accuracy · Not assessed

Reported: 51.2 pp

downstream_accuracy · Not assessed

Reported: 36.2 pp

Standard Next-Token Prediction (NTP) baseline downstream evaluation performance after 60B mid-training tokens across 15 tasks covering Reasoning, Factual Recall, and Knowledge & Commonsense.Experiment blocked; see the specific reasonReported 40.4 pp

Reported

40.4 pp

Observed

Official model checkpoints from mid-training are not publicly released, and mid-training 1B students from 4T tokens requires compute far exceeding the resource budget.

downstream_accuracy · Not assessed

Reported: 40.4 pp

downstream_accuracy · Not assessed

Reported: 29.8 pp

downstream_accuracy · Not assessed

Reported: 23.1 pp

downstream_accuracy · Not assessed

Reported: 29.9 pp

downstream_accuracy · Not assessed

Reported: 29.8 pp

downstream_accuracy · Not assessed

Reported: 3.8 pp

downstream_accuracy · Not assessed

Reported: 56.7 pp

downstream_accuracy · Not assessed

Reported: 25.5 pp

downstream_accuracy · Not assessed

Reported: 8.7 pp

downstream_accuracy · Not assessed

Reported: 43.6 pp

downstream_accuracy · Not assessed

Reported: 15.5 pp

downstream_accuracy · Not assessed

Reported: 51.1 pp

downstream_accuracy · Not assessed

Reported: 51.8 pp

downstream_accuracy · Not assessed

Reported: 51.4 pp

downstream_accuracy · Not assessed

Reported: 34.1 pp

Downstream task evaluation after post-training stage SFT for ntp_shared.Experiment blocked; see the specific reasonReported 45.2 pp

Reported

45.2 pp

Observed

Intermediate checkpoint weights after post-training stage SFT not released.

downstream_accuracy · Not assessed

Reported: 45.2 pp

downstream_accuracy · Not assessed

Reported: 28.2 pp

downstream_accuracy · Not assessed

Reported: 23.3 pp

downstream_accuracy · Not assessed

Reported: 30.2 pp

downstream_accuracy · Not assessed

Reported: 33.8 pp

downstream_accuracy · Not assessed

Reported: 6.8 pp

downstream_accuracy · Not assessed

Reported: 56.3 pp

downstream_accuracy · Not assessed

Reported: 21.7 pp

downstream_accuracy · Not assessed

Reported: 8.2 pp

downstream_accuracy · Not assessed

Reported: 39.2 pp

downstream_accuracy · Not assessed

Reported: 16.3 pp

downstream_accuracy · Not assessed

Reported: 51.1 pp

downstream_accuracy · Not assessed

Reported: 51.4 pp

downstream_accuracy · Not assessed

Reported: 51.5 pp

downstream_accuracy · Not assessed

Reported: 33.2 pp

downstream_accuracy · Not assessed

Reported: 45.8 pp

Token-Routing Knowledge Distillation (TRKD) downstream evaluation performance after mid-training with OLMo-2 13B Instruct teacher.Experiment blocked; see the specific reasonReported 47.8 pp

Reported

47.8 pp

Observed

Mid-trained checkpoint weights not released and mid-training exceeds compute limits.

downstream_accuracy · Not assessed

Reported: 47.8 pp

downstream_accuracy · Not assessed

Reported: 32.7 pp

downstream_accuracy · Not assessed

Reported: 28.4 pp

downstream_accuracy · Not assessed

Reported: 29.3 pp

downstream_accuracy · Not assessed

Reported: 34 pp

downstream_accuracy · Not assessed

Reported: 5.6 pp

downstream_accuracy · Not assessed

Reported: 53.8 pp

downstream_accuracy · Not assessed

Reported: 24.2 pp

downstream_accuracy · Not assessed

Reported: 8.3 pp

downstream_accuracy · Not assessed

Reported: 45.7 pp

downstream_accuracy · Not assessed

Reported: 15.4 pp

downstream_accuracy · Not assessed

Reported: 53 pp

downstream_accuracy · Not assessed

Reported: 56.6 pp

downstream_accuracy · Not assessed

Reported: 50.7 pp

downstream_accuracy · Not assessed

Reported: 35.2 pp

Downstream task evaluation after post-training stage DPO for 7b_rkd.Experiment blocked; see the specific reasonReported 64.5 pp

Reported

64.5 pp

Observed

Intermediate checkpoint weights after post-training stage DPO not released.

downstream_accuracy · Not assessed

Reported: 64.5 pp

downstream_accuracy · Not assessed

Reported: 40 pp

downstream_accuracy · Not assessed

Reported: 36.4 pp

downstream_accuracy · Not assessed

Reported: 34.2 pp

downstream_accuracy · Not assessed

Reported: 40.5 pp

downstream_accuracy · Not assessed

Reported: 10 pp

downstream_accuracy · Not assessed

Reported: 52.8 pp

downstream_accuracy · Not assessed

Reported: 22.3 pp

downstream_accuracy · Not assessed

Reported: 7.8 pp

downstream_accuracy · Not assessed

Reported: 48.8 pp

downstream_accuracy · Not assessed

Reported: 17.3 pp

downstream_accuracy · Not assessed

Reported: 57.4 pp

downstream_accuracy · Not assessed

Reported: 55 pp

downstream_accuracy · Not assessed

Reported: 52.5 pp

downstream_accuracy · Not assessed

Reported: 38.2 pp

downstream_accuracy · Not assessed

Reported: 64.1 pp

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B Switch Distillation post-training performance.Experiment blocked; see the specific reasonReported 77.8 pp

Reported

77.8 pp

Observed

Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.

downstream_accuracy · Not assessed

Reported: 77.8 pp

downstream_accuracy · Not assessed

Reported: 62.8 pp

downstream_accuracy · Not assessed

Reported: 51.8 pp

downstream_accuracy · Not assessed

Reported: 33.3 pp

downstream_accuracy · Not assessed

Reported: 42.8 pp

downstream_accuracy · Not assessed

Reported: 19.6 pp

downstream_accuracy · Not assessed

Reported: 54.2 pp

downstream_accuracy · Not assessed

Reported: 24.1 pp

downstream_accuracy · Not assessed

Reported: 8.4 pp

downstream_accuracy · Not assessed

Reported: 43.1 pp

downstream_accuracy · Not assessed

Reported: 18.1 pp

downstream_accuracy · Not assessed

Reported: 56.5 pp

downstream_accuracy · Not assessed

Reported: 59 pp

downstream_accuracy · Not assessed

Reported: 52.5 pp

downstream_accuracy · Not assessed

Reported: 38.1 pp

downstream_accuracy · Not assessed

Reported: 67.1 pp

Downstream task evaluation after post-training stage DPO for 7b_sd.Experiment blocked; see the specific reasonReported 70.9 pp

Reported

70.9 pp

Observed

Intermediate checkpoint weights after post-training stage DPO not released.

downstream_accuracy · Not assessed

Reported: 70.9 pp

downstream_accuracy · Not assessed

Reported: 50.8 pp

downstream_accuracy · Not assessed

Reported: 44.2 pp

downstream_accuracy · Not assessed

Reported: 35.5 pp

downstream_accuracy · Not assessed

Reported: 49 pp

downstream_accuracy · Not assessed

Reported: 15 pp

downstream_accuracy · Not assessed

Reported: 54.7 pp

downstream_accuracy · Not assessed

Reported: 23.6 pp

downstream_accuracy · Not assessed

Reported: 8.1 pp

downstream_accuracy · Not assessed

Reported: 50 pp

downstream_accuracy · Not assessed

Reported: 20.1 pp

downstream_accuracy · Not assessed

Reported: 61.2 pp

downstream_accuracy · Not assessed

Reported: 61 pp

downstream_accuracy · Not assessed

Reported: 55.5 pp

downstream_accuracy · Not assessed

Reported: 40.9 pp

downstream_accuracy · Not assessed

Reported: 62.1 pp

Downstream task evaluation after post-training stage DPO for ntp_shared.Experiment blocked; see the specific reasonReported 52.4 pp

Reported

52.4 pp

Observed

Intermediate checkpoint weights after post-training stage DPO not released.

downstream_accuracy · Not assessed

Reported: 52.4 pp

downstream_accuracy · Not assessed

Reported: 33.4 pp

downstream_accuracy · Not assessed

Reported: 28.2 pp

downstream_accuracy · Not assessed

Reported: 32.3 pp

downstream_accuracy · Not assessed

Reported: 33.5 pp

downstream_accuracy · Not assessed

Reported: 6.6 pp

downstream_accuracy · Not assessed

Reported: 55.4 pp

downstream_accuracy · Not assessed

Reported: 21.5 pp

downstream_accuracy · Not assessed

Reported: 7.7 pp

downstream_accuracy · Not assessed

Reported: 41.1 pp

downstream_accuracy · Not assessed

Reported: 16.4 pp

downstream_accuracy · Not assessed

Reported: 50.3 pp

downstream_accuracy · Not assessed

Reported: 50.6 pp

downstream_accuracy · Not assessed

Reported: 51.5 pp

downstream_accuracy · Not assessed

Reported: 34.1 pp

downstream_accuracy · Not assessed

Reported: 59.9 pp

Downstream task evaluation after post-training stage SFT for 13b_trkd.Experiment blocked; see the specific reasonReported 49 pp

Reported

49 pp

Observed

Intermediate checkpoint weights after post-training stage SFT not released.

downstream_accuracy · Not assessed

Reported: 49 pp

downstream_accuracy · Not assessed

Reported: 29.8 pp

downstream_accuracy · Not assessed

Reported: 25.4 pp

downstream_accuracy · Not assessed

Reported: 31.6 pp

downstream_accuracy · Not assessed

Reported: 35.2 pp

downstream_accuracy · Not assessed

Reported: 7.6 pp

downstream_accuracy · Not assessed

Reported: 55.4 pp

downstream_accuracy · Not assessed

Reported: 23.6 pp

downstream_accuracy · Not assessed

Reported: 8.8 pp

downstream_accuracy · Not assessed

Reported: 43.8 pp

downstream_accuracy · Not assessed

Reported: 16.7 pp

downstream_accuracy · Not assessed

Reported: 53.8 pp

downstream_accuracy · Not assessed

Reported: 55 pp

downstream_accuracy · Not assessed

Reported: 51.3 pp

downstream_accuracy · Not assessed

Reported: 36.2 pp

downstream_accuracy · Not assessed

Reported: 43.4 pp

Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q30.Experiment blocked; see the specific reasonReported 66 pp

Reported

66 pp

Observed

Threshold sweep checkpoints not released.

downstream_accuracy · Not assessed

Reported: 66 pp

downstream_accuracy · Not assessed

Reported: 53.1 pp

downstream_accuracy · Not assessed

Reported: 42.9 pp

downstream_accuracy · Not assessed

Reported: 32.5 pp

downstream_accuracy · Not assessed

Reported: 44.4 pp

downstream_accuracy · Not assessed

Reported: 11 pp

downstream_accuracy · Not assessed

Reported: 41.6 pp

downstream_accuracy · Not assessed

Reported: 54.2 pp

downstream_accuracy · Not assessed

Reported: 24.7 pp

downstream_accuracy · Not assessed

Reported: 8.6 pp

downstream_accuracy · Not assessed

Reported: 29.2 pp

downstream_accuracy · Not assessed

Reported: 49 pp

downstream_accuracy · Not assessed

Reported: 18.4 pp

downstream_accuracy · Not assessed

Reported: 60.9 pp

downstream_accuracy · Not assessed

Reported: 63.6 pp

downstream_accuracy · Not assessed

Reported: 52.9 pp

downstream_accuracy · Not assessed

Reported: 38.7 pp

downstream_accuracy · Not assessed

Reported: 47.2 pp

Downstream task evaluation after post-training stage RLVR1 for 7b_sd.Experiment blocked; see the specific reasonReported 79.3 pp

Reported

79.3 pp

Observed

Intermediate checkpoint weights after post-training stage RLVR1 not released.

downstream_accuracy · Not assessed

Reported: 79.3 pp

downstream_accuracy · Not assessed

Reported: 63.4 pp

downstream_accuracy · Not assessed

Reported: 52.8 pp

downstream_accuracy · Not assessed

Reported: 36.1 pp

downstream_accuracy · Not assessed

Reported: 48.8 pp

downstream_accuracy · Not assessed

Reported: 17.4 pp

downstream_accuracy · Not assessed

Reported: 54.1 pp

downstream_accuracy · Not assessed

Reported: 23 pp

downstream_accuracy · Not assessed

Reported: 8 pp

downstream_accuracy · Not assessed

Reported: 47.2 pp

downstream_accuracy · Not assessed

Reported: 19.7 pp

downstream_accuracy · Not assessed

Reported: 61.7 pp

downstream_accuracy · Not assessed

Reported: 59.8 pp

downstream_accuracy · Not assessed

Reported: 53 pp

downstream_accuracy · Not assessed

Reported: 39.7 pp

downstream_accuracy · Not assessed

Reported: 68.8 pp

Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 13b_q20.Experiment blocked; see the specific reasonReported 66 pp

Reported

66 pp

Observed

Threshold sweep checkpoints not released.

downstream_accuracy · Not assessed

Reported: 66 pp

downstream_accuracy · Not assessed

Reported: 54.8 pp

downstream_accuracy · Not assessed

Reported: 42.2 pp

downstream_accuracy · Not assessed

Reported: 31.6 pp

downstream_accuracy · Not assessed

Reported: 45.8 pp

downstream_accuracy · Not assessed

Reported: 12 pp

downstream_accuracy · Not assessed

Reported: 42.1 pp

downstream_accuracy · Not assessed

Reported: 54.1 pp

downstream_accuracy · Not assessed

Reported: 24.9 pp

downstream_accuracy · Not assessed

Reported: 9 pp

downstream_accuracy · Not assessed

Reported: 29.3 pp

downstream_accuracy · Not assessed

Reported: 48.5 pp

downstream_accuracy · Not assessed

Reported: 17.7 pp

downstream_accuracy · Not assessed

Reported: 58.5 pp

downstream_accuracy · Not assessed

Reported: 63.8 pp

downstream_accuracy · Not assessed

Reported: 51.2 pp

downstream_accuracy · Not assessed

Reported: 39.2 pp

downstream_accuracy · Not assessed

Reported: 46.5 pp

Mid-training ablation results for Switch Distillation with FKL relative difference across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported -2.9 pp

Reported

-2.9 pp

Observed

Ablation checkpoints not released and training exceeds compute limits.

delta_percentage_points · Not assessed

Reported: -2.9 pp

delta_percentage_points · Not assessed

Reported: -0.2 pp

delta_percentage_points · Not assessed

Reported: -1.4 pp

Downstream task evaluation after post-training stage DPO for 13b_trkd.Experiment blocked; see the specific reasonReported 53.1 pp

Reported

53.1 pp

Observed

Intermediate checkpoint weights after post-training stage DPO not released.

downstream_accuracy · Not assessed

Reported: 53.1 pp

downstream_accuracy · Not assessed

Reported: 34.9 pp

downstream_accuracy · Not assessed

Reported: 30.7 pp

downstream_accuracy · Not assessed

Reported: 32.9 pp

downstream_accuracy · Not assessed

Reported: 35.3 pp

downstream_accuracy · Not assessed

Reported: 6.8 pp

downstream_accuracy · Not assessed

Reported: 54.9 pp

downstream_accuracy · Not assessed

Reported: 22.4 pp

downstream_accuracy · Not assessed

Reported: 8.5 pp

downstream_accuracy · Not assessed

Reported: 43.4 pp

downstream_accuracy · Not assessed

Reported: 16.5 pp

downstream_accuracy · Not assessed

Reported: 53.2 pp

downstream_accuracy · Not assessed

Reported: 56 pp

downstream_accuracy · Not assessed

Reported: 52 pp

downstream_accuracy · Not assessed

Reported: 35.3 pp

downstream_accuracy · Not assessed

Reported: 61.7 pp

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 7B Switch Distillation post-training performance.Experiment blocked; see the specific reasonReported 79.8 pp

Reported

79.8 pp

Observed

Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.

downstream_accuracy · Not assessed

Reported: 79.8 pp

downstream_accuracy · Not assessed

Reported: 65 pp

downstream_accuracy · Not assessed

Reported: 52.7 pp

downstream_accuracy · Not assessed

Reported: 35.6 pp

downstream_accuracy · Not assessed

Reported: 48.2 pp

downstream_accuracy · Not assessed

Reported: 22.4 pp

downstream_accuracy · Not assessed

Reported: 53.6 pp

downstream_accuracy · Not assessed

Reported: 22.7 pp

downstream_accuracy · Not assessed

Reported: 8.2 pp

downstream_accuracy · Not assessed

Reported: 48.1 pp

downstream_accuracy · Not assessed

Reported: 19.8 pp

downstream_accuracy · Not assessed

Reported: 62.6 pp

downstream_accuracy · Not assessed

Reported: 59.4 pp

downstream_accuracy · Not assessed

Reported: 54.6 pp

downstream_accuracy · Not assessed

Reported: 39.6 pp

downstream_accuracy · Not assessed

Reported: 69.5 pp

Reverse KL Distillation (RKD, alpha=0.5) downstream evaluation performance after mid-training with OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 62.1 pp

Reported

62.1 pp

Observed

Mid-trained checkpoint weights not released and mid-training exceeds compute limits.

downstream_accuracy · Not assessed

Reported: 62.1 pp

downstream_accuracy · Not assessed

Reported: 46.2 pp

downstream_accuracy · Not assessed

Reported: 39.2 pp

downstream_accuracy · Not assessed

Reported: 31.6 pp

downstream_accuracy · Not assessed

Reported: 43.1 pp

downstream_accuracy · Not assessed

Reported: 10.4 pp

downstream_accuracy · Not assessed

Reported: 53.1 pp

downstream_accuracy · Not assessed

Reported: 24.1 pp

downstream_accuracy · Not assessed

Reported: 8.4 pp

downstream_accuracy · Not assessed

Reported: 50.1 pp

downstream_accuracy · Not assessed

Reported: 18.9 pp

downstream_accuracy · Not assessed

Reported: 62.5 pp

downstream_accuracy · Not assessed

Reported: 62.8 pp

downstream_accuracy · Not assessed

Reported: 52.5 pp

downstream_accuracy · Not assessed

Reported: 39.3 pp

Downstream task evaluation after post-training stage SFT for 7b_fkd.Experiment blocked; see the specific reasonReported 54.7 pp

Reported

54.7 pp

Observed

Intermediate checkpoint weights after post-training stage SFT not released.

downstream_accuracy · Not assessed

Reported: 54.7 pp

downstream_accuracy · Not assessed

Reported: 34.2 pp

downstream_accuracy · Not assessed

Reported: 30.7 pp

downstream_accuracy · Not assessed

Reported: 31.3 pp

downstream_accuracy · Not assessed

Reported: 39.1 pp

downstream_accuracy · Not assessed

Reported: 9.4 pp

downstream_accuracy · Not assessed

Reported: 54.1 pp

downstream_accuracy · Not assessed

Reported: 23.3 pp

downstream_accuracy · Not assessed

Reported: 8.2 pp

downstream_accuracy · Not assessed

Reported: 49.1 pp

downstream_accuracy · Not assessed

Reported: 18.5 pp

downstream_accuracy · Not assessed

Reported: 61.7 pp

downstream_accuracy · Not assessed

Reported: 59.4 pp

downstream_accuracy · Not assessed

Reported: 51.9 pp

downstream_accuracy · Not assessed

Reported: 39.7 pp

downstream_accuracy · Not assessed

Reported: 46 pp

Downstream task evaluation after post-training stage RLVR1 for 13b_sd.Experiment blocked; see the specific reasonReported 78.8 pp

Reported

78.8 pp

Observed

Intermediate checkpoint weights after post-training stage RLVR1 not released.

downstream_accuracy · Not assessed

Reported: 78.8 pp

downstream_accuracy · Not assessed

Reported: 65.6 pp

downstream_accuracy · Not assessed

Reported: 52.5 pp

downstream_accuracy · Not assessed

Reported: 33.5 pp

downstream_accuracy · Not assessed

Reported: 43.3 pp

downstream_accuracy · Not assessed

Reported: 17.8 pp

downstream_accuracy · Not assessed

Reported: 54.3 pp

downstream_accuracy · Not assessed

Reported: 24 pp

downstream_accuracy · Not assessed

Reported: 8.2 pp

downstream_accuracy · Not assessed

Reported: 38.7 pp

downstream_accuracy · Not assessed

Reported: 17.7 pp

downstream_accuracy · Not assessed

Reported: 56.2 pp

downstream_accuracy · Not assessed

Reported: 57.8 pp

downstream_accuracy · Not assessed

Reported: 51.7 pp

downstream_accuracy · Not assessed

Reported: 37.7 pp

downstream_accuracy · Not assessed

Reported: 68.9 pp

Per-task downstream accuracy for mid-training ablation sd_fkl using OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 66.7 pp

Reported

66.7 pp

Observed

Per-task ablation checkpoints not released.

downstream_accuracy · Not assessed

Reported: 66.7 pp

downstream_accuracy · Not assessed

Reported: 52 pp

downstream_accuracy · Not assessed

Reported: 43.7 pp

downstream_accuracy · Not assessed

Reported: 31.2 pp

downstream_accuracy · Not assessed

Reported: 46.4 pp

downstream_accuracy · Not assessed

Reported: 11 pp

downstream_accuracy · Not assessed

Reported: 54.8 pp

downstream_accuracy · Not assessed

Reported: 24.3 pp

downstream_accuracy · Not assessed

Reported: 8.2 pp

downstream_accuracy · Not assessed

Reported: 50.5 pp

downstream_accuracy · Not assessed

Reported: 19.1 pp

downstream_accuracy · Not assessed

Reported: 63.7 pp

downstream_accuracy · Not assessed

Reported: 62.6 pp

downstream_accuracy · Not assessed

Reported: 52.2 pp

downstream_accuracy · Not assessed

Reported: 39.6 pp

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for 13B RKD post-training performance.Experiment blocked; see the specific reasonReported 75.1 pp

Reported

75.1 pp

Observed

Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.

downstream_accuracy · Not assessed

Reported: 75.1 pp

downstream_accuracy · Not assessed

Reported: 58.1 pp

downstream_accuracy · Not assessed

Reported: 49.3 pp

downstream_accuracy · Not assessed

Reported: 33.3 pp

downstream_accuracy · Not assessed

Reported: 38.7 pp

downstream_accuracy · Not assessed

Reported: 16 pp

downstream_accuracy · Not assessed

Reported: 53.9 pp

downstream_accuracy · Not assessed

Reported: 22.3 pp

downstream_accuracy · Not assessed

Reported: 7.7 pp

downstream_accuracy · Not assessed

Reported: 38.2 pp

downstream_accuracy · Not assessed

Reported: 17 pp

downstream_accuracy · Not assessed

Reported: 56 pp

downstream_accuracy · Not assessed

Reported: 55.2 pp

downstream_accuracy · Not assessed

Reported: 51.1 pp

downstream_accuracy · Not assessed

Reported: 36.6 pp

downstream_accuracy · Not assessed

Reported: 64.1 pp

Mid-training ablation results for Random Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported -6.5 pp

Reported

-6.5 pp

Observed

Ablation checkpoints not released and training exceeds compute limits.

delta_percentage_points · Not assessed

Reported: -6.5 pp

delta_percentage_points · Not assessed

Reported: -0.8 pp

delta_percentage_points · Not assessed

Reported: -2 pp

Downstream task evaluation after post-training stage RLVR1 for 7b_trkd.Experiment blocked; see the specific reasonReported 71 pp

Reported

71 pp

Observed

Intermediate checkpoint weights after post-training stage RLVR1 not released.

downstream_accuracy · Not assessed

Reported: 71 pp

downstream_accuracy · Not assessed

Reported: 54.6 pp

downstream_accuracy · Not assessed

Reported: 46.8 pp

downstream_accuracy · Not assessed

Reported: 33 pp

downstream_accuracy · Not assessed

Reported: 37.4 pp

downstream_accuracy · Not assessed

Reported: 11.4 pp

downstream_accuracy · Not assessed

Reported: 52.8 pp

downstream_accuracy · Not assessed

Reported: 22.7 pp

downstream_accuracy · Not assessed

Reported: 8 pp

downstream_accuracy · Not assessed

Reported: 40.2 pp

downstream_accuracy · Not assessed

Reported: 17.8 pp

downstream_accuracy · Not assessed

Reported: 58.9 pp

downstream_accuracy · Not assessed

Reported: 55.6 pp

downstream_accuracy · Not assessed

Reported: 52 pp

downstream_accuracy · Not assessed

Reported: 35.9 pp

downstream_accuracy · Not assessed

Reported: 65.2 pp

Per-task downstream accuracy for mid-training ablation teacher_correct using OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 63.8 pp

Reported

63.8 pp

Observed

Per-task ablation checkpoints not released.

downstream_accuracy · Not assessed

Reported: 63.8 pp

downstream_accuracy · Not assessed

Reported: 49.5 pp

downstream_accuracy · Not assessed

Reported: 41.7 pp

downstream_accuracy · Not assessed

Reported: 31.8 pp

downstream_accuracy · Not assessed

Reported: 45.2 pp

downstream_accuracy · Not assessed

Reported: 9.8 pp

downstream_accuracy · Not assessed

Reported: 44.5 pp

downstream_accuracy · Not assessed

Reported: 19.4 pp

downstream_accuracy · Not assessed

Reported: 8.7 pp

downstream_accuracy · Not assessed

Reported: 50.6 pp

downstream_accuracy · Not assessed

Reported: 19.3 pp

downstream_accuracy · Not assessed

Reported: 63.2 pp

downstream_accuracy · Not assessed

Reported: 64.6 pp

downstream_accuracy · Not assessed

Reported: 52 pp

downstream_accuracy · Not assessed

Reported: 39.9 pp

Downstream task performance and group macro-averages for Switch Distillation under routing threshold sweep 7b_q20.Experiment blocked; see the specific reasonReported 69.7 pp

Reported

69.7 pp

Observed

Threshold sweep checkpoints not released.

downstream_accuracy · Not assessed

Reported: 69.7 pp

downstream_accuracy · Not assessed

Reported: 55.3 pp

downstream_accuracy · Not assessed

Reported: 46.1 pp

downstream_accuracy · Not assessed

Reported: 32.8 pp

downstream_accuracy · Not assessed

Reported: 49.6 pp

downstream_accuracy · Not assessed

Reported: 14.8 pp

downstream_accuracy · Not assessed

Reported: 44.7 pp

downstream_accuracy · Not assessed

Reported: 54.9 pp

downstream_accuracy · Not assessed

Reported: 24.6 pp

downstream_accuracy · Not assessed

Reported: 8.4 pp

downstream_accuracy · Not assessed

Reported: 29.3 pp

downstream_accuracy · Not assessed

Reported: 51.6 pp

downstream_accuracy · Not assessed

Reported: 19.8 pp

downstream_accuracy · Not assessed

Reported: 64.7 pp

downstream_accuracy · Not assessed

Reported: 64.2 pp

downstream_accuracy · Not assessed

Reported: 53.8 pp

downstream_accuracy · Not assessed

Reported: 41.5 pp

downstream_accuracy · Not assessed

Reported: 49.3 pp

Full downstream results after standard 4-stage post-training pipeline (SFT, DPO, RLVR1, RLVR2) for NTP post-training baseline performance.Experiment blocked; see the specific reasonReported 67 pp

Reported

67 pp

Observed

Post-trained checkpoint weights from this mid-training pipeline are not released, and executing the full 4-stage post-training pipeline requires multi-node H200 compute beyond budget.

downstream_accuracy · Not assessed

Reported: 67 pp

downstream_accuracy · Not assessed

Reported: 46.8 pp

downstream_accuracy · Not assessed

Reported: 40.6 pp

downstream_accuracy · Not assessed

Reported: 31.1 pp

downstream_accuracy · Not assessed

Reported: 32.7 pp

downstream_accuracy · Not assessed

Reported: 12.6 pp

downstream_accuracy · Not assessed

Reported: 53.3 pp

downstream_accuracy · Not assessed

Reported: 21.6 pp

downstream_accuracy · Not assessed

Reported: 7.5 pp

downstream_accuracy · Not assessed

Reported: 32.3 pp

downstream_accuracy · Not assessed

Reported: 16.6 pp

downstream_accuracy · Not assessed

Reported: 52.4 pp

downstream_accuracy · Not assessed

Reported: 52.4 pp

downstream_accuracy · Not assessed

Reported: 51.4 pp

downstream_accuracy · Not assessed

Reported: 32.3 pp

downstream_accuracy · Not assessed

Reported: 62.1 pp

Mid-training ablation results for Teacher-Correct Routing relative difference across Reasoning, Factual Recall, and Knowledge task groups.Experiment blocked; see the specific reasonReported -4.4 pp

Reported

-4.4 pp

Observed

Ablation checkpoints not released and training exceeds compute limits.

delta_percentage_points · Not assessed

Reported: -4.4 pp

delta_percentage_points · Not assessed

Reported: -5.1 pp

delta_percentage_points · Not assessed

Reported: -1 pp

Per-task downstream accuracy for mid-training ablation always_ce using OLMo-2 7B Instruct teacher.Experiment blocked; see the specific reasonReported 69.1 pp

Reported

69.1 pp

Observed

Per-task ablation checkpoints not released.

downstream_accuracy · Not assessed

Reported: 69.1 pp

downstream_accuracy · Not assessed

Reported: 55.2 pp

downstream_accuracy · Not assessed

Reported: 46.5 pp

downstream_accuracy · Not assessed

Reported: 33.3 pp

downstream_accuracy · Not assessed

Reported: 50.1 pp

downstream_accuracy · Not assessed

Reported: 12 pp

downstream_accuracy · Not assessed

Reported: 54.3 pp

downstream_accuracy · Not assessed

Reported: 24.5 pp

downstream_accuracy · Not assessed

Reported: 9.4 pp

downstream_accuracy · Not assessed

Reported: 51.1 pp

downstream_accuracy · Not assessed

Reported: 19.9 pp

downstream_accuracy · Not assessed

Reported: 63.7 pp

downstream_accuracy · Not assessed

Reported: 63.6 pp

downstream_accuracy · Not assessed

Reported: 52.7 pp

downstream_accuracy · Not assessed

Reported: 41.1 pp

Downstream task evaluation after post-training stage RLVR1 for 13b_trkd.Experiment blocked; see the specific reasonReported 72.4 pp

Reported

72.4 pp

Observed

Intermediate checkpoint weights after post-training stage RLVR1 not released.

downstream_accuracy · Not assessed

Reported: 72.4 pp

downstream_accuracy · Not assessed

Reported: 53.3 pp

downstream_accuracy · Not assessed

Reported: 47.6 pp

downstream_accuracy · Not assessed

Reported: 32.7 pp

downstream_accuracy · Not assessed

Reported: 35 pp

downstream_accuracy · Not assessed

Reported: 10.4 pp

downstream_accuracy · Not assessed

Reported: 54.7 pp

downstream_accuracy · Not assessed

Reported: 22.2 pp

downstream_accuracy · Not assessed

Reported: 8.6 pp

downstream_accuracy · Not assessed

Reported: 36.6 pp

downstream_accuracy · Not assessed

Reported: 16.6 pp

downstream_accuracy · Not assessed

Reported: 54.1 pp

downstream_accuracy · Not assessed

Reported: 54.6 pp

downstream_accuracy · Not assessed

Reported: 51.7 pp

downstream_accuracy · Not assessed

Reported: 34.5 pp

downstream_accuracy · Not assessed

Reported: 67.1 pp

Downstream task evaluation after post-training stage DPO for 7b_fkd.Experiment blocked; see the specific reasonReported 60.4 pp

Reported

60.4 pp

Observed

Intermediate checkpoint weights after post-training stage DPO not released.

downstream_accuracy · Not assessed

Reported: 60.4 pp

downstream_accuracy · Not assessed

Reported: 35.8 pp

downstream_accuracy · Not assessed

Reported: 33.1 pp

downstream_accuracy · Not assessed

Reported: 33.1 pp

downstream_accuracy · Not assessed

Reported: 39.8 pp

downstream_accuracy · Not assessed

Reported: 7.8 pp

downstream_accuracy · Not assessed

Reported: 53.6 pp

downstream_accuracy · Not assessed

Reported: 23.1 pp

downstream_accuracy · Not assessed

Reported: 8.3 pp

downstream_accuracy · Not assessed

Reported: 49.2 pp

downstream_accuracy · Not assessed

Reported: 18.5 pp

downstream_accuracy · Not assessed

Reported: 59.2 pp

downstream_accuracy · Not assessed

Reported: 58 pp

downstream_accuracy · Not assessed

Reported: 51.6 pp

downstream_accuracy · Not assessed

Reported: 39.7 pp

downstream_accuracy · Not assessed

Reported: 64.1 pp

Qualitative findings

The stage-dependent distillation tradeoff generalizes to the SmolLM family (SmolLM2 360M student, 1.7B Instruct teacher, SmolLM3 Stage-3 data), where mid-training KD exhibits the reasoning-recall deficit and Switch Distillation mitigates it.Experiment blocked; see the specific reason

Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.

Forward KL consistently exhibits higher CE-KL gradient cosine alignment than Reverse KL across pre-training and mid-training, with the gap widening at higher alpha and later training steps.Experiment blocked; see the specific reason

Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.

Gold-token probability, gradient attenuation relative to NTP, and factual recall deficits under KD hold consistently for OLMo-2 1B and 13B Instruct teachers.Experiment blocked; see the specific reason

Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.

Switch Distillation substantially accelerates reasoning acquisition during mid-training, surpassing the final 60B-token NTP reasoning macro-average within 2.5B tokens (24x fewer tokens), while standard FKD requires 5.0B tokens (12x fewer).Experiment blocked; see the specific reason

Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.

Lower teacher predictive entropy corresponds to substantially higher teacher top-1 agreement with the ground-truth token across all data domains and teacher sizes.Experiment blocked; see the specific reason

Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.

Higher teacher entropy corresponds to strictly lower gold-token probability, attenuates the gold-token gradient relative to NTP (reaching ~0.5x NTP for Q5 facts), and produces larger downstream factual recall deficits.Experiment blocked; see the specific reason

Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.

Across distillation strengths alpha in [0.0, 1.0], KL directions (forward and reverse), and teacher sizes (1B, 7B, 13B), mid-training distillation traces a reasoning-recall frontier that falls below NTP, which Switch Distillation mitigates.Experiment blocked; see the specific reason

Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.

Knowledge distillation generally improves reasoning, but its effect on factual recall changes across training stages: during pre-training, distillation improves both reasoning and factual recall over NTP, whereas during mid-training it improves reasoning at the expense of factual recall.Experiment blocked; see the specific reason

Plot-only empirical finding without extractable discrete tabular coordinates; requires full intermediate training checkpoint trajectory or continuous sweep data not provided in the paper or repository.