Reproduction outcomes

2

The latest reproduction results, one card per paper, ordered by the most recent Assessment. Skim the outcome here; open the paper's claims tab for the full evidence.

2 papers

Filters

Subjects

Computer Science

Outcome

Result updated

The paper claims that teacher predictive entropy distinguishes procedural from knowledge-intensive domains with high ROC AUC across training stages and model sizes. Our reproduction evaluated only the 1B checkpoints on a reconstructed 240-document benchmark; the 7B and 13B measurements were omitted due to size/budget limits. For the evaluated 1B results, observed ROC AUC values (e.g., 0.8079 vs reported 0.815) were broadly consistent, but the claim remains inconclusive due to incomplete multi-model and multi-stage coverage.

Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).roc_auc0.8150.8079↓ 0.9%
Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).roc_auc0.770.7967↑ 3.5%
Knowledge Distillation During Mid-Training Favors Reasoning over Factual RecallComputation and LanguageArtificial IntelligenceMachine Learning (cs.LG)

0/ 76 reproduced

Inconclusive1Not yet assessed75
Run · Match · Repeat

The paper introduces LongPIBench and reports key claims: prevention defenses show high attack success on long contexts, detection shows extreme trade-offs, GCG variants achieve high success, scope is limited to static workflows, and heuristic attacks, especially Authority spoof, substantially raise success. Our reproduction, currently inconclusive, allowed only a directional check of heuristics on document tasks: reported 1 vs observed 1 fraction attack success rate, with limitations from approximated model, corpus, and generation preventing strict matching. Other claims remain not assessed.

On long-context document tasks, heuristic prompt-injection attacks substantially increase attack success rate over no-attack baselines, and Authority spoof is generally the strongest listed heuristic.attack_success_rate11±0
LongPIBench: A Long-Context Benchmark for Prompt InjectionCryptography and SecurityComputation and LanguageMachine Learning (cs.LG)

0/ 4 reproduced

Inconclusive1Not yet assessed4
Run · Match · Repeat