Reproduction outcomes
2The latest reproduction results, one card per paper, ordered by the most recent Assessment. Skim the outcome here; open the paper's claims tab for the full evidence.
2 papers
Filters
Subjects
Computer Science
Outcome
Result updated
The paper claims that teacher predictive entropy distinguishes procedural from knowledge-intensive domains with high ROC AUC across training stages and model sizes. Our reproduction evaluated only the 1B checkpoints on a reconstructed 240-document benchmark; the 7B and 13B measurements were omitted due to size/budget limits. For the evaluated 1B results, observed ROC AUC values (e.g., 0.8079 vs reported 0.815) were broadly consistent, but the claim remains inconclusive due to incomplete multi-model and multi-stage coverage.
0/ 76 reproduced
The paper introduces LongPIBench and reports key claims: prevention defenses show high attack success on long contexts, detection shows extreme trade-offs, GCG variants achieve high success, scope is limited to static workflows, and heuristic attacks, especially Authority spoof, substantially raise success. Our reproduction, currently inconclusive, allowed only a directional check of heuristics on document tasks: reported 1 vs observed 1 fraction attack success rate, with limitations from approximated model, corpus, and generation preventing strict matching. Other claims remain not assessed.
0/ 4 reproduced