复现结果

2

按最新 Assessment 时间排列的复现结果,一篇论文一张卡。在这里速读结论;完整证据见论文的研究结论页。

2 篇论文

筛选

学科

计算机科学

结论

结果更新

The paper claims that teacher predictive entropy distinguishes procedural from knowledge-intensive domains with high ROC AUC across training stages and model sizes. Our reproduction evaluated only the 1B checkpoints on a reconstructed 240-document benchmark; the 7B and 13B measurements were omitted due to size/budget limits. For the evaluated 1B results, observed ROC AUC values (e.g., 0.8079 vs reported 0.815) were broadly consistent, but the claim remains inconclusive due to incomplete multi-model and multi-stage coverage.

Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).roc_auc0.8150.8079↓ 0.9%
Teacher predictive entropy distinguishes procedural domains (Math, FLAN) from knowledge-intensive domains (DCLM, Wikipedia, StackExchange, PeS2o) with high ROC AUC (0.744-0.826) across training stages (Base, SFT, DPO, Instruct) and model sizes (1B, 7B, 13B).roc_auc0.770.7967↑ 3.5%
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall计算与语言人工智能机器学习 (cs.LG)

0/ 76 已复现

无法判定1待评估75
Run · Match · Repeat

The paper introduces LongPIBench and reports key claims: prevention defenses show high attack success on long contexts, detection shows extreme trade-offs, GCG variants achieve high success, scope is limited to static workflows, and heuristic attacks, especially Authority spoof, substantially raise success. Our reproduction, currently inconclusive, allowed only a directional check of heuristics on document tasks: reported 1 vs observed 1 fraction attack success rate, with limitations from approximated model, corpus, and generation preventing strict matching. Other claims remain not assessed.

On long-context document tasks, heuristic prompt-injection attacks substantially increase attack success rate over no-attack baselines, and Authority spoof is generally the strongest listed heuristic.attack_success_rate11±0
LongPIBench: A Long-Context Benchmark for Prompt Injection密码学与安全计算与语言机器学习 (cs.LG)

0/ 4 已复现

无法判定1待评估4
Run · Match · Repeat