LongPIBench: A Long-Context Benchmark for Prompt Injection
PublicMore
Research claims
0 / 5 claims verified
the rest still being verified
Key metricsPaper-reported → reproduced · click a row to expand
On long-context document tasks, heuristic prompt-injection attacks substantially increase attack success rate over no-attack baselines, and Authority spoof is generally the strongest listed heuristic.Inconclusive0/1 measurements supported · 1 assessedReported 1 fraction1 fraction
Reported
1 fraction
Observed
1 fraction
Delta
±0 fraction
The available execution evidence is insufficient to confirm or challenge the paper's claim. · This does not refute the paper's claim; it means the available evidence can neither confirm nor refute it yet.
attack_success_rate · Inconclusive
Reported: 1 fraction
Observed: 1 fraction
The observed value is recorded as directional evidence, but approximate reconstruction fidelity cannot establish a strict Match against the paper
Prevention-based defenses that perform well on short-context benchmarks retain substantial attack success on the paper's long-context synthetic benchmark, although PromptLocate and MetaSecAlign 8B are lower than several simpler defenses on some tasks.Experiment blocked; see the specific reasonReported 1 fraction
Reported
1 fraction
Observed
—
The closest executable candidate is a public or independently coded prevention defense on newly generated long documents. It would replace the report-defining MetaSecAlign 8B checkpoint, pipeline, and synthetic panel, making the result proxy evidence rather than the reported ASR; the fixed paper provides none of those executable assets.
attack_success_rate · Not assessed
Reported: 1 fraction
The paper reports that detection-based defenses exhibit an extreme false-positive/false-negative trade-off on long-context inputs, with some methods flagging benign inputs and others missing attacks.Experiment blocked; see the specific reasonReported 0.43 fraction
Reported
0.43 fraction
Observed
—
The closest executable candidate is a public detector on newly generated matched benign and attacked panels with paper-like segment sizes. It would replace the named detector implementations, checkpoints, exact panel, and unspecified segmentation aggregation, producing proxy detector behavior rather than the reported comparison; those inputs are unavailable.
false_positive_rate · Not assessed
Reported: 0.43 fraction
The paper reports that GCG and its universal variant achieve high attack success across all four task suites and outperform heuristic attacks on several tasks.Experiment blocked; see the specific reasonReported 1 fraction
Reported
1 fraction
Observed
—
The closest executable candidate is an independent GCG loop against a public checkpoint on newly generated long documents. It would substitute the paper's target model, generated panel, GCG defaults, and attacker target construction; those material changes can alter ASR, so the candidate is proxy evidence rather than a comparable reconstruction and cannot satisfy this measurement.
attack_success_rate · Not assessed
Reported: 1 fraction
Other claims1
The benchmark evaluates static document-centric workflows in which the full document is supplied in one inference call and does not cover dynamic multi-step agentic workflows or the full range of automated attacks.Stated by the authors
This is an explicitly stated limitation rather than an independent empirical measurement.