LongPIBench: A Long-Context Benchmark for Prompt Injection

Public
Authors:Yupei Liu, Yuqi Jia, Neil Zhenqiang Gong, Jinyuan Jia
More
Copy repository link

AI research summary

LongPIBench introduces a benchmark for prompt-injection attacks and defenses in four document-centric long-context workflows: paper peer review, resume screening, email summarization, and code review. It combines synthetic and real-world datasets, evaluates heuristic and optimization-based attacks across eight language models, and measures both prevention-based attack success and detection-based false-positive and false-negative rates. The paper reports that attacks remain effective and that defenses that perform well on short-context benchmarks often degrade on long documents, with additional ablations over document format, injection position, attack goals, context length, and segmentation.

Generated by CiteArk.

What this means

The work targets a practical security gap: an assistant can be given a long document containing untrusted text and quietly follow an embedded instruction instead of the user’s task. This benchmark can help model developers and deployers test review, hiring, email, and code workflows under more realistic document lengths, where simple defenses may look reliable on short inputs but fail in use. Its value is conditional on reproducing the task-specific data, model versions, attack criteria, and defense pipelines; the paper itself also limits the scope to static document processing rather than multi-step agents and does not identify why context length causes each degradation.

Reproduction progress
RunMatchRepeat

Not reproduced yet

0/4

No verifiable claim has successful reproduction evidence yet.

Awaiting author or institution signature.

Research claims

The main empirical claims extracted in this research plan, with their plan coverage and live evidence state. Click a claim to expand it inline.

On long-context document tasks, heuristic prompt-injection attacks substantially increase attack success rate over no-attack baselines, and Authority spoof is generally the strongest listed heuristic.Inconclusive0/1 measurements supported · 1 assessedReported 1 fraction1 fraction

Reported

1 fraction

Observed

1 fraction

Delta

±0 fraction

The available execution evidence is insufficient to confirm or challenge the paper's claim. · This does not refute the paper's claim; it means the available evidence can neither confirm nor refute it yet.

attack_success_rate · Inconclusive

Reported: 1 fraction

Observed: 1 fraction

The observed value is recorded as directional evidence, but approximate reconstruction fidelity cannot establish a strict Match against the paper

Open claim details
Prevention-based defenses that perform well on short-context benchmarks retain substantial attack success on the paper's long-context synthetic benchmark, although PromptLocate and MetaSecAlign 8B are lower than several simpler defenses on some tasks.Experiment blocked; see the specific reasonReported 1 fraction

Reported

1 fraction

Observed

The closest executable candidate is a public or independently coded prevention defense on newly generated long documents. It would replace the report-defining MetaSecAlign 8B checkpoint, pipeline, and synthetic panel, making the result proxy evidence rather than the reported ASR; the fixed paper provides none of those executable assets.

attack_success_rate · Not assessed

Reported: 1 fraction

The paper reports that detection-based defenses exhibit an extreme false-positive/false-negative trade-off on long-context inputs, with some methods flagging benign inputs and others missing attacks.Experiment blocked; see the specific reasonReported 0.43 fraction

Reported

0.43 fraction

Observed

The closest executable candidate is a public detector on newly generated matched benign and attacked panels with paper-like segment sizes. It would replace the named detector implementations, checkpoints, exact panel, and unspecified segmentation aggregation, producing proxy detector behavior rather than the reported comparison; those inputs are unavailable.

false_positive_rate · Not assessed

Reported: 0.43 fraction

The paper reports that GCG and its universal variant achieve high attack success across all four task suites and outperform heuristic attacks on several tasks.Experiment blocked; see the specific reasonReported 1 fraction

Reported

1 fraction

Observed

The closest executable candidate is an independent GCG loop against a public checkpoint on newly generated long documents. It would substitute the paper's target model, generated panel, GCG defaults, and attacker target construction; those material changes can alter ASR, so the candidate is proxy evidence rather than a comparable reconstruction and cannot satisfy this measurement.

attack_success_rate · Not assessed

Reported: 1 fraction

The benchmark evaluates static document-centric workflows in which the full document is supplied in one inference call and does not cover dynamic multi-step agentic workflows or the full range of automated attacks.Stated by the authors

This is an explicitly stated limitation rather than an independent empirical measurement.

Reproduction and technical details1
Implementation path
Official implementation0

The experiment runs a repository and commit explicitly declared by the paper and verified by CiteArk.

CiteArk reconstruction1

No author code is used. CiteArk independently rebuilds the paper-grounded protocol and records every generated file and assumption.

Information insufficient3

A falsifiable execution still needs a missing method detail, input, split, metric, or evaluation rule.

Signed research planVerifiedThe Research Plan CAP fixes paper sources, complete Claim coverage, planned Experiments, and source-declared values.Download research plan
Generated byCiteArk Official
Modelopenai/gpt-5.6-luna
Completed
Artifact digest8ce5c6f7ca
Attestation statusTrusted signature verified
Output
View independent verification JSON