LongPIBench: A Long-Context Benchmark for Prompt Injection

Public
Authors:Yupei Liu, Yuqi Jia, Neil Zhenqiang Gong, Jinyuan Jia
More
Copy repository link

Reproduction

Start with coverage and measured results; open a run only when you need evidence or technical details.

RunMatchRepeat

0/4

claims supported by evidence

0

Supported

0

Challenged or mixed

0

Contradicted

1

Inconclusive

3

Not assessed

Run history

Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.

Claim–experiment reproduction matrix

See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 1 claims with no independent reproduction scheduled in this plan.

1 targets·1 with evidence·0 active·0 need attention

Prevention-based defenses that perform well on short-context benchmarks retain substantial attack success on the paper's long-context synthetic benchmark, although PromptLocate and MetaSecAlign 8B are lower than several simpler defenses on some tasks.

Information insufficientclaim-prevention-degradation0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

The paper reports that detection-based defenses exhibit an extreme false-positive/false-negative trade-off on long-context inputs, with some methods flagging benign inputs and others missing attacks.

Information insufficientclaim-detection-tradeoff0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

The paper reports that GCG and its universal variant achieve high attack success across all four task suites and outperform heuristic attacks on several tasks.

Information insufficientclaim-optimization-attacks0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

On long-context document tasks, heuristic prompt-injection attacks substantially increase attack success rate over no-attack baselines, and Authority spoof is generally the strongest listed heuristic.

CiteArk reconstructionclaim-heuristic-attacks-effective1 plan1 run
Scientific conclusionInconclusive

Independent approximate paper-review Authority spoof ASR on controlled long contexts

exp-reconstruct-paper-review-authority2 attempts
Evidence published

Next step

The execution path is complete; inspect the scientific Assessment next.

View local evidence pathExpand
  1. Research plan

    Claim and experiment binding established

  2. Execution target

    Evidence published

  3. Run attempt

    Evidence publication · Completed

  4. CAP evidence

    1 immutable Artifact

  5. Scientific judgment

    Assessed but inconclusive

Legacy task without target-level resource requirements