Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration

Public
Authors:Xiaoqing Wang, Keman Huang, Bin Liang, Hongyu Li, Xiaoyong Du, Wuqiong Pan
More
Copy repository link

Reproduction

Start with coverage and measured results; open a run only when you need evidence or technical details.

RunMatchRepeat

0/3

claims supported by evidence

0

Supported

0

Challenged or mixed

0

Contradicted

0

Inconclusive

3

Not assessed

Run history

Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.

No reproduction runs yet

Runs will appear here as the paper's experiment plans are executed.

Claim–experiment reproduction matrix

See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 1 claims with no independent reproduction scheduled in this plan.

2 targets·0 with evidence·0 active·2 need attention

Across five trials, DoCtOR improves initial task success rates by 22% on HotPotQA, 26% on ChartQAPro, and 27% on Mind2Web, and outperforms the reported reflection baselines.

CiteArk reconstructionclaim-doctor-improves-success1 plan0 runs
Scientific conclusionNot assessed

Independent DoCtOR success-rate reconstruction

exp-doctor-success-reconstruction3 attempts
Execution failed

Next step

Execution will continue after the required code, data, or input is supplied.

Required input unavailable·legacy.input_unavailable
View local evidence pathExpand
  1. Research plan

    Claim and experiment binding established

  2. Execution target

    Execution failed

  3. Run attempt

    Experiment execution · Failed

  4. CAP evidence

    No linked immutable Artifact yet

  5. Scientific judgment

    No Assessment yet

Legacy task without target-level resource requirements

ProFA achieves the strongest reported agent-level and step-level failure-attribution accuracies among the listed methods on the held-in HotPotQA, ChartQAPro, and Mind2Web datasets and on the held-out Algorithm-Generated and Hand-Crafted subsets.

CiteArk reconstructionclaim-profa-attribution1 plan0 runs
Scientific conclusionNot assessed

Independent ProFA failure-attribution reconstruction

exp-profa-attribution-reconstruction1 attempt
Execution failed

Next step

No automatic retry; a person must decide what to do next.

Experiment execution error·legacy.execution_error
View local evidence pathExpand
  1. Research plan

    Claim and experiment binding established

  2. Execution target

    Execution failed

  3. Run attempt

    Experiment execution · Failed

  4. CAP evidence

    No linked immutable Artifact yet

  5. Scientific judgment

    No Assessment yet

Legacy task without target-level resource requirements

Using only reasoning steps after the decisive error step produces reflection quality comparable to using the complete failure trajectory, while diagnosis, correction, and PPO each contribute to performance.

Information insufficientclaim-targeted-reflection-scope0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.