Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration

Public
Authors:Xiaoqing Wang, Keman Huang, Bin Liang, Hongyu Li, Xiaoyong Du, Wuqiong Pan
More
Copy repository link

AI research summary

The paper introduces DoCtOR, a reflection framework for large-language-model multi-agent systems. DoCtOR uses a process-reward-based ProFA diagnosis module to identify the first incorrect reasoning step and responsible agent in a failed trajectory, a counterfactual correction module to propose and score an alternative action, and a reflector that gives targeted feedback to the decisive error agent. The authors evaluate the framework on HotPotQA, ChartQAPro, and Mind2Web, compare it with Reflexion, Retroformer, and COPPER, and separately evaluate ProFA against automated failure-attribution baselines on held-in and held-out Who & When data. Additional experiments study model size, test-set size, trajectory scope, correctness threshold, and component ablations.

Generated by CiteArk.

What this means

The work targets a practical failure mode in multi-agent systems: feedback can be wasted on agents that did not cause the failure, potentially confusing future behavior. If the reported pattern holds, diagnosing the first consequential mistake and concentrating correction on its owner could make reflective systems more useful under limited inference or training budgets, especially for multi-step question answering, chart analysis, and web interaction. The value is conditional on the paper’s task-specific prompts, model choices, external services, failure annotations, and evaluation setup, and the fixed paper does not provide a verified implementation snapshot for direct reproduction here.

Reproduction progress
RunMatchRepeat

Not reproduced yet

0/3

No verifiable claim has successful reproduction evidence yet.

Awaiting author or institution signature.

Research claims

The main empirical claims extracted in this research plan, with their plan coverage and live evidence state. Click a claim to expand it inline.

Across five trials, DoCtOR improves initial task success rates by 22% on HotPotQA, 26% on ChartQAPro, and 27% on Mind2Web, and outperforms the reported reflection baselines.Awaiting reproductionReported 22 pp

Reported

22 pp

Observed

Experiment plan ready; no runs yet.

initial_success_improvement · Not assessed

Reported: 22 pp

initial_success_improvement · Not assessed

Reported: 26 pp

initial_success_improvement · Not assessed

Reported: 27 pp

Platform or execution failure · no scientific conclusion

Required input unavailable · Execution will continue after the required code, data, or input is supplied.

ProFA achieves the strongest reported agent-level and step-level failure-attribution accuracies among the listed methods on the held-in HotPotQA, ChartQAPro, and Mind2Web datasets and on the held-out Algorithm-Generated and Hand-Crafted subsets.Awaiting reproductionReported 0.85 fraction

Reported

0.85 fraction

Observed

Experiment plan ready; no runs yet.

agent_level_accuracy · Not assessed

Reported: 0.85 fraction

step_level_accuracy · Not assessed

Reported: 0.53 fraction

step_level_accuracy · Not assessed

Reported: 0.2 fraction

Platform or execution failure · no scientific conclusion

Experiment execution error · No automatic retry; a person must decide what to do next.

Using only reasoning steps after the decisive error step produces reflection quality comparable to using the complete failure trajectory, while diagnosis, correction, and PPO each contribute to performance.Experiment blocked; see the specific reason

The claim is supported by plotted curves without machine-readable scalar values, and direct evaluation requires the unavailable datasets, model assets, external agent services, and exact implementation details.

The evaluation is limited to English-language task-oriented collaborations with relatively small multi-agent teams and leaves open performance on open-ended or culturally diverse tasks and larger teams.Stated by the authors

This is an explicit limitation note, not an independent empirical measurement requiring reproduction.

Reproduction and technical details2
Implementation path
Official implementation0

The experiment runs a repository and commit explicitly declared by the paper and verified by CiteArk.

CiteArk reconstruction2

No author code is used. CiteArk independently rebuilds the paper-grounded protocol and records every generated file and assumption.

Information insufficient1

A falsifiable execution still needs a missing method detail, input, split, metric, or evaluation rule.

Signed research planVerifiedThe Research Plan CAP fixes paper sources, complete Claim coverage, planned Experiments, and source-declared values.Download research plan
Generated byCiteArk Official
Modelopenai/gpt-5.6-luna
Completed
Artifact digest92f0d0878f
Attestation statusTrusted signature verified
Output
View independent verification JSON