Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
PublicMore
AI research summary
The paper introduces DoCtOR, a reflection framework for large-language-model multi-agent systems. DoCtOR uses a process-reward-based ProFA diagnosis module to identify the first incorrect reasoning step and responsible agent in a failed trajectory, a counterfactual correction module to propose and score an alternative action, and a reflector that gives targeted feedback to the decisive error agent. The authors evaluate the framework on HotPotQA, ChartQAPro, and Mind2Web, compare it with Reflexion, Retroformer, and COPPER, and separately evaluate ProFA against automated failure-attribution baselines on held-in and held-out Who & When data. Additional experiments study model size, test-set size, trajectory scope, correctness threshold, and component ablations.
Generated by CiteArk.
What this means
The work targets a practical failure mode in multi-agent systems: feedback can be wasted on agents that did not cause the failure, potentially confusing future behavior. If the reported pattern holds, diagnosing the first consequential mistake and concentrating correction on its owner could make reflective systems more useful under limited inference or training budgets, especially for multi-step question answering, chart analysis, and web interaction. The value is conditional on the paper’s task-specific prompts, model choices, external services, failure annotations, and evaluation setup, and the fixed paper does not provide a verified implementation snapshot for direct reproduction here.
Run · Match · Repeat
Run · Match · Repeat shows the deepest evidence stage reached by any claim in this repository, not paper-wide reproduction coverage.
An outer circle appears only after a paper author or trusted institution signs the Artifact.
Not reproduced yet
0/3
No verifiable claim has successful reproduction evidence yet.
Research claims
The main empirical claims extracted in this research plan, with their plan coverage and live evidence state. Click a claim to expand it inline.
Across five trials, DoCtOR improves initial task success rates by 22% on HotPotQA, 26% on ChartQAPro, and 27% on Mind2Web, and outperforms the reported reflection baselines.Awaiting reproductionReported 22 pp
Reported
22 pp
Observed
—
Experiment plan ready; no runs yet.
initial_success_improvement · Not assessed
Reported: 22 pp
initial_success_improvement · Not assessed
Reported: 26 pp
initial_success_improvement · Not assessed
Reported: 27 pp
Required input unavailable · Execution will continue after the required code, data, or input is supplied.
ProFA achieves the strongest reported agent-level and step-level failure-attribution accuracies among the listed methods on the held-in HotPotQA, ChartQAPro, and Mind2Web datasets and on the held-out Algorithm-Generated and Hand-Crafted subsets.Awaiting reproductionReported 0.85 fraction
Reported
0.85 fraction
Observed
—
Experiment plan ready; no runs yet.
agent_level_accuracy · Not assessed
Reported: 0.85 fraction
step_level_accuracy · Not assessed
Reported: 0.53 fraction
step_level_accuracy · Not assessed
Reported: 0.2 fraction
Experiment execution error · No automatic retry; a person must decide what to do next.
Using only reasoning steps after the decisive error step produces reflection quality comparable to using the complete failure trajectory, while diagnosis, correction, and PPO each contribute to performance.Experiment blocked; see the specific reason
The claim is supported by plotted curves without machine-readable scalar values, and direct evaluation requires the unavailable datasets, model assets, external agent services, and exact implementation details.
The evaluation is limited to English-language task-oriented collaborations with relatively small multi-agent teams and leaves open performance on open-ended or culturally diverse tasks and larger teams.Stated by the authors
This is an explicit limitation note, not an independent empirical measurement requiring reproduction.
Reproduction and technical details2
The experiment runs a repository and commit explicitly declared by the paper and verified by CiteArk.
No author code is used. CiteArk independently rebuilds the paper-grounded protocol and records every generated file and assumption.
A falsifiable execution still needs a missing method detail, input, split, metric, or evaluation rule.
Signed research planVerifiedThe Research Plan CAP fixes paper sources, complete Claim coverage, planned Experiments, and source-declared values.Download research plan
92f0d0878f