Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration

公开
作者:Xiaoqing WangKeman HuangBin LiangHongyu LiXiaoyong DuWuqiong Pan
更多
复制仓库链接

AI 研究摘要

The paper introduces DoCtOR, a reflection framework for large-language-model multi-agent systems. DoCtOR uses a process-reward-based ProFA diagnosis module to identify the first incorrect reasoning step and responsible agent in a failed trajectory, a counterfactual correction module to propose and score an alternative action, and a reflector that gives targeted feedback to the decisive error agent. The authors evaluate the framework on HotPotQA, ChartQAPro, and Mind2Web, compare it with Reflexion, Retroformer, and COPPER, and separately evaluate ProFA against automated failure-attribution baselines on held-in and held-out Who & When data. Additional experiments study model size, test-set size, trajectory scope, correctness threshold, and component ablations.

由 CiteArk 生成

这意味着什么

The work targets a practical failure mode in multi-agent systems: feedback can be wasted on agents that did not cause the failure, potentially confusing future behavior. If the reported pattern holds, diagnosing the first consequential mistake and concentrating correction on its owner could make reflective systems more useful under limited inference or training budgets, especially for multi-step question answering, chart analysis, and web interaction. The value is conditional on the paper’s task-specific prompts, model choices, external services, failure annotations, and evaluation setup, and the fixed paper does not provide a verified implementation snapshot for direct reproduction here.

复现进展
RunMatchRepeat

尚未复现成功

0/3

当前还没有可验证结论获得成功复现证据。

等待作者或机构签名确认。

研究结论

本次从论文中抽取的主要经验性结论,以及它们的计划覆盖和实时证据状态。点击结论就地展开详情。

Across five trials, DoCtOR improves initial task success rates by 22% on HotPotQA, 26% on ChartQAPro, and 27% on Mind2Web, and outperforms the reported reflection baselines.等待复现报告 22 个百分点

报告

22 个百分点

观测

实验方案已生成,还没有运行记录

initial_success_improvement · 尚未评估

报告: 22 个百分点

initial_success_improvement · 尚未评估

报告: 26 个百分点

initial_success_improvement · 尚未评估

报告: 27 个百分点

平台或执行故障 · 未形成科研结论

所需输入不可用 · 等待所需代码、数据或输入补齐后继续。

ProFA achieves the strongest reported agent-level and step-level failure-attribution accuracies among the listed methods on the held-in HotPotQA, ChartQAPro, and Mind2Web datasets and on the held-out Algorithm-Generated and Hand-Crafted subsets.等待复现报告 0.85 fraction

报告

0.85 fraction

观测

实验方案已生成,还没有运行记录

agent_level_accuracy · 尚未评估

报告: 0.85 fraction

step_level_accuracy · 尚未评估

报告: 0.53 fraction

step_level_accuracy · 尚未评估

报告: 0.2 fraction

平台或执行故障 · 未形成科研结论

实验执行出错 · 不会自动重试,需要人工决定后续处理。

Using only reasoning steps after the decisive error step produces reflection quality comparable to using the complete failure trajectory, while diagnosis, correction, and PPO each contribute to performance.实验受阻,详见具体原因

The claim is supported by plotted curves without machine-readable scalar values, and direct evaluation requires the unavailable datasets, model assets, external agent services, and exact implementation details.

The evaluation is limited to English-language task-oriented collaborations with relatively small multi-agent teams and leaves open performance on open-ended or culturally diverse tasks and larger teams.论文自述边界

This is an explicit limitation note, not an independent empirical measurement requiring reproduction.

复现与技术信息2
实现路径
官方实现0

实验运行论文明确声明、且经 CiteArk 核验的代码仓库与固定版本。

CiteArk 独立重建2

不使用作者代码;CiteArk 根据论文独立实现实验协议,并记录所有生成文件和假设。

信息不足1

要形成可证伪实验,仍缺少关键方法细节、输入、数据划分、指标或评估规则。

已签名研究计划已验证Research Plan CAP 固定论文来源、完整 Claim 覆盖、计划实验与论文声明值。下载研究计划
生成方CiteArk 官方
模型openai/gpt-5.6-luna
完成时间
Artifact 摘要92f0d0878f
签名状态可信签名已验证
产出
查看独立验签 JSON