Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration

公开
作者:Xiaoqing WangKeman HuangBin LiangHongyu LiXiaoyong DuWuqiong Pan
更多
复制仓库链接

复现

优先展示覆盖情况与实测结果;需要审计时再打开运行记录和技术细节。

RunMatchRepeat

0/3

条结论获得证据支持

0

获得支持

0

受到挑战或冲突

0

遭到反驳

0

无法判定

3

尚未评估

运行记录

每一行是一条真实执行记录;命令、日志、哈希和签名都收纳在详情中。

尚无复现运行

论文中的实验方案开始执行后,运行记录会显示在这里。

结论—实验复现矩阵

逐个实验展示当前执行状态、阻塞原因、恢复动作和证据去向;技术执行与科研结论始终分开。 另有 1 条结论本次未安排独立复现。

2 个目标·0 个已有证据·0 个进行中·2 个需处理

Across five trials, DoCtOR improves initial task success rates by 22% on HotPotQA, 26% on ChartQAPro, and 27% on Mind2Web, and outperforms the reported reflection baselines.

CiteArk 独立重建claim-doctor-improves-success1 个方案0 次运行
科学结论尚未评估

Independent DoCtOR success-rate reconstruction

exp-doctor-success-reconstruction3 次尝试
执行失败

下一步

等待所需代码、数据或输入补齐后继续。

所需输入不可用·legacy.input_unavailable
查看局部证据路径展开
  1. 研究计划

    结论与实验绑定已建立

  2. 执行目标

    执行失败

  3. 运行尝试

    实验执行 · 失败

  4. CAP 证据

    尚无已关联的不可变 Artifact

  5. 科学判断

    尚无 Assessment

历史任务未记录目标级资源要求

ProFA achieves the strongest reported agent-level and step-level failure-attribution accuracies among the listed methods on the held-in HotPotQA, ChartQAPro, and Mind2Web datasets and on the held-out Algorithm-Generated and Hand-Crafted subsets.

CiteArk 独立重建claim-profa-attribution1 个方案0 次运行
科学结论尚未评估

Independent ProFA failure-attribution reconstruction

exp-profa-attribution-reconstruction1 次尝试
执行失败

下一步

不会自动重试,需要人工决定后续处理。

实验执行出错·legacy.execution_error
查看局部证据路径展开
  1. 研究计划

    结论与实验绑定已建立

  2. 执行目标

    执行失败

  3. 运行尝试

    实验执行 · 失败

  4. CAP 证据

    尚无已关联的不可变 Artifact

  5. 科学判断

    尚无 Assessment

历史任务未记录目标级资源要求

Using only reasoning steps after the decisive error step produces reflection quality comparable to using the complete failure trajectory, while diagnosis, correction, and PPO each contribute to performance.

信息不足claim-targeted-reflection-scope0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。