Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
尚未评估已列入计划研究发现claim-doctor-improves-success
Across five trials, DoCtOR improves initial task success rates by 22% on HotPotQA, 26% on ChartQAPro, and 27% on Mind2Web, and outperforms the reported reflection baselines.
来源:paper-fixed:Abstract; Sections 4.3-4.4; Figure 3; PDF pages 1 and 6
报告指标与观测值
initial_success_improvement
reported-doctor-hotpotqa-improvement
论文报告 22 percentage_points
实际观测 — percentage_points
initial_success_improvement
reported-doctor-chartqapro-improvement
论文报告 26 percentage_points
实际观测 — percentage_points
initial_success_improvement
reported-doctor-mind2web-improvement
论文报告 27 percentage_points
实际观测 — percentage_points
Assessment(0)
这条 Claim 暂无已发布的不可变 Assessment。
关联运行(0)
这条 Claim 暂无关联执行。