Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
Not assessedPlannedFindingclaim-doctor-improves-success
Across five trials, DoCtOR improves initial task success rates by 22% on HotPotQA, 26% on ChartQAPro, and 27% on Mind2Web, and outperforms the reported reflection baselines.
Source: paper-fixed:Abstract; Sections 4.3-4.4; Figure 3; PDF pages 1 and 6
Reported and observed measurements
initial_success_improvement
reported-doctor-hotpotqa-improvement
Reported 22 percentage_points
Observed — percentage_points
initial_success_improvement
reported-doctor-chartqapro-improvement
Reported 26 percentage_points
Observed — percentage_points
initial_success_improvement
reported-doctor-mind2web-improvement
Reported 27 percentage_points
Observed — percentage_points
Assessments (0)
No immutable Assessment has been published for this Claim yet.
Runs (0)
No execution has been linked to this Claim yet.