Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
Not assessedPlannedFindingclaim-doctor-improves-success

Across five trials, DoCtOR improves initial task success rates by 22% on HotPotQA, 26% on ChartQAPro, and 27% on Mind2Web, and outperforms the reported reflection baselines.

Source: paper-fixed:Abstract; Sections 4.3-4.4; Figure 3; PDF pages 1 and 6

Reported and observed measurements

initial_success_improvement

reported-doctor-hotpotqa-improvement

Reported 22 percentage_points

Observed — percentage_points

initial_success_improvement

reported-doctor-chartqapro-improvement

Reported 26 percentage_points

Observed — percentage_points

initial_success_improvement

reported-doctor-mind2web-improvement

Reported 27 percentage_points

Observed — percentage_points

Assessments (0)

No immutable Assessment has been published for this Claim yet.

Runs (0)

No execution has been linked to this Claim yet.