Across five trials, DoCtOR improves initial task success rates by 22% on HotPotQA, 26% on ChartQAPro, and 27% on Mind2Web, and outperforms the reported reflection baselines.
Independent DoCtOR success-rate reconstruction
Next step
Execution will continue after the required code, data, or input is supplied.
View local evidence pathExpandCollapse
- Research plan
Claim and experiment binding established
- Execution target
Execution failed
- Run attempt
Experiment execution · Failed
- CAP evidence
No linked immutable Artifact yet
- Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements