AGENTIC R AG-R1 reports higher or near-highest F1 than listed baselines across seven multi-hop, open-domain, and agentic QA benchmarks.
Optional independent reconstruction of the main QA benchmark
An experiment plan exists, but the current processing task has no matching execution target.