AGENTIC R AG-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing
Not assessedPlan blockedFindingclaim-main-benchmark
AGENTIC R AG-R1 reports higher or near-highest F1 than listed baselines across seven multi-hop, open-domain, and agentic QA benchmarks.
Source: paper-fixed:page 7, Table 1
Reported and observed measurements
average F1
m-main-average
Reported 33.55 percentage_points
Observed — percentage_points
F1
m-main-hotpot
Reported 44 percentage_points
Observed — percentage_points
F1
m-main-7b-bamboogle
Reported 49.21 percentage_points
Observed — percentage_points
Assessments (0)
No immutable Assessment has been published for this Claim yet.
Runs (0)
No execution has been linked to this Claim yet.