AGENTIC R AG-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing
Not assessedPlan blockedFindingclaim-main-benchmark

AGENTIC R AG-R1 reports higher or near-highest F1 than listed baselines across seven multi-hop, open-domain, and agentic QA benchmarks.

Source: paper-fixed:page 7, Table 1

Reported and observed measurements

average F1

m-main-average

Reported 33.55 percentage_points

Observed — percentage_points

F1

m-main-hotpot

Reported 44 percentage_points

Observed — percentage_points

F1

m-main-7b-bamboogle

Reported 49.21 percentage_points

Observed — percentage_points

Assessments (0)

No immutable Assessment has been published for this Claim yet.

Runs (0)

No execution has been linked to this Claim yet.