AGENTIC R AG-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing

Public
Authors:Xinke Jiang, Yue Fang, Zhibang Yang, Jiaran Gao, Zhixin Zhang, Tao Feng, Rihong Qiu, Wentao Zhang, Hongxin Ding, Ruizhe Zhang, Yongxin Xu, Yuheng Huang, Xu Chu, Junfeng Zhao, Yasha Wang
More
Copy repository link

Reproduction

Start with coverage and measured results; open a run only when you need evidence or technical details.

RunMatchRepeat

0/4

claims supported by evidence

0

Supported

0

Challenged or mixed

0

Contradicted

0

Inconclusive

4

Not assessed

Run history

Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.

No reproduction runs yet

Runs will appear here as the paper's experiment plans are executed.

Claim–experiment reproduction matrix

See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions.

This task predates target-level records. The states below are conservative projections; no attempts, resource decisions, or evidence links are invented.

AGENTIC R AG-R1 reports higher or near-highest F1 than listed baselines across seven multi-hop, open-domain, and agentic QA benchmarks.

Information insufficientclaim-main-benchmark1 plan0 runs
Scientific conclusionNot assessed

Optional independent reconstruction of the main QA benchmark

An experiment plan exists, but the current processing task has no matching execution target.

Not in current task

The paper reports increasing AGENTIC R AG-R1 F1 as the maximum reasoning-step budget grows on selected 2Wiki, Bamboogle, and TriviaQA evaluations.

Information insufficientclaim-long-horizon1 plan0 runs
Scientific conclusionNot assessed

Optional independent reconstruction of step-budget scaling

An experiment plan exists, but the current processing task has no matching execution target.

Not in current task

The paper reports lower average inference time than TC-RAG and downstream results on MedQA, DeepResearch Bench, ALFWorld, and WebShop.

Information insufficientclaim-efficiency-generalization0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

The paper reports that removing memory actions, rollout rejection, or both lowers average F1, and that removing both retrieval and memory rewards gives the lowest ablation average.

Information insufficientclaim-component-ablation1 plan0 runs
Scientific conclusionNot assessed

Optional independent reconstruction of component ablations

An experiment plan exists, but the current processing task has no matching execution target.

Not in current task