AGENTIC R AG-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing

公开
作者:Xinke JiangYue FangZhibang YangJiaran GaoZhixin ZhangTao FengRihong QiuWentao ZhangHongxin DingRuizhe ZhangYongxin XuYuheng HuangXu ChuJunfeng ZhaoYasha Wang
更多
复制仓库链接

复现

优先展示覆盖情况与实测结果;需要审计时再打开运行记录和技术细节。

RunMatchRepeat

0/4

条结论获得证据支持

0

获得支持

0

受到挑战或冲突

0

遭到反驳

0

无法判定

4

尚未评估

运行记录

每一行是一条真实执行记录;命令、日志、哈希和签名都收纳在详情中。

尚无复现运行

论文中的实验方案开始执行后,运行记录会显示在这里。

结论—实验复现矩阵

逐个实验展示当前执行状态、阻塞原因、恢复动作和证据去向;技术执行与科研结论始终分开。

当前任务创建于目标级记录上线之前。下方状态来自历史任务的保守投影,不会伪造运行尝试、资源决策或证据关系。

AGENTIC R AG-R1 reports higher or near-highest F1 than listed baselines across seven multi-hop, open-domain, and agentic QA benchmarks.

信息不足claim-main-benchmark1 个方案0 次运行
科学结论尚未评估

Optional independent reconstruction of the main QA benchmark

实验方案已经存在,但当前处理任务没有对应的执行目标。

尚未进入任务

The paper reports increasing AGENTIC R AG-R1 F1 as the maximum reasoning-step budget grows on selected 2Wiki, Bamboogle, and TriviaQA evaluations.

信息不足claim-long-horizon1 个方案0 次运行
科学结论尚未评估

Optional independent reconstruction of step-budget scaling

实验方案已经存在,但当前处理任务没有对应的执行目标。

尚未进入任务

The paper reports lower average inference time than TC-RAG and downstream results on MedQA, DeepResearch Bench, ALFWorld, and WebShop.

信息不足claim-efficiency-generalization0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

The paper reports that removing memory actions, rollout rejection, or both lowers average F1, and that removing both retrieval and memory rewards gives the lowest ablation average.

信息不足claim-component-ablation1 个方案0 次运行
科学结论尚未评估

Optional independent reconstruction of component ablations

实验方案已经存在,但当前处理任务没有对应的执行目标。

尚未进入任务