LoGo: Token-Level Dynamic Local-Global Attention

公开
作者:Yuqi PanZheng LiBohao TangZhen QinGuoqi Li
更多
复制仓库链接

复现

优先展示覆盖情况与实测结果;需要审计时再打开运行记录和技术细节。

RunMatchRepeat

0/6

条结论获得证据支持

0

获得支持

0

受到挑战或冲突

0

遭到反驳

0

无法判定

6

尚未评估

运行记录

每一行是一条真实执行记录;命令、日志、哈希和签名都收纳在详情中。

尚无复现运行

论文中的实验方案开始执行后,运行记录会显示在这里。

结论—实验复现矩阵

逐个实验展示当前执行状态、阻塞原因、恢复动作和证据去向;技术执行与科研结论始终分开。

当前任务创建于目标级记录上线之前。下方状态来自历史任务的保守投影,不会伪造运行尝试、资源决策或证据关系。

The architecture ablation table reports that removing ContextNorm fails to converge, while additive fusion, auxiliary-loss budget control, and removing P-mask reduce the reported recall average relative to LoGo.

信息不足claim-architecture-ablations1 个方案0 次运行
科学结论尚未评估

Optional independent architecture-ablation reconstruction

实验方案已经存在,但当前处理任务没有对应的执行目标。

尚未进入任务

On the paper's 32k-checkpoint short-context recall aggregate, LoGo is reported as competitive with the other methods and scores 57.65 percentage points across six recall tasks.

信息不足claim-short-context-recall1 个方案0 次运行
科学结论尚未评估

Optional independent short-context recall reconstruction

实验方案已经存在,但当前处理任务没有对应的执行目标。

尚未进入任务

In the controlled 1.5B comparison, LoGo reports the best language-modeling loss and Lambada perplexity among the listed full-attention and matched-budget static hybrid variants, with a higher commonsense-reasoning average than the full Transformer.

信息不足claim-matched-budget1 个方案0 次运行
科学结论尚未评估

Optional independent matched-budget reconstruction

实验方案已经存在,但当前处理任务没有对应的执行目标。

尚未进入任务

LoGo preserves the reported scaling behavior of full-attention Transformers from 200M through 3.3B parameters, matching or improving the baseline on the listed language-modeling and commonsense metrics.

信息不足claim-scaling1 个方案0 次运行
科学结论尚未评估

Optional independent scaling-direction reconstruction

实验方案已经存在,但当前处理任务没有对应的执行目标。

尚未进入任务

The paper reports that its query-sparse Triton kernel reaches a 1.99x forward-plus-backward speedup over the dense Triton baseline at 64k sequence length and a 0.5 attention budget.

信息不足claim-query-sparse-speedup0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

LoGo reports higher average needle-style RULER recall than the three comparison paradigms at both the 32k and 128k context-extension stages, with averages of 83.0 and 65.4 percentage points.

信息不足claim-long-range-recall1 个方案0 次运行
科学结论尚未评估

Optional independent long-range recall reconstruction

实验方案已经存在,但当前处理任务没有对应的执行目标。

尚未进入任务