LoGo: Token-Level Dynamic Local-Global Attention

Public
Authors:Yuqi Pan, Zheng Li, Bohao Tang, Zhen Qin, Guoqi Li
More
Copy repository link

Reproduction

Start with coverage and measured results; open a run only when you need evidence or technical details.

RunMatchRepeat

0/6

claims supported by evidence

0

Supported

0

Challenged or mixed

0

Contradicted

0

Inconclusive

6

Not assessed

Run history

Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.

No reproduction runs yet

Runs will appear here as the paper's experiment plans are executed.

Claim–experiment reproduction matrix

See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions.

This task predates target-level records. The states below are conservative projections; no attempts, resource decisions, or evidence links are invented.

The architecture ablation table reports that removing ContextNorm fails to converge, while additive fusion, auxiliary-loss budget control, and removing P-mask reduce the reported recall average relative to LoGo.

Information insufficientclaim-architecture-ablations1 plan0 runs
Scientific conclusionNot assessed

Optional independent architecture-ablation reconstruction

An experiment plan exists, but the current processing task has no matching execution target.

Not in current task

On the paper's 32k-checkpoint short-context recall aggregate, LoGo is reported as competitive with the other methods and scores 57.65 percentage points across six recall tasks.

Information insufficientclaim-short-context-recall1 plan0 runs
Scientific conclusionNot assessed

Optional independent short-context recall reconstruction

An experiment plan exists, but the current processing task has no matching execution target.

Not in current task

In the controlled 1.5B comparison, LoGo reports the best language-modeling loss and Lambada perplexity among the listed full-attention and matched-budget static hybrid variants, with a higher commonsense-reasoning average than the full Transformer.

Information insufficientclaim-matched-budget1 plan0 runs
Scientific conclusionNot assessed

Optional independent matched-budget reconstruction

An experiment plan exists, but the current processing task has no matching execution target.

Not in current task

LoGo preserves the reported scaling behavior of full-attention Transformers from 200M through 3.3B parameters, matching or improving the baseline on the listed language-modeling and commonsense metrics.

Information insufficientclaim-scaling1 plan0 runs
Scientific conclusionNot assessed

Optional independent scaling-direction reconstruction

An experiment plan exists, but the current processing task has no matching execution target.

Not in current task

The paper reports that its query-sparse Triton kernel reaches a 1.99x forward-plus-backward speedup over the dense Triton baseline at 64k sequence length and a 0.5 attention budget.

Information insufficientclaim-query-sparse-speedup0 plans0 runs
Scientific conclusionNot assessed

This claim has no executable experiment plan yet.

LoGo reports higher average needle-style RULER recall than the three comparison paradigms at both the 32k and 128k context-extension stages, with averages of 83.0 and 65.4 percentage points.

Information insufficientclaim-long-range-recall1 plan0 runs
Scientific conclusionNot assessed

Optional independent long-range recall reconstruction

An experiment plan exists, but the current processing task has no matching execution target.

Not in current task