LoGo: Token-Level Dynamic Local-Global Attention
Not assessedPlan blockedFindingclaim-matched-budget
In the controlled 1.5B comparison, LoGo reports the best language-modeling loss and Lambada perplexity among the listed full-attention and matched-budget static hybrid variants, with a higher commonsense-reasoning average than the full Transformer.
Source: paper-fixed:Section 4.2 and Table 2, p. 7
Reported and observed measurements
train_loss
m-table2-logo-loss
Reported 2.112 loss
Observed — loss
lambada_perplexity
m-table2-logo-lambada-ppl
Reported 7.5 perplexity
Observed — perplexity
commonsense_reasoning_average
m-table2-logo-csr
Reported 56.3 percentage_points
Observed — percentage_points
Assessments (0)
No immutable Assessment has been published for this Claim yet.
Runs (0)
No execution has been linked to this Claim yet.