LoGo: Token-Level Dynamic Local-Global Attention
Not assessedPlan blockedFindingclaim-scaling
LoGo preserves the reported scaling behavior of full-attention Transformers from 200M through 3.3B parameters, matching or improving the baseline on the listed language-modeling and commonsense metrics.
Source: paper-fixed:Table 1, p. 7
Reported and observed measurements
train_loss
m-scaling-33b-logo-train-loss
Reported 1.996 loss
Observed — loss
wikitext_perplexity
m-scaling-33b-logo-wikitext
Reported 14.785 perplexity
Observed — perplexity
commonsense_reasoning_average
m-scaling-33b-logo-csr
Reported 58.26 percentage_points
Observed — percentage_points
Assessments (0)
No immutable Assessment has been published for this Claim yet.
Runs (0)
No execution has been linked to this Claim yet.