LoGo: Token-Level Dynamic Local-Global Attention
Not assessedPlan blockedFindingclaim-scaling

LoGo preserves the reported scaling behavior of full-attention Transformers from 200M through 3.3B parameters, matching or improving the baseline on the listed language-modeling and commonsense metrics.

Source: paper-fixed:Table 1, p. 7

Reported and observed measurements

train_loss

m-scaling-33b-logo-train-loss

Reported 1.996 loss

Observed — loss

wikitext_perplexity

m-scaling-33b-logo-wikitext

Reported 14.785 perplexity

Observed — perplexity

commonsense_reasoning_average

m-scaling-33b-logo-csr

Reported 58.26 percentage_points

Observed — percentage_points

Assessments (0)

No immutable Assessment has been published for this Claim yet.

Runs (0)

No execution has been linked to this Claim yet.