LoGo: Token-Level Dynamic Local-Global Attention
Not assessedPlan blockedFindingclaim-matched-budget

In the controlled 1.5B comparison, LoGo reports the best language-modeling loss and Lambada perplexity among the listed full-attention and matched-budget static hybrid variants, with a higher commonsense-reasoning average than the full Transformer.

Source: paper-fixed:Section 4.2 and Table 2, p. 7

Reported and observed measurements

train_loss

m-table2-logo-loss

Reported 2.112 loss

Observed — loss

lambada_perplexity

m-table2-logo-lambada-ppl

Reported 7.5 perplexity

Observed — perplexity

commonsense_reasoning_average

m-table2-logo-csr

Reported 56.3 percentage_points

Observed — percentage_points

Assessments (0)

No immutable Assessment has been published for this Claim yet.

Runs (0)

No execution has been linked to this Claim yet.