LoGo: Token-Level Dynamic Local-Global Attention
尚未评估计划受阻研究发现claim-scaling

LoGo preserves the reported scaling behavior of full-attention Transformers from 200M through 3.3B parameters, matching or improving the baseline on the listed language-modeling and commonsense metrics.

来源:paper-fixed:Table 1, p. 7

报告指标与观测值

train_loss

m-scaling-33b-logo-train-loss

论文报告 1.996 loss

实际观测 — loss

wikitext_perplexity

m-scaling-33b-logo-wikitext

论文报告 14.785 perplexity

实际观测 — perplexity

commonsense_reasoning_average

m-scaling-33b-logo-csr

论文报告 58.26 percentage_points

实际观测 — percentage_points

Assessment(0)

这条 Claim 暂无已发布的不可变 Assessment。

关联运行(0)

这条 Claim 暂无关联执行。