Online Self-Weighted Fine-Tuning
Not assessedPlan blockedFindingclaim-weighting-strategy-ablation

Under K=4 rollouts on Qwen3-0.6B-Base, continuous online success-rate weighting in OSW-FT (24.00% MATH-500, 20.71% GPQA-Diamond) outperforms hard-thresholded binary rejection weighting (23.80% MATH-500, 15.15% GPQA-Diamond) and mean-matched random weighting (21.40% MATH-500, 15.17% GPQA-Diamond).

Source: source-paper:Table 3, Section 4.4, page 6

Reported and observed measurements

pass_at_1

m-ablation-hard-math500-p1

Reported 23.8 percentage_points

Observed — percentage_points

pass_at_1

m-ablation-hard-gpqad-p1

Reported 15.15 percentage_points

Observed — percentage_points

pass_at_1

m-ablation-random-math500-p1

Reported 21.4 percentage_points

Observed — percentage_points

pass_at_1

m-ablation-random-gpqad-p1

Reported 15.17 percentage_points

Observed — percentage_points

pass_at_1

m-ablation-osw-math500-p1

Reported 24 percentage_points

Observed — percentage_points

pass_at_1

m-ablation-osw-gpqad-p1

Reported 20.71 percentage_points

Observed — percentage_points

Assessments (0)

No immutable Assessment has been published for this Claim yet.

Runs (0)

No execution has been linked to this Claim yet.