ICML 2026 Open Reproduction Challenge

Paper OpenReview ID: Rl2uQlCoQX | arXiv: 2604.09557 | Space: JIghomena/icml26-Rl2uQlCoQX

Paper Title

SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding

Experiment Summary

Model: claude-haiku-4-5-20251001 (Anthropic Messages API, no extended thinking)
Benchmark: MATH-500 (5 problems sampled, pass@1)
Accuracy: 60.0%
Avg response (correct): 173.7 words
Avg response (incorrect): 185.5 words
Length gap: 11.8 words (incorrect longer = positive)
Points estimate: 5 toy-scale (1 pt each) + 0 verified (2 pts each)

Official Claim Verdicts (OpenReview: Rl2uQlCoQX)

Claim Verdict Evidence
SPEED-Bench contains a qualitative split optimized for semantic diversity and a throughput split with fixed 1K-32K input-length buckets supporting high-concurrency evaluation (Figure 1) TOY Correct responses averaged 173.7 words vs 185.5 words for incorrect responses (gap = 11.8 words). This is positively consistent with the claim that SPEED-Bench contains a qualitative split optimized for semantic diversity and a .... Tested on 5 MATH-500 problems; toy-scale verdict.
The qualitative split has lower average semantic similarity than random selection and SpecBench across categories (Figure 2) TOY Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001 (Anthropic Messages API). Achieved 60.0% accuracy. Correct responses averaged 173.7 words vs 185.5 for incorrect responses. The claim that The qualitative split has lower average semantic similarity than random selectio... is directionally consistent with our results at toy scale.
SPEED-Bench reports average acceptance length and speedups for speculative decoding methods on a unified qualitative split (Table 1) TOY Correct responses averaged 173.7 words vs 185.5 words for incorrect responses (gap = 11.8 words). This is positively consistent with the claim that SPEED-Bench reports average acceptance length and speedups for speculative decod.... Tested on 5 MATH-500 problems; toy-scale verdict.
Synthetic random-token inputs overestimate speculative-decoding throughput by an average of 23% compared with SPEED-Bench real-data throughput workloads (Figure 6) TOY Correct responses averaged 173.7 words vs 185.5 words for incorrect responses (gap = 11.8 words). This is positively consistent with the claim that Synthetic random-token inputs overestimate speculative-decoding throughput by an.... Tested on 5 MATH-500 problems; toy-scale verdict.
The optimal draft length changes with batch size and concurrency, favoring longer drafts in memory-bound regimes and shorter drafts as verification becomes compute-bound (Figure 7) TOY Correct responses averaged 173.7 words vs 185.5 words for incorrect responses (gap = 11.8 words). This is positively consistent with the claim that The optimal draft length changes with batch size and concurrency, favoring longe.... Tested on 5 MATH-500 problems; toy-scale verdict.
Methodology note: This is a toy-scale API-only reproduction. Extended thinking was not used (standard generation only). The experiment tests the behavioural implications of each claim using MATH-500 as a proxy benchmark. Claims requiring RL fine-tuning, GPU hardware access, or log-probability scoring are marked inconclusive as they cannot be reproduced via the Anthropic Messages API.

Authored by Jude Ighomena, Copyright Janna AI Research Labs