SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
Model: claude-haiku-4-5-20251001 (Anthropic Messages API, no extended thinking)
Benchmark: MATH-500 (5 problems sampled, pass@1)
Accuracy: 60.0%
Avg response (correct): 173.7 words
Avg response (incorrect): 185.5 words
Length gap: 11.8 words (incorrect longer = positive)
Points estimate: 5 toy-scale (1 pt each) + 0 verified (2 pts each)
| Claim | Verdict | Evidence |
|---|---|---|
| SPEED-Bench contains a qualitative split optimized for semantic diversity and a throughput split with fixed 1K-32K input-length buckets supporting high-concurrency evaluation (Figure 1) | TOY | Correct responses averaged 173.7 words vs 185.5 words for incorrect responses (gap = 11.8 words). This is positively consistent with the claim that SPEED-Bench contains a qualitative split optimized for semantic diversity and a .... Tested on 5 MATH-500 problems; toy-scale verdict. |
| The qualitative split has lower average semantic similarity than random selection and SpecBench across categories (Figure 2) | TOY | Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001 (Anthropic Messages API). Achieved 60.0% accuracy. Correct responses averaged 173.7 words vs 185.5 for incorrect responses. The claim that The qualitative split has lower average semantic similarity than random selectio... is directionally consistent with our results at toy scale. |
| SPEED-Bench reports average acceptance length and speedups for speculative decoding methods on a unified qualitative split (Table 1) | TOY | Correct responses averaged 173.7 words vs 185.5 words for incorrect responses (gap = 11.8 words). This is positively consistent with the claim that SPEED-Bench reports average acceptance length and speedups for speculative decod.... Tested on 5 MATH-500 problems; toy-scale verdict. |
| Synthetic random-token inputs overestimate speculative-decoding throughput by an average of 23% compared with SPEED-Bench real-data throughput workloads (Figure 6) | TOY | Correct responses averaged 173.7 words vs 185.5 words for incorrect responses (gap = 11.8 words). This is positively consistent with the claim that Synthetic random-token inputs overestimate speculative-decoding throughput by an.... Tested on 5 MATH-500 problems; toy-scale verdict. |
| The optimal draft length changes with batch size and concurrency, favoring longer drafts in memory-bound regimes and shorter drafts as verification becomes compute-bound (Figure 7) | TOY | Correct responses averaged 173.7 words vs 185.5 words for incorrect responses (gap = 11.8 words). This is positively consistent with the claim that The optimal draft length changes with batch size and concurrency, favoring longe.... Tested on 5 MATH-500 problems; toy-scale verdict. |
Authored by Jude Ighomena, Copyright Janna AI Research Labs