Coderhouse goes global

🌍

Read more

What Is Speculative Decoding and Why It Makes LLMs Generate Faster

AI Engineering

Build and ship real applications with LLMs, agents and automation.

8 weeks · Live online · View course →

Data Analytics

Go from zero to job-ready analyst with SQL, Python and BI tools.

10 weeks · Live online · View course →

Dan Patiño

AI Strategy & Innovation at Coderhouse

Artificial Intelligence

What Is Speculative Decoding and Why It Makes LLMs Generate Faster

Speculative decoding is a technique that makes large language models generate text faster without changing what they write. A small, fast “draft” model guesses the next several tokens, and the large “target” model checks all of those guesses in a single pass. Every guess the large model agrees with is kept for free; at the first disagreement, the large model supplies its own token and the cycle starts again. Because of the way acceptance is calculated, the final output follows the same distribution the large model would have produced on its own.

Why it matters now: AI agents chain many model calls together, and users expect chat and coding assistants to respond in real time. Generation speed has become a product feature and a cost lever. The original speculative decoding paper from Google researchers reported a 2X–3X speedup on T5-XXL with identical outputs, and DeepMind’s speculative sampling paper reported a 2–2.5x decoding speedup on the 70-billion-parameter Chinchilla model without compromising sample quality. Since then, the idea has become a standard option in major inference engines.

Why LLM Generation Is Slow in the First Place

Language models are autoregressive: they produce one token, append it to the input, and run again to produce the next one. Generating 500 tokens means 500 sequential passes through the model. On modern GPUs, each of those passes is usually limited less by math than by memory bandwidth, because the full set of model weights has to be moved from memory to the processor for every step. That means the hardware often has spare compute capacity while it waits. Checking several candidate tokens in one pass costs roughly the same time as generating one, and speculative decoding exploits exactly that gap.

How Speculative Decoding Works

The loop has three steps. First, the draft model proposes a short sequence, for example four or five tokens. Second, the target model scores the whole proposed sequence in parallel. Third, a verification rule accepts tokens from left to right as long as they are consistent with the target model’s probabilities, then replaces the first rejected token with one sampled from the target model. In the best case, one expensive pass yields several tokens; in the worst case, it still yields one, so you are never slower than a single step of normal decoding, aside from the small cost of running the draft.

The key metric is the acceptance rate: how often the large model agrees with the draft. Predictable text, such as code boilerplate, structured JSON, or answers that repeat parts of the prompt, tends to get high acceptance and large speedups. Creative or highly uncertain text gets lower acceptance and smaller gains.

Variants Without a Separate Draft Model

Maintaining a second model that shares the same tokenizer is a practical hurdle, so newer methods build the drafter into the main model. Medusa adds extra decoding heads that predict several future tokens at once and verifies multiple candidate continuations together using tree-based attention. Other approaches draft tokens by looking up n-grams in the prompt, which works well for summarization and code editing where the output repeats the input. Hugging Face describes the draft-model approach as assisted generation and supports it directly in its Transformers library.

Speculative Decoding vs. Other Ways to Speed Up LLMs

Teams have several tools for faster, cheaper inference, and they solve different problems. Quantization shrinks the weights to lower precision, which reduces memory and can slightly change outputs. Model distillation trains a smaller model to imitate a larger one, which is cheaper to run but is a different model. Switching to small language models trades capability for speed. Speculative decoding is different: it keeps the large model’s output quality intact and only changes how quickly tokens arrive. That is why it is often combined with the others rather than chosen instead of them.

Common Mistakes When Adopting It

The most common mistake is benchmarking on the wrong workload. Speedups measured on repetitive prompts may not hold for your real traffic, so test with production-like requests and track acceptance rate alongside latency. A second mistake is choosing a draft model that is too large; if drafting is slow, the savings disappear. A third is ignoring batch size: under very heavy concurrent load, the GPU has less idle compute to spare, and the gains from speculation can shrink. Finally, measure end-to-end latency and cost per request, not just tokens per second in isolation.

If you want to learn how to build, optimize, and serve AI models and agents in production, and how to measure their performance with data, structured practice helps. Explore Coderhouse’s courses below.

  • AI Engineering Course — learn to deploy, optimize, and evaluate LLM-powered applications and agents.

  • Data Analytics Course — build the measurement and analysis skills to benchmark latency, cost, and quality.

FAQ

Does speculative decoding change the model’s answers?

No, when implemented with the standard verification rule. The accepted tokens follow the same probability distribution as the large model alone, so quality is preserved; only speed changes.

Do I need to retrain my model to use it?

Not with the classic draft-model approach, which works with off-the-shelf models that share a tokenizer. Methods like Medusa do require training extra heads, but they avoid managing a separate draft model.

How much faster does it make generation?

Published results commonly show around 2x to 3x faster decoding, but real gains depend on how predictable your outputs are, the draft model you choose, and how heavily loaded your servers are.

Dan Patiño

I'm Dan Patiño, head of AI Strategy & Innovation at Coderhouse. My day-to-day work involves merging the tactical management of e-commerce (CRO, Email Marketing and SEO) with the development of disruptive solutions. I specialize in building internal AI-powered apps to automate tasks and boost innovation within the team. I firmly believe that technology is strategy's best ally. To dive deeper into my professional journey, I'll be waiting for you on my LinkedIn profile.

© 2026 Coderhouse. All rights reserved.

© 2026 Coderhouse. All rights reserved.

© 2026 Coderhouse. All rights reserved.