LLM Enhancements
SpecStream: Coordinating KV Offloading and Speculative Decoding for Long-Context Inference
SpecStream overlaps KV-cache transfers, verification and drafting for long-context inference. Its reported gains show why throughput per GPU matters alongside aggregate serving performance.

Speculative decoding addresses a familiar limitation of autoregressive inference: generating tokens through repeated, sequential target-model execution. Long-context serving adds a different constraint. Even when speculative execution reduces sequential work, the key-value (KV) cache can exceed GPU memory, forcing the serving system to move attention history between CPU and GPU memory. Accelerating generation then becomes a scheduling problem as well as a computation problem.[9]
SpecStream, introduced in Resource-Efficient Speculative Decoding for Long-Context LLM Serving, tackles these constraints together. It combines CPU KV-cache offloading with streamed restoration, overlapping target-model verification with cache transfers and scheduling draft-model work during transfer-related GPU compute gaps.[9]
The reported results are promising: average throughput improvements of 1.41× for Qwen3 and 1.32× for InternLM2.5 over an offloading baseline. Against parallel speculative decoding with separate target and draft GPUs, SpecStream reports 55.4% higher output throughput per GPU on average.[9] The important question for infrastructure teams is what produces those gains, and which comparisons should inform their own deployment decisions.
1. Long contexts expose a bottleneck beyond sequential decoding
Speculative decoding uses a smaller draft model to propose multiple tokens, which a larger target model then verifies in parallel. By advancing generation through proposed tokens rather than relying exclusively on repeated single-token target calls, the method reduces sequential target-model execution.[9]
That improvement does not remove the need to retain attention history. KV caching stores keys and values from previously processed tokens so the model does not have to recompute the entire history at every decoding step. The cache grows with sequence length and the number of active requests. Long contexts and greater concurrency therefore put pressure on the same limited resource: GPU memory.[9]
When the full cache no longer fits, keeping every request's attention history resident on the GPU is not an available option. CPU offloading expands effective cache capacity, but introduces transfers between CPU and GPU memory. A system can consequently relieve its memory-capacity constraint while making data movement a more prominent part of the serving pipeline.[9]
This distinction matters when interpreting speculative-decoding performance. Reducing the number of sequential target calls addresses one source of delay. It does not, by itself, establish how efficiently verification will proceed when the required KV data is offloaded. The execution schedule must account for both the model work and the availability of attention history.
SpecStream focuses on precisely this combined constraint: maintaining the benefits of speculative decoding while supporting greater concurrency under limited GPU memory.[9] Its contribution is not simply to add CPU storage to a speculative-decoding system. It changes how cache restoration, verification and drafting are coordinated so that they need not operate as isolated stages.
2. Streamed restoration lets verification start before transfers finish
The central design choice in SpecStream is streamed KV-cache restoration. Instead of restoring the complete offloaded KV history before beginning target-model verification, the system starts verification while cache data is still being transferred. This overlaps memory movement with useful computation.[9]
The difference is a scheduling barrier. With full-cache restoration, verification waits for the required history to return before it can begin. SpecStream removes that full-history barrier: verification can make progress before restoration of the entire history has finished.[9] The transfer work still exists, but it no longer has to sit entirely ahead of verification in the execution sequence.
That distinction prevents an overly broad reading of the result. SpecStream does not demonstrate that CPU-to-GPU transfers are free or that offloading has no performance cost. It demonstrates a design that overlaps those transfers with computation rather than treating restoration as a wholly separate prerequisite.[9] For infrastructure evaluation, the relevant issue is how much useful work the coordinated schedule enables during data movement.
Incremental restoration also raises an attention-semantics question. Processing only part of the KV history without an appropriate method could change the attention result. SpecStream uses online softmax to preserve full attention while processing the cache incrementally. The paper also states that each streamed KV chunk serves all queries in the round.[9]
These details distinguish streamed restoration from explicitly approximate attention. The reported design is intended to make full attention compatible with incremental cache processing, not to replace the complete history with an approximate subset.[9] Its efficiency argument rests on when cache data is processed and how that processing is shared within a round.
There is still an important boundary between a mechanism and an evaluation result. Preserving full attention through online softmax explains the intended computation. The separately reported observation that task quality remains close to SGLang speculative decoding describes the evaluated system's quality outcome.[9] Neither statement should be expanded into an unsupported guarantee about every workload or configuration.
3. Drafting uses compute gaps on the same GPUs
Streamed restoration is only part of the pipeline. SpecStream also uses GPU compute gaps created by KV transfers to run draft-model work on the same GPUs. Drafting therefore overlaps with other serving activity rather than requiring a separate accelerator allocation for the arrangement described in the paper.[9]
This changes the resource-efficiency question. Parallel speculative decoding can place the target and draft models on separate GPUs, allowing them to operate in parallel. That arrangement increases hardware requirements, however, so aggregate output throughput alone does not capture the efficiency of the accelerator resources involved.[9]
SpecStream instead exploits otherwise underused compute capacity during KV movement. Its design coordinates three activities: transferring cache data, verifying draft tokens with the target model and producing further draft tokens. The reported contribution is their overlap, not the elimination of any one activity.[9]
That is a useful distinction for engineers investigating a serving bottleneck. A draft model adds computation, while offloading introduces data movement. Evaluating either mechanism in isolation can miss the opportunity to schedule one activity during gaps created by another. SpecStream treats those interactions as part of the design rather than as independent costs.[9]
The comparison with separate target and draft GPUs should also remain specific. SpecStream reports an average 55.4% improvement in output throughput per GPU against that parallel speculative-decoding arrangement.[9] This supports a claim about resource-normalized throughput in the reported evaluation. It does not establish that separate-GPU drafting is always inferior, or that the same improvement will appear under every hardware configuration.
For infrastructure buyers, the lesson is to ask what accelerator footprint produces a throughput result. For engineers, it is to examine whether the schedule leaves useful computation waiting while cache movement is underway. Both questions are closer to SpecStream's contribution than a headline token-generation figure alone.
4. Read the reported gains against the right baselines
SpecStream's evaluation covers multiple datasets and two model families, Qwen3 and InternLM2.5. The brief reports the following outcomes:[9]
| Comparison or outcome | Reported result |
|---|---|
| Qwen3 throughput relative to an offloading baseline | 1.41× average speedup |
| InternLM2.5 throughput relative to an offloading baseline | 1.32× average speedup |
| Output throughput per GPU relative to parallel speculative decoding on separate target and draft GPUs | 55.4% average improvement |
| Concurrency under limited GPU memory | Support for more concurrent requests |
| Task quality | Close to SGLang speculative decoding |
These results answer different questions. The model-family speedups compare throughput against an offloading baseline. The per-GPU result compares resource-normalized output throughput against a separate-GPU speculative-decoding arrangement. They should not be presented as interchangeable measures of one universal acceleration factor.[9]
The concurrency result also matters independently of the headline throughput gains. SpecStream is designed for conditions in which the full KV cache does not fit in GPU memory. Supporting more concurrent requests under that constraint is part of its reported value, alongside coordinating the work required to serve them.[9] The brief does not provide a numerical concurrency increase, so none should be inferred.
Quality needs similarly careful treatment. “Close to SGLang speculative decoding” is the reported assessment, not a claim of identical task scores across all evaluations.[9] It is appropriate to describe the result as maintaining close task quality while improving serving efficiency. It would be inappropriate to convert that statement into an exact equivalence claim.
Finally, throughput is not a substitute for every latency metric. The figures summarized here do not establish corresponding numerical improvements in time to first token or time per output token. Teams should preserve those distinctions when translating research results into infrastructure requirements.
The strongest supported interpretation is therefore bounded but useful: in the evaluated models, datasets and baseline comparisons, coordinating offloaded KV transfers with verification and drafting improves throughput and accelerator efficiency.[9]
5. Related research targets different parts of the inference pipeline
Other research in the brief reinforces the importance of identifying the actual bottleneck before comparing performance numbers. Similar-looking speedups can describe substantially different changes to inference execution.
Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference reports output-throughput improvements of 1.91× to 2.19× across all prompts and reductions in time to first output of 10.0% to 14.2%. It attributes the improvement to amortization: additional token progress reduces repeated CUDA-graph execution enough to offset the added speculative-execution cost.[2]
The study also reports 56.4% to 78.1% fewer executions of the selected repeating CUDA Graph per generated token.[2] Its focus is execution amortization within GPU inference. SpecStream's focus is the interaction between speculative decoding and offloaded KV caches.[9] The two findings address different aspects of serving performance, so their throughput multipliers are not a direct ranking.
Janus, introduced in Efficient Agentic LLM Serving over SSD-based Sparse KV Storage, addresses another storage boundary. It uses SSD-centric KV storage for sparse-attention LLMs, predicts KV demand from earlier intermediate values, coalesces adjacent reads and packs scattered KV pages into sequential CPU writes. It also limits background writes while reads are active.[7]
Across three models and three agentic traces, Janus reports time-to-first-token improvements reaching 1.57× to 3.69×, with reported averages ranging from 1.22× to 1.85×, while maintaining decode efficiency.[7] Those results concern sparse KV storage and SSD-based serving, not SpecStream's CPU-memory offloading and speculative-verification pipeline.
The practical distinction is between reducing repeated execution, coordinating CPU-memory cache transfers and managing SSD-backed sparse storage. Each addresses a different constraint. The brief does not establish that these systems can be combined or that their gains would multiply.
6. Evaluate the pipeline, not just the headline multiplier
For a team assessing a SpecStream-style design, the first step is to establish whether the deployment matches the problem being addressed. Does the full KV cache exceed available GPU memory at the desired context lengths and concurrency? Is CPU offloading already necessary? Does cache restoration delay verification? These questions connect the evaluation to the system's stated purpose.[9]
The next step is to separate comparison goals. A throughput comparison against an offloading baseline evaluates the serving improvement relative to that baseline. A comparison against separate target and draft GPUs should also account for output throughput per GPU, because the accelerator footprint is central to the reported resource-efficiency claim.[9]
A useful evaluation plan would keep several questions explicit:
- Memory constraint: Is the deployment actually limited by KV-cache capacity at the intended workload?
- Pipeline overlap: Can verification begin while cache restoration continues, and can drafting use transfer-related compute gaps?
- Resource denominator: How many GPUs are involved in each configuration being compared?
- Concurrency and quality: Does the system support the required request load while meeting the team's task-quality requirements?
- Latency requirements: Are throughput gains accompanied by acceptable latency for the application?
These are deployment questions, not additional claims about SpecStream's measured behavior. The brief does not establish the result for every context length, architecture, hardware configuration or workload.[9] Local evaluation is therefore necessary to determine whether the reported mechanism addresses the team's own limiting resource.
This framing also helps keep procurement discussions precise. A percentage improvement in throughput per GPU is not automatically the same percentage reduction in total infrastructure cost. The cited result directly supports a resource-normalized performance claim, not a complete economic model.
The strongest evaluation would retain the distinctions the research makes: cache capacity versus transfer scheduling, aggregate throughput versus per-GPU throughput, and attention computation versus measured task quality. Collapsing those distinctions would make the headline simpler but the deployment decision less reliable.
7. What this means for teams running their own inference infrastructure
SpecStream's central lesson is that long-context inference efficiency depends on coordination across the serving pipeline. Speculative decoding reduces sequential target-model execution. CPU offloading expands effective KV-cache capacity. Streamed restoration and concurrent drafting make it possible to overlap some of the work those mechanisms require.[9]
For teams operating their own infrastructure, this suggests a more focused way to investigate performance. Rather than asking only whether speculative decoding is enabled or whether the cache fits, examine when verification can start, what the GPU can do during transfers and whether drafting requires additional accelerator resources. Those are the interactions SpecStream is designed to improve.
The reported results give that investigation a concrete basis: 1.41× average throughput for Qwen3 and 1.32× for InternLM2.5 relative to an offloading baseline, plus a 55.4% average improvement in output throughput per GPU against separate-GPU parallel speculative decoding.[9] They justify attention to the design, not an assumption that every deployment will reproduce the same gains.
For infrastructure buyers, the corresponding priority is to request results with clear baselines, resource counts and workload boundaries. For engineering teams, it is to evaluate throughput, concurrency, quality and application latency together under their own memory constraints.
SpecStream shows a specific way to make offloading and speculative execution work together: start verification before the entire cache has returned, preserve full attention through incremental processing and use transfer-related compute gaps for drafting.[9] The broader operational takeaway is equally specific. When GPU memory is insufficient, improving the schedule around data movement can be as important as reducing the sequential computation in generation.
Sources
- [1]arxiv.org · abs · 2609OLED-MoE: Accelerating MoE-Based dLLM Inference via...
- [2]Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference
- [3]Paper page - Hunyuan-A13B Technical Report - Hugging Face
- [4]Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference
- [5]Foundations of Large Language Models | arXiv Science
- [6]IQuest-Q1 by @IQuest_research has day-0 support in vLLM
- [7]Efficient Agentic LLM Serving over SSD-based Sparse KV Storage
- [8]Learning the Cost of Reliable Inference | arXiv Science
- [9]Resource-Efficient Speculative Decoding for Long-Context LLM Serving
- [10]Editors Pick Category - Page 92 of 1135
- [11]Large Language Model Category - Page 290 of 291
- [12]On Device Agentic Operation Caches -- Classifier-Centric NL-to-Action Generation
- [13]Editors Pick Category - Page 1135 of 1135
- [14]New Releases Category - Page 147 of 148
- [15]Agentic AI Category - Page 82 of 111
Researched from the sources above and fact-checked against them before publishing.
See what this means for your infrastructure.
Book a discovery call and get a benchmark projection for your hardware and model.
SCHEDULE A CALL