LLM Enhancements
KV-Cache Optimisation Is Becoming a Memory-Systems Discipline
Recent research targets how KV caches are compressed, retrieved and staged across memory tiers. For inference teams, the opportunity depends on matching each technique to the actual serving bottleneck.

A smaller KV cache can make long-context inference more practical, but cache size alone does not explain serving performance. The critical questions are what data each generated token requires, where that data resides, and whether the system can retrieve it without making the accelerator wait. Recent research increasingly treats these questions as one memory-systems problem rather than separate compression and offloading exercises.[4][13][14][15]
Papers posted between 25 and 30 September 2026 illustrate this shift. KV-Kaizen reports a 32× smaller decode-time cache on a 14B model while preserving accuracy. PulseInfer reports up to 4.7× higher decode throughput than SGLang. TempoKV reports up to 48.0% lower p95 time to first token, or TTFT, than an unmodified LMCache Device-DAX L1 configuration.[6][13][15]
Those results are not interchangeable or additive. They measure different interventions against different baselines. For engineers and infrastructure buyers, their collective significance is architectural: compression, selection, data movement and storage placement all influence whether retained attention history becomes useful capacity or a serving bottleneck.
1. Diagnose the memory wall before choosing a compression ratio
“The KV Cache Is the New Memory Wall” argues that long-context autoregressive inference is constrained primarily by memory bandwidth rather than arithmetic throughput. As the sequence grows, accessing cached keys and values for prior tokens becomes increasingly important relative to reading model weights.[4]
The paper identifies three performance regimes and a hardware-specific crossover. Below that crossover, weight traffic dominates, so KV-cache compression produces little speedup. Beyond it, KV traffic becomes dominant, creating an opportunity for compression and eviction to reduce bandwidth demand.[4]
This distinction changes how teams should interpret a cache-reduction claim. A smaller representation can reduce memory consumption without materially improving latency if KV traffic is not the binding constraint. Even when it is, the serving hardware must be able to exploit the reduced representation. Capacity reduction is therefore evidence of a changed memory footprint, not sufficient evidence of a faster service.[4]
Quality introduces another boundary. The same paper reports that quantisation and eviction reduce bandwidth directly, but quality degradation accelerates below 4-bit precision. Eviction can also cause discontinuous degradation on tasks sensitive to token position.[4] The resulting trade-off is not necessarily smooth: removing slightly more retained context may have a disproportionate effect on a particular task.
For evaluation, the implication is to establish the operating regime first. Teams should compare the contexts and serving configurations they intend to support, then determine whether cache optimisation changes the resource that limits performance. Accuracy testing should accompany that work, especially where successful completion depends on information appearing at particular positions in the history. A single compression ratio cannot describe both the performance opportunity and the quality risk.
2. Adaptive compression makes the policy part of the system
“KV-Kaizen: Learning Context-Adaptive Cache Compression Choices” addresses a limitation of fixed compression policies: different contexts and tasks do not require identical treatment. Its selector chooses compression decisions based on context rather than applying the same compression ratio to every request.[6]
On instruction-following and reasoning tasks, the paper reports that its selectors reach the Pareto frontier of accuracy against cache size compared with learning-free and post-hoc baselines. Within the reported comparison, the relevant operating points cannot simply be improved upon by choosing an evaluated baseline that delivers both higher accuracy and a smaller cache.[6]
The strongest cache-size result comes from long-context evaluation. KV-Kaizen improves on eviction and can be combined with it, reaching a 32× smaller decode-time cache on a 14B model while preserving accuracy. The paper also reports that a 4× cache-size reduction incurs no accuracy degradation from 7B parameters upward.[6]
These are distinct findings, not a universal compression setting. The 32× result is attached to the reported 14B evaluation. The 4× result describes another reported operating range. Neither establishes that every model or workload can use those reductions without quality loss.
A further comparison matters for infrastructure planning: at the same cache size, the compressed model is reported to be more accurate than a smaller uncompressed model.[6] This challenges the assumption that selecting a smaller model is always the best way to fit a memory budget. Adaptive cache compression offers another variable to test before reducing model capacity.
The practical evaluation question is therefore broader than “How many bytes can we remove?” It is whether context-sensitive compression offers a better quality and memory trade-off for the intended workload. Teams should compare candidate policies with both eviction and smaller-model alternatives, using the same task requirements. KV-Kaizen supports that comparison as a promising direction, not a universal ranking across models, hardware and applications.
3. Sparse offloading succeeds or fails on data movement
PulseInfer and AVSG approach the cache problem from complementary directions. PulseInfer targets long-context decoding in which large KV caches restrict batch size and leave GPUs underutilised. AVSG targets the overhead of gathering sparse KV entries from offloaded storage and transferring them back to GPU memory.[13][14]
Implemented on SGLang, PulseInfer reports up to 4.7× higher decode throughput than SGLang and up to 2.6× higher decode throughput than the best existing offloading baseline. It also reports up to 76% lower time per output token, or TPOT, with near-lossless accuracy in its evaluation.[13]
The two throughput comparisons answer different questions. The SGLang comparison measures the reported improvement over that serving baseline. The offloading comparison measures improvement over an alternative approach to moving KV state outside GPU memory. Neither factor should be presented without its baseline, and neither is a guarantee for an arbitrary deployment.
PulseInfer's central systems contribution is to frame sparse KV offloading as an I/O problem. Its results depend on the evaluated workload, cache sparsity, storage hierarchy and serving configuration.[13] Choosing fewer entries is only part of the work. The serving stack still has to retrieve the selected entries efficiently enough to benefit decoding.
AVSG makes that second requirement explicit. On a serving stack processing a real-world production dataset, it reports a 39% reduction in TPOT and a 1.27× increase in output throughput relative to the same KV-offload layout without HBM-buffer reuse.[14] That comparison isolates a different intervention from simply adopting sparse offloading.
The paper attributes its improvement to efficient matching, balanced transfer and effective residency in high-bandwidth memory. Reusing data already resident in HBM can reduce repeated transfers and improve the practical value of sparse offloading.[14]
Together, the papers suggest a useful evaluation discipline: assess the selection policy and the retrieval path separately. A promising sparsity result should prompt questions about gather overhead, transfer behaviour and reuse, not just the percentage of cache retained. Their reported gains should not be multiplied into a combined throughput projection.
4. Tiered caching is also a scheduling problem
“TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash” focuses on when cache data enters faster storage tiers. It does not primarily ask how aggressively KV entries can be compressed. Instead, it addresses the timing of staging and the capacity consumed while staged data waits to be useful.[15]
Across two models and three prefix-cache ratios, TempoKV reduces protected fast-tier byte-time per request by 63% to 91% compared with immediate staging, while retaining much of the serving benefit of advance staging.[15] Byte-time measures capacity occupied over time. It therefore captures a resource cost that a static cache-size measurement misses.
Against an unmodified LMCache Device-DAX L1 configuration, TempoKV reports up to 48.0% lower p95 TTFT and up to 27.8% higher output throughput.[15] These results connect staging decisions to both request-start latency and serving output, but they remain tied to the evaluated configuration, models and prefix-cache ratios.
The architectural lesson is that earlier placement in a fast tier is not automatically the most efficient use of that tier. Immediate staging can occupy protected capacity for longer. TempoKV demonstrates an alternative that reduces this occupancy measure while retaining much of the advantage of staging ahead of use.[15]
For teams evaluating tiered cache systems, this argues for measuring duration as well as volume. How much cache data is held in the protected tier, and for how long? Does the staging policy improve p95 TTFT under the intended prefix-reuse conditions? Those questions are more informative than evaluating fast-tier capacity in isolation. They also keep a byte-time reduction from being mistaken for an equivalent reduction in physical memory requirements or end-to-end latency.
5. Agentic serving brings SSD behaviour into the latency budget
Agentic workloads add another reason to retain KV state: sessions can contain long histories and repeated interactions. “Efficient Agentic LLM Serving over SSD-based Sparse KV Storage” presents Janus, which treats the resulting storage access patterns as part of the serving system rather than a secondary persistence concern.[9]
Janus runs the model's own KV-selection module on earlier intermediate values to predict future KV demand without additional training. It also coalesces adjacent SSD reads, packs scattered KV pages into sequential CPU writes, and limits background writes while reads are active.[9]
These mechanisms address different parts of the retrieval path. Prediction anticipates the state that will be needed. Read coalescing and page packing change how scattered cache data is accessed and written. Controlling background writes addresses interference while reads are in progress.[9] The contribution is therefore broader than simply placing retained history on SSDs.
Across three models and three agentic traces, Janus reports TTFT improvement factors ranging up to 1.57× to 3.69× over existing work, with average improvement factors of 1.22× to 1.85×, while maintaining decode efficiency.[9] These results concern the evaluated agentic traces, not a general latency multiplier for SSD-backed inference.
For teams running persistent agent sessions, Janus suggests evaluating storage-backed retention against realistic interaction histories. A useful assessment should include both the latency of making retained state available and the efficiency of subsequent decoding. Treating those as separate metrics helps avoid selecting a storage policy that improves one phase while obscuring its effect on another.
6. Weight quantisation complements cache optimisation, but measures something different
“Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference” addresses low-precision model inference rather than KV-cache storage. It belongs in this discussion because reducing weight-memory use can leave more GPU memory available for concurrent sequences or retained KV history.[4][7]
The paper reports 99.35% and 100.84% question-weighted recovery from BF16 across seven benchmarks on Qwen3.5-397B-A17B and Llama-3.3-70B-Instruct, respectively.[7] Those figures describe the reported benchmark recovery, not a universal accuracy guarantee for NVFP4.
On the 397B model, its infrastructure reduces measured per-layer time by 15.17× compared with ModelOpt and 23.14× compared with LLM Compressor, while using less memory per GPU.[7] These implementation results should not be relabelled as decode-throughput gains. They measure something different from PulseInfer's serving throughput or TempoKV's p95 TTFT.
The complementary opportunity is a shared memory budget. Weight quantisation changes the space occupied by model weights; cache compression changes the space required to retain attention history. The memory-wall analysis indicates that the binding traffic shifts with workload conditions, so improving one does not establish that the other has stopped mattering.[4]
A sound assessment should therefore track the two interventions separately before evaluating them together. Teams need to distinguish benchmark quality, memory consumption, quantisation implementation measurements and serving performance. Combining techniques may be valuable, but the brief's results do not establish a combined speedup or quality outcome.
7. What this means for teams running their own inference infrastructure
The common thread across this research is not a single winning compression method. It is the need to coordinate representation, selection, movement and placement around the workload's actual constraint. Adaptive compression changes how retained state is represented. Sparse retrieval changes what is fetched. Gather optimisation changes the cost of fetching it. Timely staging changes where data waits and for how long.[6][9][13][14][15]
For infrastructure teams, a practical evaluation programme should follow five principles:
- Identify the limiting resource first. Establish whether the relevant operating points are dominated by weight traffic, KV bandwidth, decode-time cache capacity or tiered retrieval. Do not assume a capacity problem and a bandwidth problem require the same intervention.[4][13]
- Keep metrics and baselines explicit. Cache size, decode throughput, TPOT, p95 TTFT and protected fast-tier byte-time describe different outcomes. Record the serving stack and comparison configuration alongside each result.
- Test quality under the intended policy. Include tasks that exercise long-context retention and position sensitivity. The reported acceleration of degradation below 4-bit precision and discontinuous eviction effects make quality a deployment constraint, not an afterthought.[4]
- Evaluate the complete retrieval path. For offloaded state, examine sparse gathers, HBM reuse, staging timing and read/write interference. The cited systems show that these mechanisms can materially affect serving results.[9][14][15]
- Compare alternatives within the same budget. Test adaptive compression against eviction and smaller-model options rather than assuming model downsizing is the only route to a workable memory footprint.[6]
Infrastructure buyers should likewise ask which bottleneck a proposed optimisation removes, on which workloads, and against which baseline. A headline cache reduction is not a service-level commitment, and results from separate papers cannot be assembled into one aggregate performance claim.
The opportunity is nevertheless substantial: retained attention history is becoming a resource that can be managed adaptively across representations and storage tiers. Teams running their own inference infrastructure should evaluate these techniques as coordinated changes to a memory system, with workload-specific quality and latency requirements determining which changes are worth deploying.
Sources
- [1]Resource-Efficient Speculative Decoding for Long-Context LLM Serving
- [2]arxiv.org · abs · 2609OLED-MoE: Accelerating MoE-Based dLLM Inference via...
- [3]Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference
- [4][2609.30854] The KV Cache Is the New Memory Wall
- [5]Greenpixie's AI Token Methodology: Assessing the Energy, Water and $\mathrm{CO_2\text{-}eq}$ Impact of AI Tokens for Open and Closed Weight Models
- [6]KV-Kaizen: Learning Context-Adaptive Cache Compression Choices
- [7]Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference
- [8]amarcu.github.io · ai-paper-digestAI Paper Digest, 2026-09-25
- [9]Efficient Agentic LLM Serving over SSD-based Sparse KV Storage
- [10]Dynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUs
- [11]Spexis: Speculative Lookahead Scheduling for LLM Inference
- [12]flozi.net › en › guidesvLLM vs SGLang: Architecture, Overhead, and When to Pick ...
- [13]PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding
- [14]AVSG: Accelerated Vectorized Sparse Gather for Efficient KV Cache Offload in Sparse-Attention LLM Serving
- [15]TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash
Researched from the sources above and fact-checked against them before publishing.
See what this means for your infrastructure.
Book a discovery call and get a benchmark projection for your hardware and model.
SCHEDULE A CALL