LLM Enhancements

What Is LLM Inference? How It Works, What Limits It, and How to Optimize It
Short answer: LLM inference is the process of running a trained large language model to turn a prompt into new tokens. It happens in two phases: prefill, which reads the whole prompt in parallel, and decode, which generates one token at a time. Decode is usually limited by GPU memory bandwidth rather than raw compute.
Training builds a model once. Inference is the recurring bill: every chat reply, code completion, document summary and agent step is an inference request. For teams that own GPUs, inference efficiency decides how many users, tokens and agents a node can serve before you need more hardware. This guide explains how LLM inference works, which metrics matter, what actually limits speed, and which optimization techniques have published evidence behind them.
What is LLM inference?
Most popular LLMs are decoder-only models trained as next-token predictors. At inference time they take a sequence of input tokens and generate new tokens autoregressively until they hit a stop condition, such as a token limit or an end-of-sequence token (NVIDIA Technical Blog, "Mastering LLM Techniques: Inference Optimization", Nov 17, 2023). The weights do not change during inference. The model only reads them.
LLM inference vs training
Training updates billions of parameters using large batches of data and backpropagation. Inference uses the frozen weights to answer new prompts. Training is a large, periodic project. Inference runs continuously for as long as the product is in use, so small per-token gains compound into real capacity.
How LLM inference works, step by step
1. Tokenization
Text is converted into tokens before it reaches the model. NVIDIA notes that one token is roughly four English characters, and that different tokenizers make tokens-per-second figures hard to compare across models (NVIDIA, Nov 17, 2023).
2. Prefill: processing the prompt
In the prefill phase the model processes all input tokens at once and computes the intermediate keys and values needed to produce the first output token. Because the full input is known, this is a matrix-matrix operation that is highly parallel and keeps the GPU busy (NVIDIA, Nov 17, 2023). Long prompts, such as retrieval-augmented generation (RAG) contexts, make prefill heavier.
3. Decode: generating the answer
In the decode phase the model generates output tokens one at a time. Each step behaves like a matrix-vector operation that underuses GPU compute. NVIDIA describes decode as memory-bound: the speed at which weights, keys and values move from memory dominates latency, not the arithmetic itself (NVIDIA, Nov 17, 2023). The vLLM authors reach the same conclusion and add that decode is responsible for most of the latency of a single request (Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention", SOSP 2023).
4. The KV cache
To avoid recomputing attention for every previous token at every step, servers cache the key and value tensors in GPU memory. NVIDIA gives the per-token size as 2 x layers x hidden size x bytes per value. For Llama 2 7B in 16-bit precision with a 4,096-token sequence and batch size 1, that is about 2 GB of KV cache, on top of roughly 14 GB for the weights (NVIDIA, Nov 17, 2023). The cache grows linearly with batch size and sequence length, which is why long contexts and many concurrent users collide with memory limits. For a deeper look at this problem, see our analysis of KV-cache optimization as a memory-systems discipline.
Why memory bandwidth sets the speed limit
A simple calculation shows why bandwidth matters. At batch size 1, each decode step must read the model weights from GPU memory. An NVIDIA H100 SXM provides 3.35 TB/s of memory bandwidth (NVIDIA H100 product page, accessed Oct 9, 2026). Reading the roughly 14 GB of Llama 2 7B weights once takes about 4.2 ms, which caps a single stream near 240 tokens per second before KV-cache reads and software overhead are counted. On an H200 with 4.8 TB/s (NVIDIA H200 product page, accessed Oct 9, 2026), the same ceiling rises to roughly 340 tokens per second. These are theoretical upper bounds for illustration, not benchmarks.
The practical lesson is that serving one user at a time wastes most of a GPU. Batching lets many requests share each pass over the weights, so throughput rises until KV-cache memory or compute runs out. That is why the hardware comparison that matters most for inference is usually memory capacity and bandwidth, as covered in our H100 vs H200 comparison.
LLM inference metrics that matter
NVIDIA's benchmarking documentation defines the metrics most teams use (NVIDIA NIM LLM Benchmarking: Metrics, accessed Oct 9, 2026):
| Metric | What it measures | What drives it |
|---|---|---|
| Time to first token (TTFT) | Wait from request to first output token | Queuing, prefill, network |
| Inter-token latency (ITL), also called TPOT | Average gap between output tokens | Decode speed, memory bandwidth, KV-cache handling |
| End-to-end latency | TTFT plus generation time | Prompt length, output length, load |
| Tokens per second per system | Total output throughput across all requests | Batching, scheduling, hardware |
| Tokens per second per user | Output speed seen by one client | Approaches 1 / ITL for long outputs |
| Requests per second | Completed requests per second | All of the above |
The same documentation notes that as concurrency increases, total system throughput rises while per-user throughput falls, and that different tools define these metrics differently, so results are only comparable when definitions match. A single "tokens per second" number without concurrency, prompt length and output length is not enough to judge a system.
LLM inference optimization techniques
Continuous batching
Static batching makes every request wait for the longest one in the batch. Continuous, or in-flight, batching removes finished sequences after each iteration and adds new ones immediately, which NVIDIA says can greatly increase GPU utilization in real workloads (NVIDIA, Nov 17, 2023).
PagedAttention and KV-cache management
The vLLM team measured that earlier serving systems used only 20.4% to 38.2% of their KV-cache memory for actual token states, with the rest lost to reservation and fragmentation. PagedAttention stores the cache in fixed-size blocks, like virtual-memory pages, and vLLM reported 2 to 4 times higher throughput than FasterTransformer and Orca at the same latency (Kwon et al., SOSP 2023).
Attention variants and FlashAttention
Multi-query and grouped-query attention reduce how many key and value heads are stored, shrinking the KV cache. FlashAttention reorders the attention computation to make better use of the GPU memory hierarchy while remaining mathematically exact (NVIDIA, Nov 17, 2023).
Quantization
Lower-precision weights and activations take less memory and move faster over the same bandwidth, which helps bandwidth-limited decode. Activation outliers make activations harder to quantize than weights, so quality must be tested per model (NVIDIA, Nov 17, 2023).
Speculative decoding
A small draft model proposes several tokens and the large model verifies them in parallel. The original paper reported 2X to 3X faster decoding on T5-XXL with identical outputs, and notes the method helps most when memory bandwidth, not compute, is the bottleneck (Leviathan et al., "Fast Inference from Transformers via Speculative Decoding", arXiv, Nov 2022).
Model parallelism
Tensor and pipeline parallelism split a model across GPUs so larger models fit and per-device memory drops, at the cost of communication between devices (NVIDIA, Nov 17, 2023).
These gains are reported against specific baselines, models and workloads. They are not additive, and none should be read as a guarantee for a different deployment.
Hardware vs software: where capacity really comes from
Buying newer GPUs raises the memory and bandwidth ceiling. Software decides how close you get to that ceiling. Two nodes with identical GPUs can deliver very different useful output depending on batching, scheduling, KV-cache handling, decoding strategy and kernels. That is why the right question for an inference team is not only "which GPU?" but "how much of the GPU we already own are we using?"
Getting more inference from hardware you already own
EucaX Cortex is performance software that installs on top of existing AI infrastructure, from a single HGX H100, HGX H200 or B200 node to full fleets. EucaX states that operators typically unlock 3x or more serving capacity, up to 6x on HGX-class systems, depending on model, workload and configuration. There is no generic multiplier: Cortex is benchmarked against your own baseline on your own setup. If you want to see what your nodes can deliver, request a live test.
FAQ
What is the difference between LLM inference and training?
Training updates a model's weights using large datasets. Inference uses the finished, frozen weights to generate answers for new prompts. Training happens periodically; inference runs every time a user or agent sends a request, so it drives ongoing cost.
Why is LLM inference slow?
Output is generated one token at a time, and each step must read model weights and the KV cache from GPU memory. NVIDIA describes this decode phase as memory-bound, so memory bandwidth, batching and cache management usually matter more than peak compute.
What GPU do I need for LLM inference?
Start with memory: weights need about two bytes per parameter in 16-bit precision, plus KV cache that grows with context length and concurrency. Memory bandwidth then sets decode speed. An H100 SXM offers 80 GB at 3.35 TB/s; an H200 offers 141 GB at 4.8 TB/s.
How do you measure LLM inference performance?
Track time to first token, inter-token latency, end-to-end latency, tokens per second per system and per user, and requests per second, always at a stated concurrency, prompt length and output length so results are comparable.
Sources
- [1]NVIDIA Technical Blog, "Mastering LLM Techniques: Inference Optimization" (Verma and Vaidya), published Nov 17, 2023
- [2]Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention", SOSP 2023 (arXiv 2309.06180, Sep 2023)
- [3]Leviathan, Kalman and Matias, "Fast Inference from Transformers via Speculative Decoding", arXiv 2211.17192, Nov 2022
- [4]NVIDIA NIM LLM Benchmarking documentation, "Metrics", accessed Oct 9, 2026
- [5]NVIDIA H100 GPU product page and specifications, accessed Oct 9, 2026
- [6]NVIDIA H200 GPU product page and specifications, accessed Oct 9, 2026
See what this means for your infrastructure.
Book a discovery call and get a benchmark projection for your hardware and model.
SCHEDULE A CALL