# Solving LLM Serving Latency Interference

Via prefill-decode disaggregation and hybrid scheduling

Sep 28, 2026 · Inference · 6 min read · https://arclyx.ai/blog/solving-llm-serving-latency-interference

## Summary

As LLM prompt context lengths expand to tens and hundreds of thousands of tokens, serving architectures encounter an acute latency interference problem: compute-heavy prompt prefill bursts preempt and stall the strictly time-sensitive autoregressive decode iterations. This conflict creates erratic Time-Per-Output-Token (TPOT) spikes, while prioritizing decode severely inflates Time-To-First-Token (TTFT) [Sources 1, 4]. To resolve this operational deadlock, modern high-throughput inference frameworks are moving beyond monolithic batching toward Prefill-Decode (PD) disaggregation, chunked prefill, and hybrid latency-shifting architectures that optimize end-to-end "goodput" under stringent Service-Level Objectives (SLOs) [Sources 1, 3].

## Key Takeaways

- **Prefill vs. Decode Heterogeneity**: Prefill is compute-bound, saturating GPU tensor cores with parallel matrix multiplications over prompt tokens, while decode is memory-bandwidth-bound, repeatedly loading model weights and cached Key-Value states for single-token autoregressive generation [Sources 1, 4].
- **Latency Interference in Aggregated Serving**: In conventional co-located architectures, bursty prefill phases preempt or delay ongoing decode iterations, causing severe TPOT violations and inter-token streaming jitter [Sources 1, 4].
- **Architectural Disaggregation**: Physically separating GPU instances into dedicated Prefill clusters and Decode clusters eliminates cross-phase interference, but risks instance underutilization when workloads do not exhibit extreme single-metric skew [Source 1].
- **Hybrid Latency-Shifting (TaiChi)**: Recent research demonstrates that unifying aggregation and disaggregation through configurable instance ratios and request-level latency shifting (flowing decode and length-aware prefill) boosts hardware goodput by up to 77% under balanced SLO regimes [Source 1].
- **Production Cache & Memory Optimizations**: Industry systems leverage complementary techniques—Anthropic's prompt caching (cutting input latency by up to 85%), DeepSeek's Multi-Head Latent Attention (compressing KV cache footprints), and vLLM's PagedAttention (eliminating fragmentation)—to mitigate memory bandwidth saturation [Sources 2, 3, 4].

## Problem Background

Modern enterprise generative AI applications—ranging from multi-turn coding assistants to long-document analysis workflows—routinely process input prompts spanning 10,000 to over 100,000 tokens. Autoregressive transformer serving divides execution into two structurally mismatched phases:

1. **Prefill (Prompt Processing)**: The model processes all prompt tokens simultaneously in parallel matrix multiplications to generate the initial Key-Value (KV) cache and emit the first token. This phase is **compute-bound** and governed by the Time-To-First-Token (TTFT) Service-Level Objective (SLO) [Sources 1, 4].
2. **Decode (Token Generation)**: The model generates subsequent tokens sequentially, one by one. Each forward pass requires reading all model weights and the cumulative KV cache from High-Bandwidth Memory (HBM) to compute a single new token. This phase is **memory-bandwidth-bound** and governed by the Time-Per-Output-Token (TPOT) SLO [Sources 1, 4].

When prefill and decode are executed within the same GPU workers (Prefill-Decode Aggregation), an incoming long-prompt prefill burst consumes the streaming multiprocessors. Active decoding sequences batched on the same GPUs are forced to wait, causing severe TPOT latency spikes and noticeable streaming jitter for end users [Source 1].

Conversely, statically dedicating separate GPU pools to prefill and decode (Prefill-Decode Disaggregation) isolates execution and stabilizes TPOT, but introduces load-balancing fragility. If the incoming traffic shifts, one cluster starves while the other forms request queues, degrading TTFT and hardware goodput [Source 1].

## Proposed Experiments / Examples

To evaluate interference patterns and quantify the trade-offs between serving configurations, engineers can benchmark continuous batching, chunked prefill, and disaggregated topologies.

### 1. Workload Profiling Setup (Simulation)

```python
from dataclasses import dataclass

@dataclass
class ServingRequest:
    req_id: str
    prompt_tokens: int
    output_tokens: int
    ttft_slo_ms: float
    tpot_slo_ms: float

# Heterogeneous benchmark workload mixing long RAG contexts and conversational queries
benchmark_workload = [
    ServingRequest("rag_doc_1", prompt_tokens=32768, output_tokens=128, ttft_slo_ms=1500.0, tpot_slo_ms=40.0),
    ServingRequest("chat_turn_1", prompt_tokens=512, output_tokens=256, ttft_slo_ms=300.0, tpot_slo_ms=30.0),
    ServingRequest("code_gen_1", prompt_tokens=4096, output_tokens=512, ttft_slo_ms=600.0, tpot_slo_ms=35.0),
]
```

### 2. Deployment Configurations in vLLM

```bash
# Setup A: Aggregated continuous batching with chunked prefill
# Chunks long prompts into manageable token blocks (e.g., 2048) to avoid stalling decode steps
vllm serve meta-llama/Llama-3.1-70B-Instruct \
    --tensor-parallel-size 4 \
    --enable-chunked-prefill \
    --max-num-batched-tokens 2048 \
    --gpu-memory-utilization 0.90

# Setup B: Disaggregated serving topology
# Prefill-dedicated node
vllm serve meta-llama/Llama-3.1-70B-Instruct \
    --port 8100 \
    --tensor-parallel-size 4 \
    --gpu-memory-utilization 0.92 \
    --kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_producer"}'

# Decode-dedicated node (receives KV cache over high-speed interconnect)
vllm serve meta-llama/Llama-3.1-70B-Instruct \
    --port 8200 \
    --tensor-parallel-size 4 \
    --gpu-memory-utilization 0.92 \
    --kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_consumer"}'
```

### 3. Benchmarking Matrix: SLO Attainment & Latency

*Illustrative figures, not measurements: the table shows the expected shape of the trade-offs, not results from a benchmark run.*

| Serving Architecture | Mean TTFT (Prompt=32K) | Mean TPOT (Decode) | TPOT P99 Jitter | Goodput (% Requests Meeting All SLOs) |
|---|---|---|---|---|
| **Vanilla Continuous Batching** | 1,120 ms | 48 ms | ±34 ms (Severe preemption) | 56% (Frequent TPOT violations) |
| **Chunked Prefill (Chunk=2048)** | 1,320 ms | 28 ms | ±7 ms (Controlled) | 82% (Balanced execution) |
| **Static Disaggregation (Splitwise style)** | 1,650 ms | 22 ms | ±2 ms (Near-deterministic) | 77% (Prefill queue bottleneck) |
| **Unified Latency-Shifting (TaiChi)** | 1,220 ms | 24 ms | ±4 ms (Optimized) | **94%** (Optimal cross-request balance) |

## How Big Companies Solve This

Leading AI platforms and labs address KV cache scaling and latency contention across algorithmic, framework, and hardware layers:

1. **Anthropic (Prompt Prefix Caching)**: Anthropic introduced prompt caching for Claude API workloads, allowing developers to cache frequently reused contexts (such as system instructions and few-shot examples) [Source 2]. By reusing precomputed KV states across API calls, prompt caching reduces prompt processing latency by up to 85% and cuts input token costs by up to 90%, drastically reducing the prefill load on serving infrastructure [Source 2].
2. **DeepSeek (Multi-Head Latent Attention - MLA)**: Rather than expanding physical GPU memory footprint to accommodate massive KV caches, DeepSeek (in DeepSeek-V2, V3, and R1) introduced Multi-Head Latent Attention [Sources 3, 5]. MLA compresses keys and values into a low-rank latent representation during caching, which is projected back during attention computation. This reduces the per-token KV cache memory footprint by multiple folds compared to standard Multi-Head Attention, significantly alleviating memory bandwidth bottlenecks during autoregressive decoding [Sources 3, 5].
3. **vLLM & Academic Research (PagedAttention & System Co-design)**: vLLM introduced PagedAttention, an algorithm that partitions KV caches into non-contiguous memory blocks inspired by OS virtual memory paging [Source 4]. This eliminates external memory fragmentation and minimizes internal fragmentation (reducing memory waste from over 60% down to near-zero) and enables flexible cross-request cache sharing [Source 4]. Building on this, modern engines combine chunked prefill and disaggregated KV transfers over RDMA or NVLink to decouple prefill latency from decode responsiveness [Source 1].

## Discussion

The architectural shift in LLM serving reveals that hardware capacity alone cannot resolve latency contention. Even the fastest accelerators face a structural dichotomy between compute-bound matrix multiplications and memory-bound autoregressive decoding.

While static Prefill-Decode disaggregation provides clean physical isolation, it imposes an infrastructure tax through underutilized GPU instances when request arrival patterns vary. As demonstrated by recent research in unified architectures like TaiChi, the next frontier in AI inference engineering is dynamic, SLO-aware scheduling [Source 1]. By treating prefill chunk size and cluster allocation as dynamic control sliders, serving systems can shift latency margins from requests that easily meet their deadlines toward requests at risk of SLO violations [Source 1]. When paired with architectural KV compression like MLA and memory management like PagedAttention, inference clusters can achieve both high token throughput and predictable, low-jitter latency [Sources 3, 4].

## Conclusion

Resolving LLM serving latency interference requires a multi-layered engineering approach. At the model architecture level, low-rank attention mechanisms like MLA shrink the physical memory footprint of cached states [Source 3]. At the memory runtime layer, virtualized paging mechanisms like PagedAttention maximize GPU memory density [Source 4]. At the cluster scheduling layer, chunked prefill and hybrid prefill-decode disaggregation prevent compute-heavy prefill bursts from degrading interactive streaming latency [Source 1]. Together, these techniques transform LLM serving from a rigid, monolithic batching process into an elastic, SLO-driven distributed system.

## Sources

1. Wang et al. (2025). *Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving*. arXiv:2508.01989. [https://arxiv.org/html/2508.01989v1](https://arxiv.org/html/2508.01989v1)
2. Anthropic. (2024). *Prompt caching with Claude*. Anthropic Engineering & News. [https://www.anthropic.com/news/prompt-caching](https://www.anthropic.com/news/prompt-caching)
3. Meng et al. (2025). *TransMLA: Multi-head Latent Attention Is All You Need*. arXiv:2502.07864. [https://arxiv.org/html/2502.07864v1](https://arxiv.org/html/2502.07864v1)
4. Kwon et al. (2023). *Efficient Memory Management for Large Language Model Serving with PagedAttention*. Proceedings of ACM SOSP '23 / arXiv:2309.06180. [https://ar5iv.labs.arxiv.org/html/2309.06180](https://ar5iv.labs.arxiv.org/html/2309.06180)
5. DeepSeek-AI et al. (2024). *DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model*. arXiv:2405.04434. [https://arxiv.org/abs/2405.04434](https://arxiv.org/abs/2405.04434)
