Details
## Affected
- **Ecosystem / package:** pip / `vllm`
- **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the flash late-interaction query cache reaches.
## Summary
On late-interaction `/score` and `/rerank` deployments with flash late interaction enabled (the default for supported models), the worker caches per-request query embeddings under a key derived from the **caller-controlled** `X-Request-Id` header. A second concurrent request that reuses the victim's header value replaces the victim's cached query embedding before document scoring — so the victim's documents are scored against the **attacker's** query. Because the data-parallel router pins all requests sharing a cache key to the same engine, the collision is deterministic for an attacker who reuses the victim's `X-Request-Id`. Depending on timing, one request can also consume the shared use counter and force the other request into a late-interaction cache-miss error.
This is a remotely reachable, request-controlled cross-request integrity break on the standard scoring and reranking endpoints. It requires only that flash late interaction be enabled, which is the default for supported models.
## Affected code
Links pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1):
- [`vllm/entrypoints/serve/engine/serving.py#L117-L124`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/serve/engine/serving.py#L117-L124) — `_base_request_id()` copies the public `X-Request-Id` header directly.
- [`vllm/entrypoints/pooling/base/serving.py#L109`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/pooling/base/serving.py#L109) — the frontend request id is `f"{self.request_id_prefix}-{self._base_request_id(raw_request)}"`.
- [`vllm/entrypoints/pooling/scoring/serving.py#L211`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/pooling/scoring/serving.py#L211) — `flash_late_interaction()` (at [L191](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/pooling/scoring/serving.py#L191)) derives worker cache keys directly from that id: `query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]`.
- [`vllm/v1/pool/late_interaction.py#L30-L36`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/pool/late_interaction.py#L30-L36) — the data-parallel routing helper pins all requests sharing a `query_key` to the same engine via `crc32(query_key)`, making collisions deterministic.
- [`vllm/v1/worker/gpu/pool/late_interaction_runner.py#L95`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/worker/gpu/pool/late_interaction_runner.py#L95) — the worker stores query embeddings in a process-local cache keyed only by that string: `self._query_cache[query_key] = output.clone()`.
The caller-controlled header enters as the request id, and the flash late-interaction path derives the worker cache key directly from it:
```python
# vllm/entrypoints/serve/engine/serving.py Lines 116-126
@staticmethod
def _base_request_id(
raw_request: Request | None, default: str | None = None
) -> str | None:
"""Pulls the request id to use from a header, if provided"""
if raw_request is not None and (
(req_id := raw_request.headers.get("X-Request-Id")) is not None
):
return req_id
return random_uuid() if default is None else default
```
```python
# vllm/entrypoints/pooling/scoring/serving.py Lines 207-212
n_queries = ctx.n_queries
n_docs = len(ctx.engine_inputs) - n_queries
query_engine_inputs = ctx.engine_inputs[:n_queries]
query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]
query_uses = [n_docs if n_queries == 1 else 1] * n_queries
```
The worker then stores and reads the query embedding under that string with no check that the reader owns the entry — a colliding key returns another request's cached query, or (once the use counter is exhausted) raises a cache-miss error:
```python
# vllm/v1/worker/gpu/pool/late_interaction_runner.py Lines 91-107
if mode == LATE_INTERACTION_MODE_CACHE_QUERY:
assert query_uses is not None
# `output` can be a view into the current step's hidden-states
# buffer, so clone it before storing across scheduling steps.
self._query_cache[query_key] = output.clone()
self._query_uses[query_key] = query_uses
outputs[i] = torch.zeros((), device=output.device, dtype=torch.float32)
continue
if mode == LATE_INTERACTION_MODE_SCORE_DOC:
query_output = self._query_cache.get(query_key)
if query_output is None:
raise ValueError(
"late-interaction query cache miss for key "
f"{query_key!r}. Ensure query requests are executed "
"before their paired document requests."
)
```
The bug is specific to the flash late-interaction path. Non-flash late-interaction scoring computes MaxSim directly from one request's in-memory outputs and does not create a cross-request worker cache key.
## Impact
A network client of the standard scoring API can, on a flash late-interaction `/score` or `/rerank` deployment:
1. **Corrupt another user's results** — by reusing the victim's `X-Request-Id`, the attacker's query embedding overwrites the victim's cached entry, so the victim's documents are scored against the attacker's query (a cross-request integrity break).
2. **Induce errors** — depending on timing, one request consumes the shared use counter and forces the other request into a late-interaction cache-miss error.
Both consequences follow deterministically from reusing the victim's header value, because same-key work is pinned to one engine. This is reachable through normal request handling and does not depend on any trusted inter-node network.
## Suggested Fix
Derive the flash late-interaction query-cache key from a server-generated, unforgeable per-request identifier (a `random_uuid()` namespace) rather than the caller-supplied `X-Request-Id`, and thread that key through the `PoolingServeContext` to the doc-scoring pass so both passes reuse the same key and a caller cannot address another request's cache entry.
In `vllm/entrypoints/pooling/scoring/serving.py`, the encode-queries pass mints a fresh namespace and stashes the keys on the context:
```python
- query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]
+ query_namespace = random_uuid()
+ query_keys = [
+ f"late-interaction-{query_namespace}-query-{i}" for i in range(n_queries)
+ ]
+ ctx.late_interaction_query_keys = query_keys
```
and the encode-docs pass reads those stored keys instead of re-deriving them from `ctx.request_id`:
```python
- query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]
+ query_keys = ctx.late_interaction_query_keys
+ if query_keys is None:
+ raise RuntimeError("Late-interaction query keys were not initialized.")
```
This requires adding the `late_interaction_query_keys: list[str] | None = None` field to `PoolingServeContext` (`vllm/entrypoints/pooling/typing.py`). Because the namespace is a server-generated UUID, colliding `X-Request-Id` values no longer produce a shared cache key; a regression test asserting exactly that (colliding request ids yield distinct query-cache keys) accompanies the change.
## Credit
**Reported by:** Patch the Planet (Trail of Bits + OpenAI collaboration)
This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
---
**Proposed fix:** a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51445