Details
## Affected
- **Ecosystem / package:** pip / `vllm`
- **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the tool-continuation re-submission has omitted the salt.
## Summary
On the GPT-OSS "Harmony" path (`POST /v1/responses`), a request that uses a built-in or MCP tool runs as a multi-turn loop: after each tool call vLLM re-renders the full next-turn Harmony prompt and re-submits it to the engine. Turn 1 correctly carries `request.cache_salt`, but the tool-continuation re-submission rebuilds the engine input via `tokens_input(token_ids)` with **no** `cache_salt`. The continuation prefix is therefore cached in the global *unsalted* namespace even though the caller opted into salting. A second tenant who can guess the low-entropy post-tool history submits the reconstructed continuation (unsalted) and reads exact per-turn cached-token counts from the Responses usage — restoring the prompt-membership oracle that `cache_salt` is documented to prevent.
Silently dropping a preserved salt *after* the supported tool workflow is enabled is a broken isolation control: the caller enabled salting and every turn should stay isolated, but continuation turns leak into the shared cache.
This is distinct from [GHSA-4qjh-9fv9-r85r](https://github.com/vllm-project/vllm/security/advisories/GHSA-4qjh-9fv9-r85r) ([CVE-2025-46570](https://nvd.nist.gov/vuln/detail/CVE-2025-46570)): that advisory is the prefix-cache membership oracle for which `cache_salt` is the documented mitigation, and its PR-17045 fix does not close this site — the Harmony tool continuation silently drops the preserved salt, caching in the unsalted namespace and leaking exact `cached_tokens_per_turn` counts from a different sink (the Responses serving continuation, not general TTFT timing).
## Affected code
Links pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1):
- **The drop (sink):** [`vllm/entrypoints/openai/responses/serving.py#L712-L713`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/responses/serving.py#L712-L713) — `token_ids = context.render_for_completion()` then `engine_input = tokens_input(token_ids)`, with no `cache_salt`.
- **Correct turn-1 call for contrast:** [`vllm/entrypoints/openai/responses/serving.py#L755`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/responses/serving.py#L755) — `tokens_input(prompt_token_ids, cache_salt=request.cache_salt)`.
- **`tokens_input` stores the salt only if passed:** [`vllm/inputs/engine.py#L51-L66`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/inputs/engine.py#L51-L66) (`if cache_salt is not None: inputs["cache_salt"] = cache_salt`).
- **The engine request copies only the current input's salt:** [`vllm/v1/engine/input_processor.py#L380`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/engine/input_processor.py#L380) (`cache_salt=decoder_inputs.get("cache_salt")` → `None` for the continuation).
- **Prefix-cache hashing keys on the salt only when present:** [`vllm/v1/core/kv_cache_utils.py#L560-L561`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/core/kv_cache_utils.py#L560-L561) (`[request.cache_salt] if (start_token_idx == 0 and request.cache_salt) else []`).
- **The oracle the attacker reads:** [`vllm/entrypoints/openai/responses/serving.py#L909`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/responses/serving.py#L909) (`cached_tokens_per_turn`).
- **The documented control being defeated:** [`vllm/entrypoints/openai/responses/protocol.py#L235`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/responses/protocol.py#L235) (`cache_salt` field).
The tool-continuation re-submission rebuilds the engine input with no `cache_salt`:
```python
# vllm/entrypoints/openai/responses/serving.py Lines 711-715
if isinstance(context, HarmonyContext):
token_ids = context.render_for_completion()
engine_input = tokens_input(token_ids)
sampling_params.max_tokens = max_model_len - len(token_ids)
```
Contrast with the correct turn-1 call, which does preserve the caller's salt:
```python
# vllm/entrypoints/openai/responses/serving.py Lines 754-755
prompt_token_ids = render_for_completion(messages)
engine_input = tokens_input(prompt_token_ids, cache_salt=request.cache_salt)
```
`tokens_input` stores the salt on the engine input only when it is passed, so the continuation input carries none and lands in the unsalted namespace:
```python
# vllm/inputs/engine.py Lines 51-66
def tokens_input(
prompt_token_ids: list[int],
*,
prompt: str | None = None,
cache_salt: str | None = None,
) -> TokensInput:
"""
Construct [`TokensInput`][vllm.inputs.engine.TokensInput]
from optional values.
"""
inputs = TokensInput(type="token", prompt_token_ids=prompt_token_ids)
if prompt is not None:
inputs["prompt"] = prompt
if cache_salt is not None:
inputs["cache_salt"] = cache_salt
```
## Impact
An authenticated tenant of a shared deployment can recover whether a guessed post-tool prompt or history was processed by another tenant, with exact cached-token counts rather than noisy latency — the exact prompt-membership oracle `cache_salt` is documented to prevent. It defeats the multi-user prefix-cache isolation guarantee for salted Harmony tool sessions.
Preconditions: a GPT-OSS Harmony model on `/v1/responses`; prefix caching enabled (default); an operator-enabled built-in or MCP tool server; the victim sets `cache_salt` and triggers at least one tool continuation; and the attacker can reconstruct the post-tool history closely enough to match the token prefix. The `AC:H` metric reflects that guessable-history precondition.
## Suggested Fix
Propagate `request.cache_salt` into every Harmony (and Parsable) tool-continuation re-submission — at the continuation call site call `tokens_input(token_ids, cache_salt=request.cache_salt)`, mirroring the correct turn-1 call. Carry the salt on the `HarmonyContext` (thread the originating `request` into the context) so no continuation path can omit it:
```diff
# vllm/entrypoints/openai/responses/serving.py
if isinstance(context, HarmonyContext):
token_ids = context.render_for_completion()
- engine_input = tokens_input(token_ids)
+ engine_input = tokens_input(
+ token_ids,
+ cache_salt=(
+ context.request.cache_salt
+ if context.request is not None
+ else None
+ ),
+ )
```
with `HarmonyContext.__init__` gaining a `request: ResponsesRequest | None = None` parameter (stored as `self.request`) that `_create_responses` passes when constructing the context. The continuation prefix is then cached in the victim's salted namespace, mirroring turn 1.
Suggested regression test: assert `cached_tokens_per_turn == 0` for a different-salt probe against a salted victim continuation (the four-way control from the proof of concept).
## Credit
**Reported by:** Patch the Planet (Trail of Bits + OpenAI collaboration)
This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
---
**Proposed fix:** a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51818