Details
### Summary
An unauthenticated remote attacker can exhaust the memory of the vLLM API-server process by raising the request-level `media_io_kwargs.video.max_frames` and `fps` knobs on any deployment serving a Qwen2-VL or Qwen3-VL model. 74 extra bytes of JSON took the server's peak RSS from 2 271 MiB to 13 629 MiB over unauthenticated `POST /tokenize`.
The `num_frames` ceiling reported in GHSA-vxqj-p4gw-9h4c and fixed by open PR #51969 does not reach this path: the Qwen samplers do not read `num_frames` at all. The same knobs were already capped upstream for `GLMGAVideoBackend` as an accepted security fix in `8b6de0eb9` (PR #54935, merged 2026-09-04); that cap never reached Qwen.
### Details
#### Relationship to GHSA-vxqj-p4gw-9h4c and PR #51969 (read this first)
GHSA-vxqj-p4gw-9h4c reported that request-level `media_io_kwargs.video.num_frames` overrides the engine frame-count ceiling, and open PR #51969 fixes it by clamping `num_frames` inside `VideoMediaIO.merge_kwargs`.
That clamp does not reach the Qwen samplers. `Qwen2VLVideoBackend` and `Qwen3VLVideoBackend` do not read `num_frames` at all — `Qwen2VLVideoBackend`'s own docstring says so ("``num_frames`` is ignored (fps-driven, like the Qwen3-VL loader)"). They bound on `max_frames`, read from the same merged dict with no ceiling:
```python
# vllm/multimodal/video.py — Qwen3VLVideoBackend.compute_frames_index_to_sample
min_frames = kwargs.get("min_frames", 4)
max_frames = kwargs.get("max_frames", 768)
num_frames = int(total_frames_num / original_fps * fps)
num_frames = min(max(num_frames, min_frames), max_frames, total_frames_num)
```
With `max_frames` raised from the request, the only remaining bound is `total_frames_num` — every frame in the container.
I applied PR #51969's patch locally and re-ran both paths through the real merge layer (`merge_media_io_kwargs` → `VideoMediaIO.merge_kwargs` → `MediaConnector`), against `main` @ `b23433088`:
| request `media_io_kwargs.video` | merged kwargs after #51969 | frames decoded | peak RSS |
|---|---|---|---|
| *(absent)* | `None` | 32 | 706 MiB |
| `{"video_backend":"opencv","num_frames":-1}` | `{…,"num_frames":32}` | **32 — fixed** | 706 MiB |
| `{"video_backend":"qwen3_vl"}` | `{…,"num_frames":32}` | 60 | 779 MiB |
| `{"video_backend":"qwen3_vl","max_frames":1e9,"fps":1e6}` | `{…,"max_frames":1000000000,"fps":1000000,"num_frames":32}` | **900 — survives** | **2 994 MiB** |
#51969 does exactly what it claims for `num_frames`; the clamp writes `num_frames: 32` into the merged dict and the Qwen sampler ignores it, while `max_frames` and `fps` pass through untouched.
#### The codebase already has the fix pattern, on other backends
This is not a new control being proposed. Commit `8b6de0eb9` — "[Security] Cap GLMGA video sampling to prevent request-driven resource exhaustion" (PR #54935, merged 2026-09-04, same author as #51969) — caps precisely these two knobs for `GLMGAVideoBackend`:
```python
_MAX_FRAMES: ClassVar[int] = 640
_MAX_FPS: ClassVar[int] = 30
...
target_fps = min(target.fps, cls._MAX_FPS)
max_frames = min(kwargs.get("max_frames", cls._MAX_FRAMES), cls._MAX_FRAMES)
```
Its description states the root cause as "both `target_fps` and `max_frames` are controllable via request-level `media_io_kwargs`", and says it follows "the pattern established by `GLM46VVideoBackend`" (which caps via `_MAX_FRAME_COUNT_DYNAMIC = 640` and `_MAX_DURATION = 2400`).
So two backends cap request-controlled `fps`/`max_frames` as an accepted security measure. `Qwen2VLVideoBackend` and `Qwen3VLVideoBackend` — the most widely deployed video models on vLLM — cap neither. In GLMGA the uncapped knobs sized an intermediate *index list*; in the Qwen samplers they size the *decoded frame buffer*, which is larger by the per-frame pixel count.
#### Reachability — default configuration, no authentication
* `media_io_kwargs` is a request body field on `ChatCompletionRequest` and is carried by `/v1/chat/completions`, `/v1/embeddings`, `/v1/responses`, `/tokenize` and `/invocations`. No flag gates it.
* `api_key` defaults to `None` (`vllm/entrypoints/launchers/cli_args.py`), and `AuthenticationMiddleware` is installed only when a key is configured — a default `vllm serve` is entirely unauthenticated.
* Even with `--api-key` set, `GUARDED_PREFIX = ("/v1", "/v2", "/inference", "/cohere")` (`vllm/entrypoints/serve/middleware/authenticate.py:11`), so `/invocations` — which validates the same `ChatCompletionRequest` body — and `/tokenize` — which performs full media ingestion — remain unauthenticated.
* `MediaConnector.fetch_video` applies the model's registered sampler only when the request did not name one (`if "video_backend" not in video_io_kwargs`), so `video_backend: "qwen3_vl"` is selectable on any deployment; request-level selection of stock sampler subclasses is already established as reachable by GHSA-j682-9xp5-rrf3. On a Qwen deployment no `video_backend` key is needed at all.
* Default-on for any video-capable Qwen model.
**The Rust frontend is not affected.** `rust/src/server/src/routes/openai/chat_completions/validate.rs:90` rejects `media_io_kwargs` with "media_io_kwargs is not supported." No second front is needed.
### Impact
Unauthenticated remote denial of service by memory exhaustion of the API-server process. The decode runs in the frontend during chat parsing, before scheduling or admission control, so every tenant on the instance is affected. `--limit-mm-per-prompt` does not apply — it bounds media *items*, not frames within an item. CWE-770 / CWE-400.
Measured over unauthenticated HTTP: 1.43 MiB request body, 74 extra bytes of JSON, server peak RSS 2 271 → 13 629 MiB. In-process, the same request takes the decode from 60 to 900 frames. The frame count is bounded only by `total_frames_num`, which is the attacker's choice of video, and decoded bytes are `frames × H × W × 3`.
**Context, measured on the `num_frames` path (GHSA-vxqj-p4gw-9h4c's path), not this one** — these figures show what an unbounded frame count costs once the source video is chosen for it, and they transfer to this path because both converge on the same `_read_frames_no_recovery` allocation:
* 5.74 MiB request body → 9.27 GiB decoded, 9.73 GiB of new resident memory (1 736×).
* 117.6 MiB `video/jpeg` payload → process OOM-killed: `Out of memory: Killed process 182667 (python) total-vm:38665704kB, anon-rss:21121616kB`
The two were not re-run at the larger sizes through the Qwen sampler; the 900-frame figure above is what I measured on this path.
### Suggested fix
Extend the ceiling to the sampler-side knobs rather than clamping the single `num_frames` key. Two options, either acceptable:
1. **Per-backend caps, matching `8b6de0eb9`.** Give `Qwen2VLVideoBackend` and `Qwen3VLVideoBackend` the `_MAX_FRAMES` / `_MAX_FPS` treatment already applied to `GLMGAVideoBackend`, so `kwargs.get("max_frames", …)` and `target.fps` are clamped to class constants. Smallest change; consistent with the accepted precedent. It leaves `Glm5NextVideoBackend`, `Molmo2VideoBackend`, `NemotronVLVideoBackend`, `DynamicVideoBackend` and `OpenCVDynamicOpenPanguVideoBackend` to be audited one by one, which is the current trajectory (#55727, #56207, #56390).
2. **Strip the frame-count knobs at the merge boundary.** In `VideoMediaIO.merge_kwargs`, drop `max_frames` / `min_frames` / `fps` from `runtime_kwargs` the way `hw_decoders`, `pool_size`, `device` and unconfigured GPU backends are already dropped there. That treats the whole frame-count family as startup-only in one place and is robust to future sampler subclasses, at the cost of removing a request-level knob some users may rely on. A clamp-don't-strip variant (request may lower, never raise) preserves the feature.
Option 2 composes with #51969 and needs no per-backend audit; I would favour it, but option 1 is the more conservative change and matches what has already been merged.
### Affected versions
`>= 0.24.0`, through v0.29.1rc0 and `main` @ `b23433088`.
Lower bound established by probing release tags through the GitHub contents API; no version below is inferred, each was read out of the file at that tag.
**Confirmed present — Qwen samplers reading unclamped `max_frames`** (`max_frames = kwargs.get("max_frames", 768)` inside `Qwen2VLVideoBackend` / `Qwen3VLVideoBackend` in `vllm/multimodal/video.py`):
| ref | `Qwen3VLVideoBackend` present | unclamped `max_frames` |
|---|---|---|
| v0.23.0 | no (class does not exist) | n/a |
| v0.24.0 | yes | yes |
| v0.25.0, v0.26.0, v0.27.0, v0.28.0, v0.29.0, v0.29.1rc0 | yes | yes |
| `main` @ `b23433088` | yes | yes |
**Confirmed present — request-level `media_io_kwargs`** (field on the chat request model, and `VideoMediaIO.merge_kwargs` present): every ref probed, v0.19.0 through v0.29.1rc0 and `main` @ `b23433088`.
**Confirmed absent — any `num_frames` ceiling clamp** (PR #51969 unmerged): every ref probed, v0.19.0 through v0.29.1rc0 and `main` @ `b23433088`.
**Not resolved, and why:**
1. The introducing commit/PR for `Qwen2VLVideoBackend` / `Qwen3VLVideoBackend`. My clone is shallow (`--depth=300`), so `git log -S'max_frames'` cannot reach it; the v0.23.0 → v0.24.0 boundary above is the tightest bound I established by probing release tags.
2. The rc tags between v0.23.0 and v0.24.0, to tighten the lower bound to a specific release candidate.
3. Whether a `max_frames` knob on a *differently named* pre-v0.24.0 backend is separately affected — at v0.23.0 `GLMGAVideoBackend` already carried `max_frames = kwargs.get("max_frames", 640)` with no clamp, and that clamp was only added on 2026-09-04 by `8b6de0eb9`. Versions between are plausibly affected through GLMGA rather than Qwen; I did not test that path.
4. Whether GHSA-vxqj-p4gw-9h4c's own affected range differs, which I cannot see — the advisory returns 404 to me.
No dependency versions are asserted anywhere in this report.
### Weaknesses of this report, stated upfront
* **No GPU was used.** This host has no CUDA device, so the HTTP results come from the in-tree GPU-less render server rather than a full `vllm serve`. That server runs the real frontend — the same request model, the same `merge_media_io_kwargs` → `VideoMediaIO.merge_kwargs` → `MediaConnector.fetch_video` chain, the same `/tokenize` route — and media decoding happens entirely in the frontend, so I do not believe the engine's presence changes the result. I have not confirmed that on a GPU deployment, and a reviewer may reasonably want that repeated under `vllm serve`.
* The in-process measurements (the #51969 comparison table) were taken by importing the tree directly at `b23433088`, verified by `__file__`, with no `vllm` wheel installed.
* **Amplification figures are a floor.** The OpenCV build available here offers only `mp4v`/`XVID`, giving ~2 200× compression on static content. An attacker using x264/x265 would do materially better for the same frame count.
* The two large figures in the Impact section (9.73 GiB RSS; the OOM kill) were measured on the `num_frames` path, not this one, and are labelled as such.
* Per-frame pixels remain bounded by `VLLM_MAX_IMAGE_PIXELS`; this concerns the unbounded frame *count*, which multiplies it.
* `VLLM_MAX_MEDIA_DOWNLOAD_SIZE_MB` (default 256) caps the compressed size of an HTTP-fetched video but not a `data:` URI, which `MediaConnector.load_from_url` dispatches before any size logic is reached.
---
*Prepared with AI assistance (Claude), per the repository's contributing guidance on disclosing AI-assisted contributions. All findings were verified by executing vLLM's own code at the commits cited; the analysis and the claims are my own.*
*Reported by Eva Crystal / 0xiviel (XSource Security).*