Details
## Overview
`probe-image-size` scans the SVG header with a searching regular expression, `/<[-_.:a-zA-Z0-9][^>]*>/`. On input that contains many `<` characters but no `>`, the engine restarts the `[^>]*` scan at every `<` position and runs to end of input each time, giving quadratic time complexity.
Both the synchronous and the streaming parser are affected.
## Impact
Every entry point that reaches the SVG parser is affected: `probe.sync()`, `probe(stream)` and `probe(url)`. The URL form is the most exposed one — the input is fetched from a remote host, so an attacker only needs to supply a link.
Processing a crafted buffer blocks the Node.js event loop at 100% CPU for the whole duration. In production environments such as upload validators, image proxies or link unfurl services, a small number of concurrent requests is enough to deny service.
## Root Cause Analysis
Two independent problems.
1. **Absence of input size cap in the sync path.** `lib/parse_sync/svg.js` copied the entire buffer into a string and matched against it. There was no size limit at all, so cost scaled with the size of the attacker-supplied buffer.
2. **Repeated rescanning in the stream path.** `lib/parse_stream/svg.js` did cap accumulated data at 64 KB, but called `parseSvg(str)` on the whole accumulated string on *every* chunk, giving `O(chunks × N²)`. The cap does not help here: the more chunks the input is split into, the more times the quadratic scan is repeated.
Chunk size is influenced by the sender. `highWaterMark` (16 KB) is a buffering threshold, not a lower bound — a socket read returns whatever has arrived. A server that writes one byte at a time produces one-byte chunks; this was confirmed against the real `needle` pipeline with default options.
The original report identified (1) only, and stated that the 64 KB cap mitigates the streaming path. It does not.
## Proof of Concept (PoC)
Synchronous:
```js
const probe = require('probe-image-size')
// ~200 KB of '<a' — contains '<' but never '>'
probe.sync(Buffer.from('<a'.repeat(100000), 'latin1'))
```
Streaming — the same payload split into chunks, slower per byte than the synchronous form:
```js
const { Readable } = require('stream')
const probe = require('probe-image-size')
const payload = Buffer.from('<a'.repeat(32768), 'latin1')
const chunks = []
for (let i = 0; i < payload.length; i += 4096) chunks.push(payload.subarray(i, i + 4096))
await probe(Readable.from(chunks))
```
Measurements on the maintainer's machine:
| path | input | time |
| --- | --- | --- |
| `probe.sync()` | 25 KB | 0.9 s |
| `probe.sync()` | 50 KB | 5.5 s |
| `probe.sync()` | 100 KB | 18 s |
| `probe.sync()` | 200 KB | 54 s |
| `probe(stream)` | 64 KB, 1 chunk | 1.6 s |
| `probe(stream)` | 64 KB, 4 chunks | 2.9 s |
| `probe(stream)` | 64 KB, 16 chunks | 9.6 s |