Details
### Summary
When remote fetching is enabled (`enable_remote_fetch=True`, together with `fetch_images=True` or HTML `render_page=True`), docling's protection against requests to private and loopback addresses can be bypassed.
- The URL check resolves the host name once, to a single IPv4 address. The HTTP request then resolves it again. A name that returns a public A record together with an internal AAAA record, or one that changes its answer between the two lookups (DNS rebinding), reaches internal addresses.
- The check and the HTTP request parse the URL differently. With a backslash in the authority, such as `http://127.0.0.1:8080\@1.1.1.1/x.png`, the check validates `1.1.1.1` while the HTTP client connects to `127.0.0.1:8080`.
- In HTML browser-rendering mode (`render_page=True`), requests made by the page are allowed for any `http(s)` URL without an IP check.
### Details
- `validate_url_safety` in `docling/backend/utils/image_resource_loader.py` calls `socket.gethostbyname()`, which returns one IPv4 address. It checks that address and then passes the original URL string to `requests`, which parses it again with urllib3 and resolves the name independently through `getaddrinfo`. `urllib.parse` treats a backslash in the authority as part of the host section, while urllib3 ends the authority there, so the two can pick different hosts. This loader is shared by the HTML, Markdown, AsciiDoc, OpenDocument and JATS backends.
- `docling/backend/html_backend.py` routes browser requests in render mode. When `enable_remote_fetch` is set, it continues any `http(s)` request, including iframes and images, without resolving or checking the destination.
JavaScript is disabled during rendering, so the browser path can only show internal responses passively in the page screenshot. It cannot read them and send them elsewhere.
### Affected configurations
- Affected: callers that enable `enable_remote_fetch` (CLI: `--html-image-fetch remote` or `all`) and process untrusted documents.
- Not affected: the default configuration (`enable_remote_fetch=False`).
The main-URL fetch for `DocumentConverter.convert("https://...")` is implemented in docling-core and is tracked there.
### Impact
Requests from the converting host to internal services such as cloud metadata endpoints and loopback services. Response content is exposed only when it decodes as an image or is rendered into the page screenshot.
### Patches
Fixed in docling 2.132.0 by [#4420](https://github.com/docling-project/docling/pull/4420). The image loader now resolves the host once, requires every resolved address (IPv4 and IPv6, including IPv4-mapped, 6to4 and NAT64 forms) to be globally routable, and connects to one of the validated addresses, using the same parsed host for the check and the connection. Remote requests made during HTML browser rendering are downloaded by the same loader, and the browser itself stays offline. When a proxy is configured through environment variables, requests go through the proxy, which is then responsible for filtering destinations.
### Workarounds
Upgrade to 2.132.0. For older versions:
Keep `enable_remote_fetch=False` for untrusted documents, or block egress to private, link-local and loopback ranges at the network level.
EPSS, exploit probability
Low0.19%
estimated chance of real-world exploitation in the next 30 days, higher than 7.9% of every CVE FIRST.org scores
Refreshed 10/7/2026, via FIRST.org's EPSS model, not CVSS, this measures likelihood of exploitation, not how severe it would be.