Which host does this URL point to?
>>> from urllib.parse import urlsplit
>>> u = "https://api.trusted-weather.com:443@evil.com/data"
>>> urlsplit(u).netloc.split(":")[0] # what our consent screen used
'api.trusted-weather.com'
>>> urlsplit(u).hostname # what our provisioning code used
'evil.com'
The second answer is the right one. The host is evil.com, and everything before the @ is a username and a password. Our pre-release review found that the screen where an app owner approves a domain read that URL one way, and the code that wires the API key read it the other way. The URL is written by a model, so the attacker only needs to reach the text the model reads. The fix is two rules: reject any URL that carries userinfo, and build what you display from the same parsed field as what you act on.
The flow that was under review
The code under review belongs to an extension builder that my team at GoodBarber had not yet switched on in production. Our own review found the defect below, and the fix landed before activation. In that builder, a model writes a widget. When the widget needs a third-party API with a private key, the model declares the API and its URL. The app's owner sees a consent screen that names the domain, approves it, and pastes the key into a masked field. From then on a proxy of ours holds the key and injects it into every call the widget makes to that host.
The consent screen is the owner's only look at where the key will go. On July 22, eight days before the feature was switched on in production, a review of the whole branch came back with a verdict of not ready. This was its first finding.
One URL, two interpretations
Go back to the two answers at the top of this page.
The display path took the raw network location and cut it at the first colon to drop the port. That line was written with host:443 in mind. On this URL it returns the username. The result was then reduced to its registrable domain, and the owner read trusted-weather.com on the consent screen.
The provisioning path asked the standard library for the hostname and got the correct answer: in that URL, api.trusted-weather.com is a username and 443 is a password.
One reading was wrong and one was right. The danger is that they disagreed. The owner approves a domain they trust, pastes a live key, and the proxy injects that key into every request to evil.com. The SSRF checks on the proxy do not help. evil.com is a public, routable host with a valid certificate, which is exactly what those checks allow.
An old trick with a new author
None of this is new. RFC 3986 described it in 2005, in its section on semantic attacks: because userinfo is rarely used and sits before the host, it can build a URI that appears to name one trusted authority while identifying another. Orange Tsai's 2017 talk on URL parsers that disagree turned that family of differences into a method.
What is new is who writes the URL. In our flow it is not the owner and it is not us. It is a model, and a model writes from what it reads: a prompt, a pasted snippet, a template, a page of documentation. OWASP calls the resulting risk indirect prompt injection: content from an external source changes what the model produces. The attacker does not need an account on our platform. They need one sentence, somewhere the model will look, that says the weather API lives at that address.
To be precise about what was tested: the disagreement between the two functions was reproduced, with running code. The injection path is a scenario. Nobody demonstrated a poisoned document steering our model into writing that URL, and the fix does not depend on how the string arrives.
The misled human of the RFC used to be someone squinting at a link. Here it is someone doing the right thing, on a screen built for that purpose, reading a domain our own code printed. A consent screen is a security boundary. It has a parser, and it had a different one from the code it was guarding.
The fix, and when it landed
Two rules, committed on July 30 at 12:44.
Reject what you do not need. A URL with userinfo is refused at validation and again at provisioning, including the bare @ with nothing before it:
def url_has_userinfo(url: str) -> bool:
try:
parsed = urlsplit(url.strip())
except ValueError:
return True
return bool(parsed.username or parsed.password or "@" in parsed.netloc)
No API we integrate needs credentials in the authority part of a URL, and RFC 3986 itself deprecates the user:password form. When a syntax is rare and dangerous, interpreting it carefully is the wrong ambition. Refuse it.
One parse, many readers. The host shown on the consent screen is now built from the parsed hostname and port, never from the raw network location, so the approved host and the provisioned host are byte-identical by construction. The same commit binds a key to the root domain the owner approved, so a generic key name reused by two APIs cannot be provisioned onto a host nobody approved.
Nine regression tests came with it. The configuration change that switched the feature on in production is dated the same day, 14:01.
How the review found it
The method matters more than the bug, because the bug is one line and the method finds the next one.
The branch was 49 files and 6,707 added lines. Seven reviewer agents read it, each with one specialty, followed by an adversarial pass whose only job was to break things. That produced claims. Claims are cheap, so the second round was verification: ten claims, one verifier per claim, each returning CONFIRMED, PLAUSIBLE or REFUTED, with the code quoted. This finding was not accepted because it sounded right. It was accepted because a verifier ran the two functions under Django on the crafted URL and printed two different hosts.
A finding that is not reproduced is an opinion. That rule cuts both ways. One claim did not survive its verifier and sits in the report as refuted, with the reason. And "not ready" became impossible to argue with.
The test worth writing
The lessons are in the sections above. What I would add to any product with an approval screen is one test, and it is not about userinfo. It compares the two ends: what a person is shown, and what the network layer will use.
HOSTILE = [
"https://api.trusted-weather.com:443@evil.com/data",
"https://api.trusted-weather.com@evil.com/data",
"https://@evil.com/data",
]
def test_approved_host_is_the_provisioned_host():
for url in HOSTILE + ORDINARY_URLS:
if url_has_userinfo(url):
continue # refused before anyone is asked to approve it
assert displayed_host(url) == provisioned_host(url)
displayed_host and provisioned_host stand for whatever your screen prints and whatever your HTTP client connects to. If that assertion is awkward to write because the two values come from different code, you have found the bug before a reviewer does.
Top comments (1)
Requiring an executable reproduction test under Django before accepting the reviewer agent's claim is the detail that makes this work. Automated security reviewers spit out endless plausible-sounding hallucinations; forcing the verifier to run both functions on the payload turns review noise into an immediate blocker.
On the parser side, the subtle trap that often follows this fix is re-serializing the URL back into a string for the downstream HTTP client. If Python's urlsplit and whatever library actually opens the socket disagree on IDNA punycode normalization or trailing dots, the string boundary can still desync. Passing the pre-validated hostname and port directly into the socket transport or an egress allowlist avoids parsing the raw string twice.