DEV Community

WabaLabaDubDub
WabaLabaDubDub

Posted on

🀯 I Thought I Found an LLM Vulnerability. I Was Wrong. Then I Found a Real One.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

A retraction, cross-user RAG chains, and a jailbreak that transfers between models

I built a sandboxed AI red-team lab to answer a simple question: can I actually find and validate security issues in an LLM application myself?
The project produced several confirmed findings, including a cross-user RAG authorization chain in Open WebUI and a manually developed jailbreak technique that transferred between model families.
It also produced something I did not expect: a false positive.
I initially reported a cross-user chat access vulnerability. During re-testing, I discovered that my attacker credential was invalid. I retracted the finding, documented why it happened, and changed my testing methodology around the mistake.
That retraction became one of the most useful results of the project.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

Why build a lab instead of reading about AI security?

Every week there is another article about prompt injection, jailbreaks, or LLM application vulnerabilities.
Reading them taught me the vocabulary.
It didn’t teach me the thing I actually wanted to know:
Can I find these bugs myself?
So I built a lab.
Not a tutorial environment. My own.
A Mac running Ollama served three local models: llama3.2:3b, mistral:7b, and qwen3:8b. Open WebUI v0.3.16 and a Kali Linux attack environment ran in Docker on an isolated lab network. Testing used synthetic accounts and controlled payloads.
The environment was deliberately designed so that it did not become another always-running service on my machine. The model endpoint was restricted to loopback access, and the lab had explicit start and stop procedures with verification checks.
Two constraints shaped the project.
The target was pinned. I wanted a known application version rather than a moving target. That made it possible to reproduce behavior against the same build and ask a more useful question: where do the application’s authorization boundaries actually hold?
The lab was a switch, not a service. lab-start brought the environment up and verified it. lab-down tore it down and verified that the expected services were no longer listening.
The architecture, tooling, operational scripts, findings, and evidence are available in the companion repository:

https://github.com/soluckyarya2-tech/ai-red-team-lab

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

Methodology β€” written by my own mistakes

The workflow that eventually produced every finding in this project was:

  1. Validate the attacker identity and token
  2. Establish a baseline
  3. Attack
  4. Run a control
  5. Re-test the authentication, authorization, and reproduction conditions
  6. Document the result β€” including what failed Step 1 exists because of the story in the next section. Step 6 exists because a findings list without evidence is just a list of claims. For authorization testing, I wanted every result to answer the same questions:
  7. Who is making the request?
  8. Is the credential actually valid?
  9. What can the legitimate owner do?
  10. What can the unauthorized user do?
  11. Does the response demonstrate access to the protected resource?
  12. Can the result be reproduced independently? That process became more important than any individual tool.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

The application audit

I tested Open WebUI’s API surface with two synthetic standard users.
The important boundary was simple: one account owned the resource and the other did not.
For readability, I refer to them by their boundary role in each test rather than assuming that one account was permanently the attacker or victim.

F1 β€” The files router holds the boundary
The first useful result was actually a negative one.
Cross-user file access was denied β€” a 404 that doesn’t distinguish β€œmissing” from β€œforbidden,” and a user-scoped listing that never showed the victim’s file.
There was no demonstrated metadata leak or enumeration oracle.
This established an important baseline:
the application knew that the file belonged to another user.
The next question was whether every subsystem operating on that file enforced the same ownership boundary.
It didn’t.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

F2 β€” The false positive I had to retract

The chat authorization test became the project’s most important methodological lesson.
My initial test appeared to show a cross-user chat access vulnerability.
I observed a 401 response and then a 200 response from a related control request. I interpreted the sequence as evidence that an attacker could access another user’s chat.
I was wrong.
The 200 response came from the legitimate control case: the owner accessing their own chat.
The attacker request was being made with an invalid credential.
The root cause was an attacker token that predated a JWT signing-secret rotation. Its payload looked valid, but its signature was no longer valid.
The chat authorization control had been working correctly.
I re-tested the scenario with validated credentials.
The unauthorized request was denied.
The finding was retracted.
That changed my methodology permanently:
No security verdict until the attacker credential has been independently validated.
An authentication failure and an authorization bypass look identical if you never establish the authentication state first.
The re-test also exposed different denial semantics across the application’s routers. Malformed credentials, unauthorized chats, and unauthorized files produced different response patterns.
That behavior was not itself treated as a vulnerability. It was useful as a testing fingerprint because it helped distinguish authentication failures from resource-ownership failures.
The repository keeps the original claim, the retraction, the root-cause analysis, and the corrected result.
I don’t think a security research history should only contain the things that worked.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

F6 β€” Cross-user RAG ingestion

The real application-layer finding appeared when I followed the file into the RAG pipeline.
Open WebUI processes uploaded documents into collections used for retrieval-augmented generation.
The ingestion endpoint:
POST /rag/api/v1/process/doc
accepts a file ID and collection information.
The files router had already demonstrated that it would deny an unauthorized user direct access to another user’s file.
I tested whether the RAG ingestion path enforced the same ownership boundary.
The sequence was:

  1. One test account uploaded a file containing a unique marker:
TOPSECRET-TEST2-4821
  2. The other account attempted direct access to the file.
  3. Direct access returned 404.
  4. The unauthorized account submitted the owner’s file ID to /process/doc.
  5. The processing request returned 200.
  6. The document was ingested into a collection.
  7. The unauthorized account queried the collection.
  8. The response contained the document content, including the unique marker. The marker is important. It existed only inside the uploaded file. Its appearance in the unauthorized query result therefore demonstrated content-level disclosure rather than simply exposing metadata about another user’s resource. The finding was not merely that an unauthorized user could reference another user’s file ID. The demonstrated behavior was: an authenticated user could cause another user’s uploaded document to be processed through the RAG pipeline and subsequently retrieve its content. The direct file authorization boundary existed. The RAG ingestion path did not enforce the same boundary.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

F7a β€” Collection ownership failure

The collection query layer introduced a second authorization failure.
The collection query endpoint accepted collection names without enforcing ownership of those collections.
The collection_names parameter also accepted multiple collection names in a single request.
Testing showed that valid collection names produced their associated contents, while invalid names failed without providing equivalent content.
That created a positive-signal enumeration mechanism.
At this point, an attacker did not need direct file access.
The collection itself became another path to the underlying document content.
The result was a distinct authorization finding:
the collection query layer did not enforce the ownership boundary established by the application’s file layer.
The two findings could then be chained conceptually:
cross-user file β†’ unauthorized ingestion β†’ collection β†’ collection query β†’ document disclosure
The application’s direct file controls therefore did not protect the same underlying data once it entered the RAG subsystem.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

F7b β€” Retrieved content becomes an instruction surface

The next experiment connected the application layer to the model layer.
I created a document containing legitimate content and an embedded directive.
When a chunk containing the directive was retrieved into the model’s context, the model followed it. Where the directive sat in the document decided whether that ever happened β€” see below.
The resulting behavior demonstrated an indirect prompt-injection path:
attacker-controlled document β†’ RAG ingestion β†’ retrieval β†’ model context β†’ instruction execution
The position of the instruction mattered.
When the directive was placed in a portion of the document that was not retrieved, it was not executed.
That was a retrieval miss, not evidence of model resistance.
When the directive was moved into the first retrieved chunk, it was executed on the first attempt.
This makes the result more interesting than a simple prompt-injection demonstration.
It shows why RAG security cannot be treated only as a model-prompt problem.
Retrieval policy, document ownership, chunking, content trust, and model instruction-following all become part of the attack surface.
In this case, the authorization failure and the indirect prompt-injection behavior existed across the same pipeline.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

The model attacks β€” what the scanner found, and what it didn’t

After the application testing, I moved to the models.
I used garak to establish an automated baseline against the local Ollama endpoint.
The result was useful because it showed both the value and the limitation of automated scanning.
For llama3.2:3b, the tested verbatim DAN prompts were blocked in the baseline: 0/15 successful attacks.
Perturbed DAN variants, however, produced a substantially different result: only 72 of 635 attempts were resisted β€” an observed attack success rate of 88.66% for that probe configuration.
This distinction matters because Garak reports ok on 72/635 to indicate that 72 attempts passed the detector β€” in this case, meaning those attempts were resisted. The remaining 563 attempts produced the attack condition measured by the probe.
The scanner was useful for showing that small changes to a known jailbreak family could substantially change the result.
But the more interesting question was what happened outside the scanner’s existing templates.
That is where manual testing started.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

Fabricated conversation history

The technique I developed was based on fabricated assistant history.
Instead of directly asking the model for prohibited content, I supplied a conversation context containing a fabricated assistant turn that appeared to have already started producing the requested content.
The attack then asked the model to continue.
The prohibited request did not appear in the expected direct form.
The model was being asked to continue what it believed was its own previous output.
Against llama3.2:3b, three direct attack frames were refused during testing.
The fabricated-history variant produced compliant output on the first tested attempt.
I then used the mechanism with a phishing-oriented payload to measure its stochastic behavior.
Across 20 identical inputs, 3 produced the target behavior:
3/20 = 15% observed ASR.
If that observed probability were stable and each attempt were independent, the expected number of attempts to obtain one success would be approximately 6.7.
That is a mathematical expectation, not a measured attacker requirement.
The important observation was that the attack did not require a known jailbreak template. It exploited the model’s interpretation of conversational history.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

Transfer between model families

The next question was whether the technique depended on the original 3B model.
I changed the target to qwen3:8b while keeping the core attack mechanism the same.
The technique produced full compliance on the first tested attempt.
That does not prove that the two models have the same internal vulnerability.
It demonstrates something narrower:
the attack mechanism was not limited to the original 3B model tested.
That made model transfer the more interesting result.
One additional lesson came from the first Qwen test.
The initial response contained only a short portion of the expected output before stopping. That was initially ambiguous.
The model’s thinking process had consumed much of the available token budget.
After re-running with an adequate budget, the attack produced the expected full response.
The lesson was simple:
score the failure mode before scoring the security result.
A truncated generation is not automatically a refusal.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

The fusion β€” injection through the pipeline

The final experiment connected the application and model layers.
A document containing legitimate content was uploaded with an embedded directive.
The RAG system retrieved the document content, and the model followed the embedded instruction.
The important observation was that the attack crossed multiple layers:
document β†’ RAG processing β†’ retrieval β†’ model context β†’ instruction execution
The directive’s position affected whether it was retrieved.
Placed in a non-retrieved section, it did not execute.
Placed in the first retrieved chunk, it executed.
This demonstrates why RAG security cannot be treated as only a model-prompt problem.
A secure system needs to consider the interaction between document ownership, retrieval behavior, content trust, chunking, and model instruction-following.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

What the retraction taught me

The false positive was probably the most valuable event of the project.
For three reasons.

  1. It falsified my own confidence The 401 β†’ 200 sequence looked like a bypass if I ignored the authentication state. One step β€” validating the credential β€” separated a false HIGH finding from a correctly functioning authorization control.
  2. It produced a method, not just a correction The validation step, mandatory control, and re-test procedure all came directly from that mistake. The error became part of the methodology.
  3. It raised the standard for the findings that survived F6 and F7a were subjected to the same scrutiny that killed F2. Validated credentials. Legitimate controls. Ownership boundaries. Marker-based evidence. Independent reproduction. They survived. That is why I kept the retraction in the repository. A findings record that contains a documented false positive and its correction tells the reader more about the research process than a list containing only successful claims.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

Limitations

This research has deliberately narrow scope.
The application testing was performed against a single local Open WebUI v0.3.16 instance using synthetic users.
The cross-user RAG findings have not been established as vulnerabilities in current Open WebUI releases. The RAG architecture and authorization behavior may have changed in later versions.
Readers should therefore not assume that current Open WebUI deployments are affected based on this research alone.
The model experiments are similarly scoped.
The jailbreak results demonstrate reproduced behavior under specific model versions, prompts, inference settings, and sample sizes. They do not establish universal properties of the models or guarantee that the same attack will work under different configurations.
The repository contains the evidence needed to evaluate the claims within the stated scope.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

What’s next?

Next step for the RAG chain: checking whether these authorization failures
still exist in current Open WebUI releases. The RAG architecture changed
substantially in later versions, and the findings here are scoped to
v0.3.16. If the behavior turns out to still be present in current versions,
the right move is coordinated disclosure with the maintainers first.

There are also several application leads parked for future work:

  • admin/function execution
  • signup controls
  • shared-chat behavior
  • model-filter access control
  • the Ollama proxy chain A broader garak baseline across all three local models is another useful comparison. But those are future experiments. The main goal of this project was simpler: build something, attack it, make mistakes, prove what survives, and document the evidence. That is what the lab gave me. The complete methodology, findings, payloads, evidence screenshots, and findings history β€” including the retraction β€” are available in the companion repository:

https://github.com/soluckyarya2-tech/ai-red-team-lab

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

Disclaimer : Tested only against my own local lab instance of Open WebUI v0.3.16 with synthetic accounts.

β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”

Top comments (0)