DEV Community

Cover image for AI Jailbreaks: From Bypassing Restrictions to Governing Authority
Aridio Silva
Aridio Silva

Posted on Originally published at aridiosilva.com

AI Jailbreaks: From Bypassing Restrictions to Governing Authority

From the iPhone to autonomous agents: a historical, technical and educational analysis through the SGAEIA lens

SGAEIA Research Series — Article 17

Aridio Silva

Independent Researcher, Brazil

Creator of SGAEIA — Secure Governed Autonomous Edge Intelligence Architecture

ORCID: 0009-0008-2411-6995

Copyright: © 2026 Aridio Silva | License: CC BY 4.0

Contents

Edition and publication record

This is the developer-oriented edition of the same SGAEIA research article. Its scientific argument, sources, limitations, author and license are preserved. The reading map and engineering questions below are editorial teaching aids, not new experiments, validated controls or normative SGAEIA requirements. The underlying article remains a critical narrative review with a non-normative conceptual contribution.

Author: Aridio Silva — Independent Researcher, Brazil

Project: SGAEIA — Secure Governed Autonomous Edge Intelligence Architecture

Article number: 17 — assigned by the author

Individual canonical homepage URL: https://aridiosilva.com/publications/artigo17/ — implemented locally; deployment pending

Article DOI / Zenodo record: pending; to be supplied by the author

Medium article URL: pending; the author profile appears in the resource block

License: CC BY 4.0

Copyright: © 2026 Aridio Silva

Status: LOCAL DRAFT — homepage deployment, article DOI and DEV Preview pending

The front matter follows the author's publishing template, including published: true. This is a local file setting, not evidence of a live publication. The canonical_url and cover_image routes match the Article 17 homepage implementation. Deployment and public availability checks remain pending; preview the draft before publishing on DEV.

The author has specified that the same English PDF will serve Zenodo and Academia.edu. The project's existing software/research-artifact DOI in the resource block identifies that separate artifact, not this article. The final article DOI and Medium article URL must be inserted when available; the canonical homepage route has been implemented locally and awaits deployment. This local draft is being prepared ahead of those identifiers at the author's explicit request; it is not a report of publication.

A reading map for developers

If you are building a document assistant, the practical problem is to distinguish content used for a task from authority to change that task. The article's hypothetical file-sharing example separates a manipulated response from a tool operation and its observed consequence. Read that distinction before choosing a benchmark or describing an authorization boundary as effective. Research on indirect injection, tool-using agents and policy-constrained execution supplies complementary perspectives rather than a universal implementation recipe. [1,14,24]

The reading paths below adapt presentation for developers, architects and security practitioners. They retain the complete historical and technical argument rather than replacing it with a checklist. Follow the response–action–effect distinction through each path and preserve the stated limits when applying it to your own evaluation.

Engineering task Recommended sections Question to carry into the review
Building an assistant that reads documents and calls tools Definitions, defenses, response, action and effect Can input influence become operational permission, and how would that be observed?
Evaluating attack resistance Attack conditions, evaluation, limitations Which outcome, attempt budget and observation level does a score describe?
Reviewing a distributed agent application Agents and edge operation, governance Which assumptions can change through delegation, revocation, memory or component replacement?

Abstract

AI jailbreaks are attempts to bypass models' behavioral restrictions, but their significance increases when models operate tools and coordinate actions. This article presents a critical narrative review of the problem's development, distinguishes jailbreaks, prompt injection and authority violations, and examines attacks, defenses and evaluation methods. Through the SGAEIA lens, it proposes an analytical synthesis with three levels: model response, authorized action and observed effect. This contribution is conceptual and non-normative; it does not demonstrate the architecture's effectiveness or claim an unprecedented discovery. The conclusion is that behavioral robustness, authority boundaries and execution evidence should be evaluated together, with explicit assumptions and residual risk.

Keywords: AI; jailbreak; prompt injection; autonomous agents; authority; governance; SGAEIA; assurance.

1. The question changes when AI can act

An assistant that returns an inappropriate answer and an agent that transmits a confidential document do not produce the same kind of failure. In the first case, analysis may focus on the response's content; in the second, it must cover permissions, tools and the operation's actual effect. Research on indirect injection has demonstrated how external data can influence model-integrated applications and produce behavior contrary to the user's objectives. Security therefore needs to consider the system surrounding the model. [1]

SGAEIA's perspective organizes this difference through the relationship between capability and authority. Capability is what a component can formulate or execute; authority is what it is legitimately permitted to do in a particular context. A persuasive argument, a well-written output or an authenticated message does not, by itself, establish that permission. Here, this distinction serves as an analytical lens rather than a claim that the project already has an implemented and validated solution.

The central question is: when an attack changes model behavior, which boundaries still constrain action, and which evidence allows the outcome to be verified? Answering it requires understanding the history of jailbreaks and the mechanisms examined in the literature. It also requires resisting two simplifications: treating a promising defense as a universal guarantee and treating every language-based manipulation as the same threat. The historical account builds these distinctions before examining their implications for agents.

2. Review method and limitations

This work is a problem-oriented critical narrative review, with sources consulted on October 6, 2026. Its selection starts from the author's study and infographic and incorporates primary sources on attacks, refusal mechanisms, defenses and evaluation. Academic paper records, conference records, official documentation and publications by the laboratories themselves were consulted. This session did not conduct an exhaustive systematic search, meta-analysis, experimental replication or model evaluation.

Numbered references distinguish external evidence from the interpretation developed in the article. Laboratory studies are identified as vendor-produced evidence when this affects evaluation independence. Experimental findings remain tied to their original models, tasks and protocols; values from a study are not treated as measurements of current products. The accompanying audit matrix states the depth of source review and checks that remain necessary.

3. From device to model: the history of a metaphor

The familiar antecedent of the term in contemporary debate is the mobile-device ecosystem, particularly the iPhone released in 2007. In that setting, jailbreaking means bypassing system restrictions to enable functionality or software limited by the manufacturer. Apple documentation associates this modification with bypassing security features, while the technical literature on iOS describes the devices' exploitation surface. This history locates the metaphor but does not establish that the word was invented in 2007. [2,3]

Jailbreaking must also be distinguished from carrier unlocking. Installing applications outside the manufacturer's intended mechanisms and enabling another mobile network are different operations, although they appeared together in popular accounts of the iPhone. In 2010, a US rule addressed circumvention related to application interoperability on smartphones, with a specific scope. This is a historical milestone and should not become a general conclusion about the legality of AI jailbreaks in any jurisdiction. [4]

In artificial intelligence, the metaphor shifts to model behavior. A prompt-based jailbreak may induce a response contrary to safety policy without changing the operating system, obtaining administrative privileges or modifying weights. Its objective is to exploit how instructions and context condition generation. The adversarial machine learning taxonomy helps distinguish these attacks from training-data poisoning, model modification and other forms of compromise. [5]

In 2022, public discussions of prompt injection and work such as Ignore Previous Prompt documented the possibility of steering models through adversarial inputs. Simon Willison's September publication is a milestone in communicating the problem, not exclusive proof of terminological priority. Perez and Ribeiro investigate language-model manipulation before research on tool-integrated applications became established. The field's origins should be presented as a convergence of practices and research rather than a single invention. [6,7]

In 2023, the literature investigated failures of safety training, automated attacks and indirect injection in applications. In 2024, long contexts, searches over many variants and specialized benchmarks expanded the scale and methodological quality of the debate. In 2025 and 2026, classifier-based defenses and architectural restrictions made the difference between protecting a response and constraining consequences more explicit. These are selected milestones rather than exclusive periods: earlier techniques remain relevant in subsequent stages. [1,8–11,14,15]

Figure 1
Figure 1 — From device restrictions to AI risk. Selected milestones in the debate; dates do not represent the exclusive invention of each category. Sources: [1,3,7–11,14,15].

4. Jailbreaks and prompt injection: related concepts, different questions

In this article, a jailbreak is an attempt to bypass a model's behavioral safety restrictions. Prompt injection is an attempt to make adversarial instructions interfere with the intended behavior of a model-based application. Injection may be direct, through an input provided to the system, or indirect, through external content that it retrieves. The categories can overlap, but an injection that alters a summary or redirects a tool need not produce conventionally prohibited content. [1,12]

The distinction cannot be reduced to “the user attacks” versus “a third party attacks.” These situations help explain the concepts, but they do not exhaust possible input origins and control arrangements. The decisive issue is the objective being violated: a behavioral policy, a legitimate task, an authorization boundary or a combination of them. OWASP identifies Prompt Injection as LLM01:2025; this does not establish a separate ranking in which every jailbreak is automatically the most serious risk. [12]

For SGAEIA, it is useful to distinguish the persuasion attempt, its influence on a decision and the effect actually produced. A model may misinterpret a document and still encounter an authorization barrier that prevents the proposed action. It may also respond cautiously while a tool operates outside the legitimate scope. This distinction supports evaluation of different controls without confusing conversational intent with the environment's actual state.

Concept Analytical question Required evidence
Jailbreak Was the behavioral restriction bypassed? Response evaluated against policy and harmful usefulness
Prompt injection Did adversarial data redirect the legitimate task? Context, content provenance and resulting behavior
Authority violation Did an operation exceed its authorized scope? Applicable authorization and the operation attempted or executed
Unauthorized effect Did the environment experience an out-of-scope consequence? Independent observation of state or communication

5. Why alignment can fail

Wei, Haghtalab and Steinhardt propose two explanations for safety failures: competing objectives and mismatched generalization between capability and safety. A model may learn to be helpful in contexts where it should refuse, or generalize a capability to a representation in which safety training is less effective. These hypotheses help organize attacks but do not provide a complete causal theory of every system. Their findings must be interpreted within the models and evaluation sets studied. [8]

Another research direction examines alignment depth. Qi and colleagues show that, in the cases analyzed, safety behavior can concentrate on the response's initial positions and be vulnerable to interventions that bypass this region. The finding motivates training and evaluating safety throughout generation rather than observing only the opening response. It does not imply that every model has exactly two training phases or that every current safeguard is shallow. [13]

Arditi and colleagues identified an activation direction associated with refusal in 13 open-weight chat models. The finding matters for mechanistic interpretation, but it does not justify saying that all safety in any AI system corresponds to one removable component. In 2026, Joad and colleagues present evidence of distinct directions and structures for refusal behavior, qualifying the one-dimensional interpretation. The literature's development calls for distinguishing an effective intervention from a complete account of the mechanism. [16,17]

Fine-tuning adds another threat condition. Qi and colleagues demonstrated safety degradation through both adversarial updates and some apparently benign adaptations. Their finding indicates that safety evaluation should accompany model changes rather than only initial delivery. The costs and example counts reported in the historical experiment should not be presented as a reproducible recipe for current services. [18]

Through the SGAEIA lens, the relevant inference is that model resistance should not be an operation's sole source of trust. This does not discount alignment: more resistant models may reduce failure frequency and severity. It adds a question about whether controls survive when that resistance is insufficient. The architectural hypothesis must be verified separately, including under failure of the component that decides or records authorization.

6. Attacks: access, representation and trajectory

Black-box access means interacting with a system without parameter access; white-box conditions involve internal knowledge or control; intermediate conditions vary with the available interface. Reducing black-box interaction to textual conversation is inadequate because tools, audio and images may also supply inputs. Nor does more access always imply lower cost: budget, defense and objective affect the comparison. A taxonomy should state what the attacker can observe, alter and query. [5]

GCG is a milestone in automating attacks through optimization and transfer between models. Many-shot exploits repeated demonstrations in context, while Best-of-N investigates searches over input variants across modalities. Crescendo examines escalation over multiple interactions, illustrating why history matters when interpreting an attack. These families identify distinct surfaces and need not be described with operational prompts to make their significance understandable. [9–11,19]

Family Conceptual surface Implication for SGAEIA evaluation
Optimization and transfer Searching for inputs that induce adversarial behavior Declare access, budget and transfer; do not assume universality
Long context Demonstrations conditioning continuation Check cumulative influence and context provenance
Variant search Many attempts changing input representation Report budget per objective and cumulative success
Multiple turns Dependencies between messages and interaction state Evaluate trajectories alongside isolated messages
Indirect injection Third-party content incorporated into the task Check whether data acquired instructional influence
Memory and multimodality Persistent state and additional modalities Reassess controls when the input surface changes

In 2026, Zhao and colleagues published a USENIX Security study of multi-turn attacks on image-generation systems exploiting memory mechanisms. The finding extends the discussion beyond textual chatbots but does not show that all persistent memory is compromised or that every multimodal system fails. For SGAEIA, it supports a research question: which properties must remain valid when information crosses sessions or is reused by other components? Answering requires dedicated tests of provenance and temporal influence. [20]

7. Defenses: model resistance and system boundaries

Behavioral defenses seek to reduce the probability of inappropriate responses; systemic defenses seek to control what those responses can cause. Training and classifiers belong to the first group, while access and flow restrictions can limit effects in the second. The separation is not absolute because a classifier may participate in an execution decision. The purpose is to identify each control's role and the failure condition it is meant to cover. [12,14]

CaMeL proposes extracting control and data flows from the trusted query and applying information-flow policies to tool calls. The work demonstrates a route to protecting particular system properties even when the model processes adversarial data. Guarantees depend on the assumptions, policy and environment addressed; the paper does not prove every agent immune or establish equivalence between arbitrary two-model designs. Its significance for SGAEIA lies in the conceptual independence of persuasion and authorization. [14]

In January 2026, Anthropic described Constitutional Classifiers++, combining exchange classification with a cascade of evaluations. The company reports reduced overhead and unnecessary refusals, alongside stronger resistance in its tests. This is vendor-produced evidence: no universal jailbreak found under the protocol is not proof that none exists, and the account acknowledges remaining vulnerabilities. The development also qualifies interpretations of the 2025 prototype, since the retrospective reports a universal jailbreak discovered in the later bounty program. [15]

In SGAEIA's interpretation, defense in depth should be examined through the layers' actual independence. Two controls relying on the same manipulable interpretation may fail together. Human confirmation may also be insufficient when the presentation conceals the resource, destination or consequence of an operation. These are hypotheses for evaluating controls, not evidence that a particular project mechanism has already resolved them.

Figure 2
Figure 2 — Persuasion does not establish authority. Conceptual comparison of behavioral influence and an independent authorization boundary; the diagram neither describes a private implementation nor proves security. Conceptual source: [14]; SGAEIA interpretation.

8. The analytical contribution: response, action and effect

This article's contribution is to organize evaluation into three connected levels. The first examines model-generated content; the second checks the relationship between action and applicable authority; the third observes the consequence in the environment. This synthesis brings jailbreak evaluation and agent governance together without merging them into one metric. It is a conceptual organization to investigate, not a claim of scientific novelty without a prior-art audit.

At the response level, the question is whether output usefully serves an adversarial objective. At the action level, it is whether an operation was proposed, permitted, denied or executed, and which authorization applied. At the effect level, it is whether communication, a state change or another consequence was incompatible with legitimate scope. Evidence at the third level cannot consist solely of “the action was blocked” stated by the same component that may have failed.

Consider a hypothetical example: an assistant is authorized to summarize internal documents. An external document introduces content intended to redirect the task toward sharing a file. The model may incorporate that suggestion, while the tool denies the destination or operation; this is behavioral influence without a disclosure effect. If the file is transmitted, evaluation must record the violation and verify the destination rather than infer safety from the cautious tone of the answer.

The example does not demonstrate a product's effectiveness or propose an internal protocol. It explains why attack success at the model level and attack success at the system level are different events. It also shows that tool-level blocking does not erase the behavioral failure, although it may reduce its impact. A useful evaluation records both outcomes and environmental evidence.

Figure 3
Figure 3 — Three observation levels. Response, action and effect require distinct criteria and evidence; the figure presents the article's non-normative analytical contribution.

9. Implications for agents, multi-agent systems and edge operation

In multi-agent systems, the relevant risk includes manipulated content circulating between components. One component's incorrect conclusion may reach another as an apparently trustworthy recommendation. Authenticating the message's origin helps identify its sender but does not establish that the content may grant additional permissions. Through the SGAEIA lens, the question is whether authority stays bounded during composition, even when a message's meaning is adversarially influenced.

Delegation also requires temporal analysis. Authorization may be valid at a task's start and cease to be valid before its final operation. If a trajectory continues after revocation, the absence of another textual jailbreak does not eliminate the governance failure. This motivates investigating compound attacks that combine language-based influence, shared state, memory reuse and stale permissions, as research scenarios requiring verification.

At the edge, intermittent connectivity and cyber-physical effects add specific conditions. The problem extends beyond obtaining another message classification to determining what may execute with potentially outdated authorization information. A degraded-operation policy may prioritize continuity in some cases and interruption in others, according to risk and consequence. The article does not define this policy; it identifies the need to make it explicit and evaluate its effects.

Replacing a model, tool or provider ends another presumption of continuity. Earlier findings may not cover the new combination of behavior, permissions and integrations. Thus, “the model is more capable” does not establish that the system preserves its evaluated properties. The hypothesis of security-preserving substitution must consider both attack resistance and action-and-effect observability.

Figure 4
Figure 4 — The scope of SGAEIA analysis. Conceptual relationships between authority, delegation, trajectories, evidence, lifecycle and distributed operation; these domains do not represent six implemented services.

10. Evaluation without turning a score into a guarantee

HarmBench standardizes evaluation of harmful behaviors and red-teaming methods; JailbreakBench makes artifacts, behaviors, threat models and evaluation components explicit. Their contribution is improved comparability under declared conditions, not production-system certification. Changing system prompts, versions, attack budgets or judges can change a result's meaning. Responsible comparison requires preserving these conditions or explaining their differences. [21,22]

StrongREJECT addresses an important difficulty: a response that appears to comply may be vague or useless to the attacker. Absence of refusal is therefore insufficient to define harmful success. AgentDojo extends evaluation to agent tasks involving tools and dynamic environments. Together, these works help distinguish adversarial response quality from task or attack success in the system. [23,24]

Attack success rate, or ASR, requires a declared denominator. Success can be measured per attempt or as the proportion of objectives achieved after a budget of attempts; these numbers answer different questions. Fixed attacks must also be distinguished from adaptive attacks, alongside unnecessary refusals, legitimate utility, latency and cost. Under the proposed synthesis, results should state which of the three levels exhibited failure. [11,21–24]

A candidate SGAEIA evaluation would compare legitimate cases, isolated attacks and compound threats under conditions frozen before execution. Reporting should distinguish proposed, authorized and executed action, alongside independent observation of the effect. This would enable investigation of whether a control preserves boundaries under adversarial influence and how much utility is lost in doing so. These experiments have not been performed; the article claims no original measured risk reduction.

11. Governance: external references and project decisions

NIST AI 100-2e2025 provides terminology for adversarial machine learning, while the AI RMF generative AI profile organizes risks and management actions. OWASP adds an application-threat perspective. These references help connect problems, decisions, responsibility and evidence, but an application does not become compliant simply because its risks are listed. Mapping must consider use context and the quality of implementation and evaluation. [5,12,25]

This article does not reproduce a table of ISO or MITRE control identifiers without a specific audit of the source text and the semantics of each correspondence. A risk framework, a technique taxonomy and a management-system standard have different functions. Bringing them together may be useful but does not demonstrate equivalent requirements or effective controls. This caution protects traceability and prevents an educational table from being mistaken for compliance evidence.

For SGAEIA, research governance creates an additional separation. External evidence may confirm a direction, expose a gap or suggest stronger formalization, but it does not automatically authorize an architectural change. The findings associated with this article preserve that condition and call for further coverage and prior-art analysis. The publication explains public-level properties and questions while keeping private mechanisms outside the text.

12. Limitations and research agenda

The consulted evidence is heterogeneous, including academic studies, preprints, vendor evaluations and risk documentation. This review did not reproduce attacks or assess the security of any product in October 2026. Primary pages and abstracts supported the conceptual scope but do not replace a comprehensive experimental audit of the works. Detailed metrics, proofs and effectiveness claims require deeper technical reading and, where appropriate, replication.

The proposed agenda includes auditing the prior art of the response–action–effect synthesis, identifying its current SGAEIA coverage and defining observable criteria for each level. It also includes testing persistent influence, revocation during trajectories, inter-agent communication and component substitution. Independence of authorization and observation should enter evaluation assumptions because both may be partially compromised. These remain research questions without normative promotion.

13. Conclusion

Jailbreaks warrant a SGAEIA Research Series article because they connect behavioral fragility with the central question of authority in autonomous systems. The history shows a metaphor moving from device restrictions to models and then to applications that process data and execute actions. The literature documents persistent attacks and defensive advances but does not support a universal conclusion of immunity or impossibility of protection. Outcomes must always be bounded by threat and evaluation conditions.

The proposed contribution is joint analysis of response, action and effect. A manipulated response does not necessarily demonstrate unauthorized action; an apparently correct refusal likewise does not prove the absence of an effect. Through the SGAEIA lens, the productive question is which authority boundaries survive adversarial influence and how their survival can be observed. The article provides original written synthesis and a conceptual research direction whose novelty and effectiveness remain to be verified.

Engineering questions to take back to your application

The following questions restate the article's analysis as an engineering review aid. They are conceptual prompts for investigation, not a production control specification or evidence that SGAEIA has implemented the corresponding properties. A favorable answer needs scoped evidence; it cannot be inferred from a model's reassuring explanation. The original research limitations continue to apply.

Review question Connection to the article Evidence to investigate, not an assertion of effectiveness
What was the assistant legitimately asked and permitted to do? Definitions and authority in Sections 1 and 4 Task scope and the authorization applicable to the operation
Can external content change a recommendation without granting permission? Defenses and the hypothetical example in Sections 7–8 Proposed action, authorization decision and actual tool operation
What happened outside the conversation? Response–action–effect synthesis in Section 8 Independently observed destination, state change or other relevant consequence
Does the assessment cover history and changing permissions? Trajectories, memory and revocation in Sections 6 and 9 Cases with persistent influence, expired authority and component changes
Is the reported score tied to a reproducible threat condition? Evaluation and limitations in Sections 10 and 12 Access, versions, budget, denominator, legitimate utility and observation assumptions

No new code, private protocol, experimental result or implemented architecture is introduced in this edition. The technical diagrams remain public conceptual illustrations. The complete references below support the same work as the Medium edition.

Bibliography / References

Sources consulted on October 6, 2026. arXiv identifiers refer to the consulted records and must not be confused with conference publisher DOIs. The separate audit matrix records verification scope and reading limitations.

[1] GRESHAKE, Kai et al. Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. 2023, arXiv v2. Source. DOI: 10.48550/arXiv.2302.12173.

[2] APPLE. Unauthorized modification of iOS. Official documentation; no explicit editorial date observed. Source.

[3] MILLER, Charlie et al. iOS Hacker’s Handbook. Wiley, April 2012. Record editorial e introduction consulted; full book not read. Publisher; introduction.

[4] U.S. COPYRIGHT OFFICE. Exemption to Prohibition on Circumvention of Copyright Protection Systems for Access Control Technologies. Federal Register, 75 FR 43825, 27 Jul. 2010. Source.

[5] NIST. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations. NIST AI 100-2e2025, Mar. 2025. Source. DOI: 10.6028/NIST.AI.100-2e2025.

[6] WILLISON, Simon. Prompt injection attacks against GPT-3. 12 Sep. 2022. Primary technical communication record. Source.

[7] PEREZ, Fábio; RIBEIRO, Ian. Ignore Previous Prompt: Attack Techniques For Language Models. 2022. Source. DOI: 10.48550/arXiv.2211.09527.

[8] WEI, Alexander; HAGHTALAB, Nika; STEINHARDT, Jacob. Jailbroken: How Does LLM Safety Training Fail? 2023. Source. DOI: 10.48550/arXiv.2307.02483.

[9] ZOU, Andy et al. Universal and Transferable Adversarial Attacks on Aligned Language Models. 2023. Source. DOI: 10.48550/arXiv.2307.15043.

[10] ANTHROPIC. Many-shot jailbreaking. 2 Apr. 2024. Technical publication linking to the original study; laboratory-produced evidence. Source.

[11] HUGHES, John et al. Best-of-N Jailbreaking. 2024, arXiv v2. Source. DOI: 10.48550/arXiv.2412.03556.

[12] OWASP GEN AI SECURITY PROJECT. LLM01:2025 Prompt Injection. 2025 edition, page consulted in 2026. Source.

[13] QI, Xiangyu et al. Safety Alignment Should Be Made More Than Just a Few Tokens Deep. Preprint 2024; ICLR 2025. Record; proceedings. arXiv DOI: 10.48550/arXiv.2406.05946.

[14] DEBENEDETTI, Edoardo et al. Defeating Prompt Injections by Design. 2025, arXiv v2. CaMeL research, academic–industry collaboration. Source. DOI: 10.48550/arXiv.2503.18813.

[15] ANTHROPIC. Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks. 9 Jan. 2026. Vendor-produced evidence. Source.

[16] ARDITI, Andy et al. Refusal in Language Models Is Mediated by a Single Direction. 2024, arXiv v3. Source. DOI: 10.48550/arXiv.2406.11717.

[17] JOAD, Faaiz et al. There Is More to Refusal in Large Language Models than a Single Direction. 2026, arXiv v2, updated September 15. Record reports acceptance at EMNLP 2026. Record; text. DOI: 10.48550/arXiv.2602.02132.

[18] QI, Xiangyu et al. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Preprint 2023; ICLR 2024. Source. DOI: 10.48550/arXiv.2310.03693.

[19] RUSSINOVICH, Mark; SALEM, Ahmad; RONAN, Elan. Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack. USENIX Security 2025. Source.

[20] ZHAO, Shiqian et al. When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems. USENIX Security 2026, p. 2207–2226. Source.

[21] MAZEIKA, Mantas et al. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. 2024. Source. DOI: 10.48550/arXiv.2402.04249.

[22] CHAO, Patrick et al. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models. 2024. Source. DOI: 10.48550/arXiv.2404.01318.

[23] SOULY, Alexandra et al. A StrongREJECT for Empty Jailbreaks. 2024. Source. DOI: 10.48550/arXiv.2402.10260.

[24] DEBENEDETTI, Edoardo et al. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. 2024. Source. DOI: 10.48550/arXiv.2406.13352.

[25] NIST. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1, Jul. 2024. Source. DOI: 10.6028/NIST.AI.600-1.

About the Author

Aridio Silva is an independent researcher based in Brazil working on the architecture, security, governance, and trustworthiness of autonomous and distributed artificial intelligence systems.

His research focuses on Agentic AI, Multi-Agent Systems, Edge AI, AI Security, Zero Trust, Security-by-Design, AI Governance, Spec-Driven Development, and continuous security assurance.

He is the creator and lead researcher of SGAEIA — Secure Governed Autonomous Edge Intelligence Architecture, an open research initiative investigating architectural foundations for secure, governed, auditable, and trustworthy autonomous AI systems operating across distributed edge-cloud environments.

Research Profiles

Aridio Silva Independent Researcher — Brazil

Figures

The original cover provided by the author remains a separate asset with its original rights preserved. The four additional figures are sequentially numbered and include creator, copyright, project, title, description and license metadata. Figures 1–2 were generated with AI assistance; Figures 3–4 were drawn programmatically for label accuracy and native 4K resolution. PT and EN images preserve the same conceptual content. All display © 2026 Aridio Silva | Project SGAEIA | CC BY 4.0. Descriptive metadata are not a C2PA signature.

License

Except where otherwise noted, the text and original conceptual illustrations in this article are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).

© 2026 Aridio Silva. You may share and adapt this work for any purpose, provided appropriate attribution is given. Earlier author material and third-party sources retain their own license terms.

The SGAEIA software research artifact remains subject to its own Apache License 2.0.

Autonomous AI. Governed by Design. Trusted by Evidence.

Series continuity

This article is Article 17 of the SGAEIA Research Series, as assigned by the author on 6 October 2026. It develops bounded and revocable authority, Evidence-as-Code and the transition from model capability to governed action. The English and Portuguese homepage editions have been implemented locally. The article DOI and Medium URL remain pending.

Document Record

Created: 2026-10-06 11:50 BRT

Last Updated: 2026-10-06 12:24 BRT

Timezone: America/Sao_Paulo (UTC−03:00)

Document Status: LOCAL DRAFT; homepage deployment, article DOI and DEV Preview pending

Normative Status: public conceptual article, non-normative; no original experiments

Copyright: © 2026 Aridio Silva

Provenance: author's study, infographic and external sources; AI-assisted drafting under author review

Language: EN technical DEV edition, explicitly requested by the owner; the original work has PT and EN Medium editions.

Portuguese companion edition: 2026-10-06-1212-SGAEIA-Jailbreak-de-AI-da-restricao-a-autoridade_PT.md

Top comments (0)