DEV Community

Mustafa Salih Berk
Mustafa Salih Berk

Posted on Originally published at meetcyber.net on

Building RedPatch: How I Built an AI-Powered AppSec Playground with FastAPI & Docker

RedPatch Banner (Created via ChatGPT)

Introduction

Most CTF (Capture The Flag) platforms heavily prioritize offensive mechanics-finding vulnerabilities and extracting flags. But does this model actually help developers understand the root cause of security bugs? More importantly, how effectively does it bridge the gap between offensive exploitation and secure coding practices?

Driven by these questions, I set out to build RedPatch : an open-source, hybrid application security (AppSec) playground engineered for both developers and security researchers. RedPatch adopts a dual approach combining Red Team (offensive) and Blue Team (defensive) workflows. Rather than focusing solely on exploitation, it forces you to analyze the underlying source code and write secure patches.

By integrating Large Language Models (LLMs), RedPatch automatically evaluates whether a submitted patch successfully mitigates the security flaw, serving real-time, actionable feedback directly to the user.

System Architecture

Architecture Schema
Architecture Schema

Project Structure

To maintain a clean, modular architecture and encourage community contributions, the core orchestration engine is strictly decoupled from individual lab environments via isolated Docker containers across separate repositories.

redpatch/
├── app/
│ ├── main.py
│ │
│ ├── core/
│ │ ├── config.py
│ │ └── config.json
│ │
│ ├── labs/
│ │ └── manifest.json
│ │
│ ├── services/
│ │ ├── ai/
│ │ ├── container_services/
│ │ └── module_manager/
│ │
│ ├── static/
│ └── templates/
│
├── docker-compose.yaml
├── requirements.txt
├── CONTRIBUTING.md
├── SECURITY.md
└── LICENSE
Enter fullscreen mode Exit fullscreen mode

Grounded in this decoupled design, external developers can seamlessly register custom lab environments using the project’s manifest.json specification:

{
  "labs": {
    "SQLi": {
      "description": "SQL Injection",
      "submodules": [
        {
          "id": "sqli-0",
          "title": "SQL Injection - Authentication Bypass",
          "category": "web",
          "image_tag": "redpatch-lab/sqli-0:v1.0.0",
          "port": 5000,
          "dev_path": "./labs/sqli-0",
          "download_url": "....tar.gz" 
        }
      ]
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Tech Stack Breakdown

Component Technology Technical Role
Backend Engine FastAPI (Python) High-performance asynchronous API, routing, and core orchestration
Lab Isolation Docker Desktop / Engine Isolated containerized execution environments for vulnerable targets
AI Auditor Google Gemini API Automated security auditing and patch validation engine
Frontend UI Jinja2 / HTML5 / Tailwind UI layout, integrated Monaco code editor, and interactive terminal interface

To prevent tight coupling with a single AI ecosystem, the LLM provider engine is abstractly interfaced through a RedTeamAgent class. The active provider instance is dynamically returned at runtime based on the config.json parameters:

class RedTeamAgent:
    def __init__ (self):
        self.provider: BaseLLMProvider = self._get_provider()

    def _get_provider(self) -> BaseLLMProvider:
        provider_name = settings.LLM_PROVIDER.lower()

        if provider_name == "gemini":
            return GeminiProvider()
        else:
            raise ValueError(f"Unsupported or undefined LLM provider: {provider_name}")

    async def run_attack(self, code: str, vulnerability_type: str, routes: list, lab_link: str) -> VulnerabilityAnalysis:
        return await self.provider.analyze_code(code, vulnerability_type, routes, lab_link)
Enter fullscreen mode Exit fullscreen mode

Technical Details & Core Mechanics

1. Containerized Workspace Isolation

Every lab environment spawns as an independent Docker container separate from the primary engine. Real-time code modifications applied within the embedded editor are dynamically mounted into a temporary workspace volume inside the target container.

2. Dual-Mode Workflow

  • Pentester Mode: Uncover vulnerabilities, construct payloads, execute exploits, and retrieve flags.
  • Coder Mode: Inspect raw source code within the embedded Monaco Editor, diagnose structural flaws, and refactor code to enforce secure coding practices.

3. Automated Patch Verification Engine

At runtime, RedPatch inspects the target application’s route structures and feeds the entire modified source code  — along with the active target URL and vulnerability classification — to the Gemini API. By leveraging Gemini’s native response_schema feature, the engine guarantees strictly typed structural output adhering to a JSON specification:

{
"type": "object",
"properties": {
    "vulnerability_found": {"type": "boolean"},
    "target_line": {"type": "integer"},
    "explanation": {"type": "string"},
    "exploit_request": {
        "type": "object",
        "properties": {
            "path": {"type": "string", "description": "The HTTP endpoint path, e.g., /login-vulnerable"},
            "method": {"type": "string", "description": "HTTP Method in uppercase: POST, GET, PUT, DELETE"},
            "headers": {"type": "string", "description": "JSON string of headers or empty string"},
            "params": {"type": "string", "description": "URL query string or empty string"},
            "data": {"type": "string",
                     "description": "Form payload string, e.g. username=admin&password=123, or empty string"},
            "json_body": {"type": "string", "description": "JSON body string or empty string"}
        },
        "required": ["path", "method"]
    }
},
"required": ["vulnerability_found", "target_line", "explanation", "exploit_request"]
}
Enter fullscreen mode Exit fullscreen mode

This structural output allows users to fire the AI-generated attack vector against their modified target with a single click in the UI. If the patch successfully mitigates the attack vector, the AI acknowledges the remediation and flags the module as solved.

Engineering Challenges & Lessons Learned

Architecting RedPatch involved navigating several non-trivial system design hurdles:

1. Decoupled Docker Container Management

RedPatch was my first major project built with Docker, and I prioritized two primary constraints:

  • Complete process and network isolation for target lab instances.
  • Effortless environment setup by allowing the core platform itself to deploy within a container.

A primary engineering challenge was handling workspace syncing across container boundaries. Transmitting code changes over HTTP endpoints introduced unnecessary overhead and failure points. Instead, I implemented a temporary host workspace mounted directly into the target lab container. This guaranteed hot-reloading when code edits occurred in the web editor.

Additionally, to accommodate execution environments both inside and outside Docker containers, I added an ENV IS_DOCKER=true variable within the Dockerfile. A runtime helper function checks this flag to construct valid target workspace paths across environment configurations.

2. Enforcing Deterministic AI Outputs

In my initial implementation, I passed Pydantic models directly into the Gemini API’s native response_schema parameter. However, the model struggled to interpret the constraints of complex Pydantic schemas accurately—it frequently misconstrued critical fields as optional and omitted them, causing constant verification failures during downstream parsing.

To resolve this, I refactored the pipeline by feeding a simplified, explicit JSON Schema directly into response_schema at the API level. This ensured strict structural adherence directly during the model's generation stage. I then passed the resulting raw JSON payload back to Pydantic purely for the final validation pass.

Getting the model to generate accurate exploit HTTP requests was another hurdle. It required heavy prompt tuning, but dynamically injecting extracted application routes and signature parameters into the context window finally resolved payload inconsistencies.

Future Roadmap & Conclusion

Looking ahead, planned updates for RedPatch include:

  • Expanding the vulnerable lab inventory to cover broader OWASP Top 10 classifications.
  • Patching identified bugs and refining system execution pipelines.
  • Introducing multi-user capabilities and competitive dynamic scoring leaderboards.

Contributing & Supporting the Project

If you are interested in contributing to RedPatch:

  • Check out CONTRIBUTING.md, build custom lab modules, and submit a Pull Request to the labs repository.
  • Propose feature enhancements or report system bugs via GitHub Issues.
  • Support project development by dropping a star on GitHub! ⭐

GitHub Repositories:

📌 References & Community

If you want to check out my other security research, tools, or open-source projects, feel free to explore the links below:


Top comments (0)