DEV Community

Cover image for Hephaestus: Local-First, Open-Source AI Agents That Train ML Models While You Go Outside
Love Yadav
Love Yadav

Posted on

Hephaestus: Local-First, Open-Source AI Agents That Train ML Models While You Go Outside

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

Hephaestus is a local-first, autonomous coordinator for machine learning pipelines. You describe a model in plain language, for example "train a CNN to classify these images", and a team of AI agents takes it from there: they read the research, find a dataset, write the PyTorch code, run the training in a sandbox, fix their own bugs, and remember what worked for next time.

How it gets people off the screen. This isn't a hiking app, and I won't pretend it is. The "grass" here is the hours that ML engineers spend babysitting: hunting for papers, reading shape-mismatch tracebacks, re-running a script that died on a missing import, and watching a GPU in case it runs out of memory. Hephaestus is built so the human's screen time is the shortest part of the job. You write one prompt, the agents run the whole loop unattended (research, data, code, training, debugging), and the dashboard exists to be glanced at, not stared at. When a run finishes, the best scripts are saved as reusable skills, so the next run starts further along.

Who it's for: students and researchers who have a GPU and a deadline, and who would rather go outside than debug tensor dimensions. It's also for anyone who wants an agent system that keeps code, data and API keys on their own machine.

Demo

Hephaestus is local-first by design, so it runs on your own machine instead of at a public URL, which keeps your data and keys off any server you don't control. You can start it in four commands (Docker is required, plus Ollama for local models):

cp .env.example .env
docker compose up -d redis mongodb searxng
cd backend && uv sync && uv run python main.py     # API on :8000
cd frontend && pnpm install && pnpm dev            # dashboard on :3000
Enter fullscreen mode Exit fullscreen mode

Then open http://localhost:3000. The dashboard shows a live GPU memory graph, a streaming terminal of the training logs, and a multi-agent chat where each message is tagged with the agent that sent it (Research, MLOps, Orchestrator). The repo README has the full setup for both native and full-Docker modes.

Code

GitHub logo Ewan-Dkhar / hephaestus

Hephaestus is an end-to-end, local-first autonomous machine learning pipeline coordinator. It features a multi-agent system designed to manage the entire lifecycle of training complex neural networks, from data curation to PyTorch deployment, while strictly managing consumer hardware limits to prevent Out-Of-Memory (OOM) crashes.

Project Hephaestus: Quickstart

Getting Started with Docker Compose (Still working on this)

Follow these steps to spin up the entire Hephaestus stack locally using Docker Compose.

Prerequisites

  • Docker Desktop installed and running.
  • Docker Compose installed (included with Docker Desktop).
  • NVIDIA drivers and the NVIDIA Container Toolkit (if you plan to use GPU acceleration).

1. Setup Environment Variables

First, create a .env file from the example template:

cp .env.example .env        # macOS / Linux
copy .env.example .env       # Windows (cmd / PowerShell)
Enter fullscreen mode Exit fullscreen mode

Open the .env file and set the required variables (like SEARXNG_SECRET).

2. Build and Start the Containers

Run the following command from the root of the project to build the images and start the services in detached mode:

docker compose up --build -d
Enter fullscreen mode Exit fullscreen mode

This will spin up:

  • Next.js Frontend (Port: 3000)
  • Python/LangGraph Backend (Port: 8000)
  • Redis
  • MongoDB
  • SearXNG (Port: 8080)

3. Access the Application

Once the containers are…




How I Built It

The open-source AI at its core:

  • Local inference with Ollama (default) or vLLM. The default model is qwen2.5-coder:14b, an open-weight coding model running on the user's own GPU. No cloud account is needed to use Hephaestus.
  • LangGraph for the multi-agent workflow, FastAPI for the backend, and Next.js 16 / React 19 for the dashboard.
  • FastMCP (Model Context Protocol) tool servers give the agents bounded abilities: search and download Hugging Face datasets, save code, write an append-only audit log, log experiments to MongoDB, and use a Redis scratchpad.
  • SearXNG (self-hosted metasearch) and the arXiv API for research, so literature search doesn't depend on a paid search API.
  • Docker for the training sandbox.

How a request flows:

prompt -> Orchestrator -> Triage Router
   simple script -> General Coder ----------------------+
   ML pipeline  -> Research (arXiv + SearXNG, in parallel)
                -> Data Ingestion (Hugging Face Hub -> Arrow files)
                -> Skill retrieval (past successful scripts)
                -> DataPrep agent -> Architecture agent -> Integration agent
                -> Best-of-N candidates (3) -> AST pre-flight gate
                                                          |
   evict LLM from VRAM -> Docker sandbox training <-------+
        |  failure (up to 3 retries)        | success
        v                                   v
   Analyzer -> Patcher -> re-run     Parse metrics -> save skill if better
Enter fullscreen mode Exit fullscreen mode

The design decisions that matter most

  1. Split the coding job into three agents. Data loading, model architecture and the training script are written in separate prompts (DataPrep, Architecture, Integration). This avoids the "context dilution" that makes a single prompt invent neural layers inside a data loader. Each agent has to return its code in XML-tagged sections (<imports>, <model_class>, <training_loop>), so nothing leaks between parts.
  2. Check code before running it. Every generated script goes through an AST syntax gate and an import audit that finds names like nn or DataLoader and injects the missing imports. These checks take microseconds, compared with the seconds a Docker start-up costs, so trivial mistakes never reach the sandbox. For quality, the system generates three candidate scripts and scores them on syntax, import completeness and the presence of the metrics line.
  3. Heal in two steps. When a script crashes, an Analyzer agent reads about 15 lines around the error and classifies it (import error, shape mismatch, device error, out-of-memory, and so on) without writing code. A separate Patcher agent then applies the diagnosis and the fix is syntax-checked before the retry. Splitting diagnosis from repair avoids the common failure where a model "fixes" a bug by repeating it. Retries are capped at 3.
  4. Remember what worked. Scripts print a standard metrics line (HEPHAESTUS_METRICS::{...}). If a run beats the previous baseline by at least 0.02, the script is saved to MongoDB as a learned skill, and later runs retrieve matching skills as reference code.
  5. Treat the GPU as a shared resource. A hardware manager reads VRAM and temperature via pynvml (falling back to simulated numbers on machines without a GPU) and tells Ollama to unload the language model (keep_alive: 0) before training starts, so the LLM and PyTorch never fight over memory.
  6. Sandbox everything. Generated code runs in a Docker container with an 8 GB memory cap, 4 CPU cores, no network access, user data mounted read-only, and a watchdog that kills any run whose output stays silent for 300 seconds.

User API keys, for the optional cloud providers, are encrypted at rest with Fernet (AES-128 plus HMAC) and support key rotation through MultiFernet. The backend has 22 pytest modules covering the guardrails, sandbox, encryption, tools and end-to-end pipeline.

Why Does Open Innovation Matter?

Hephaestus works because the pieces underneath it are open. Here is what that made possible that a closed API wouldn't.

  • Control over the model's memory. The GPU-eviction trick depends on owning the inference engine: Hephaestus can tell Ollama to unload the model to free VRAM for training, then load it back for debugging. A closed, hosted API gives you no handle on that. This is the clearest example in the project of open infrastructure enabling a feature, not just replacing a cost.
  • Your code and data stay on your machine. Research topics, datasets and generated training code can contain unpublished work. With a local model, none of it is sent to a third party.
  • Best-of-N without a bill. Sampling several candidate scripts, plus an analyzer and a patcher on every failure, can mean dozens of model calls per run. Locally that costs electricity, not tokens. Hephaestus even generates candidates one after another on Ollama to protect VRAM, a trade-off you can only make when you control the hardware.
  • Swap models and engines freely. Every agent talks to a single client interface, so moving from Ollama to vLLM, or trying a different open-weight model, is a configuration change. The same interface also supports Groq, OpenAI, Anthropic and Gemini as optional extras for people without a GPU, but none are required.
  • Auditable by design. Because the tools, sandbox and agent graph are open code, you can read exactly what the agents are allowed to do, and the append-only audit log records what they did.

The honest trade-off: cloud models are often stronger, and the 14B default needs a reasonably capable GPU. The open approach asks for hardware and gives back control, privacy and zero marginal cost.

## Prize Categories

Best Use of Render ($200 USD) : The Express backend (RAG ingestion, retrieval, generation, and the hourly commit job) and the managed PostgreSQL database both run on Render.

Top comments (0)