Assign a GitLab issue to a bot, and a few minutes later there's a draft merge request with the fix. Open an MR, and the same bot reviews it before a human gets to it. That's langgraph-harness, a self-hosted agent platform for GitLab I've been building and running on real projects since June.
What it does
- MR review. A webhook fires when an MR opens or changes. The agent reads the diff and the code around it, drafts comments, screens them, and posts what survives.
- Issue → draft MR. Assign an issue to the bot and it works in its own Kata Containers VM, runs the project's build and tests, and opens a draft MR. Comment on the MR and it picks the work back up on the same branch.
- Tasks. The same sandboxed loop, started from a plain instruction instead of an issue. Run it once or on a schedule.
- Chat. A GitLab-aware assistant that can read your projects, browse, and search past reviews. Anything sensitive waits for your approval first.
It all runs on your own servers: your code doesn't go to a third-party service, and you choose the model (OpenRouter, Gemini, Groq, Ollama or sglang). There's a React admin UI for watching runs live, and an analytics page that tracks whether people actually act on what the bot says.
So far: a little over 1,000 merge requests reviewed, 67% of its comments resolved by the people it reviewed, and 133 releases.
How it's built
The agent loop was the easy part: LangGraph.js, createAgent, a model, GitLab tools and a system prompt. What took the months was making it trustworthy enough to leave running unattended. Four pieces did most of that.
Runs that crash pick up where they stopped
A review makes around 26 model calls, and somewhere in there a provider times out or the server restarts for a deploy. Every run is a LangGraph thread backed by the Postgres checkpointer, so a failed run retries from its last checkpoint instead of starting over and paying for the same tokens twice. The same checkpoints are what let Chat pause for your approval and continue once you answer.
Guardrails that catch the agent mid-run
The common failures weren't wrong answers. They were the agent calling the same tool over and over, ending a run with "Would you like me to...?" when no one is there to answer, or wrapping up without checking for new replies on the MR. LangChain v1's createMiddleware hooks into every model and tool call, so each failure became a small guard. A trimmed version of one:
const TRAILING_QUESTION = /\?[*_~\s]*$/;
export function trailingQuestionGuardMiddleware() {
return createMiddleware({
name: 'TrailingQuestionGuard',
wrapModelCall: async (request, handler) => {
const result = await handler(request);
if (result.tool_calls?.length) return result;
if (!TRAILING_QUESTION.test(messageContent(result.content).trim())) return result;
// ask once to restate it as a finished summary
return handler({
...request,
messages: [...request.messages, result, new HumanMessage(CORRECTIVE_MESSAGE)],
});
},
});
}
None of these were planned. Each one exists because a real run went wrong in a specific way. Review comments get a guard too: a separate critic pass reads every draft before it posts and can drop it. It has dropped 81 of 310 so far.
Sub-agents that check one file at a time
On a large MR, a reviewer that has read 40 files reviews the last one worse than the first. So the reviewer can hand one file to an isolated verifier that investigates on its own and returns findings only. It can't post anything or spawn more agents; the main reviewer decides what to keep. For code-changing runs, a separate check at the end classifies the run as done, partial or failed, rather than trusting the agent's own last message.
Memory that stays useful
A bot that reviews the same project every week shouldn't relearn its conventions every week. During a run, the agent flags things worth remembering: a build command it verified, a repo convention, a false positive it keeps hitting. At the end of the run, a curator pass decides what to add, update or retire. That's 422 added, 22 updated and 2 retired so far, and you can see and edit all of it in the admin UI.
Try it
You'll need Node 22+, Docker, a GitLab access token and OAuth app, and an LLM key (or a local Ollama/sglang).
npm install
cp backend/.env.example backend/.env
npm run dev
The getting started guide covers OAuth, webhooks and per-repo config.
Repo (MIT): https://github.com/vrajpal-jhala/langgraph-harness
Docs: https://vrajpal-jhala.github.io/langgraph-harness/
GitHub support is next on the roadmap. If you'd run something like this on your team, what would you want it to do first?
Top comments (0)