DEV Community

Cover image for 195 repos, one map: the meta-repo, and why your AI agents need it
Maxime Sahroui
Maxime Sahroui

Posted on

195 repos, one map: the meta-repo, and why your AI agents need it

A client asked a simple question this week: “Can the agent see the whole platform?”

The platform is roughly 195 git repositories. Frontends, services, Terraform, a few libraries nobody remembers writing. No agent sees all of that. Neither does a new developer.

So overrepo got built. A manifest, a selector, and a map.

Why I built it

Agents don't work alone anymore. They hand off, call each other, and try to connect the dots across a codebase to be efficient.

But connecting dots needs dots. Without a picture of what exists, every agent starts from zero: grepping, guessing, rediscovering the same structure on every task.

overrepo gives them a solid starting point. A written map of the fleet that any agent can read before it touches anything.

It also covers a situation I hit constantly. Maybe you build with T3 Turbo, Nx, Turborepo or pnpm workspaces. Your apps and packages are neatly organised inside one repo. Then a task comes in that crosses the boundary: the API in one repository, the infra in another, a shared SDK in a third. Suddenly you're juggling several clones by hand. That's exactly the gap a meta-repo fills.

And if you're further along, with agents running in cloud VMs, it gets very practical. You don't want a fresh VM cloning 195 repositories for a two-repo fix. With overrepo, the agent clones only the repositories the current task needs, selected by tag or path, and gets to work.

Where the meta-repo comes from

The problem is old. Android hit it first at scale: hundreds of git repositories that had to be checked out and worked on together. Google's answer was the repo tool: one manifest file listing every repository, its path and its remote, and a command that syncs them all.

Later, the meta npm package brought the same idea to everyday teams. A .meta file in a parent folder lists child repos, and one command runs across all of them.

That's the meta-repo in one sentence: keep many repositories, but give them one clone and one operations surface.

What it enables:

  • One clone. A new machine, a new teammate, a CI job: one command and the whole fleet is on disk, at the right paths.
  • One place to act. Pull, check status, run a script, grep across everything. No shell loop pasted from a wiki.
  • One description of the fleet. The manifest is the list of what exists, where it lives and what it's for.

And what it doesn't change: each repository keeps its own history, CI, permissions, owners and release cycle.

Monorepo, submodules, meta-repo

The monorepo merges everything into one repository. One history, one CI, atomic changes across projects. Great when you choose it on day one. Hard to adopt when 195 repos already exist, each with its own pipeline and owners. Migrating is a project in itself.

Git submodules keep repos separate but pin each child to a commit inside a parent repository. That's useful for vendoring a dependency at a known version. For 200 independently released services, it means a parent repo that is always out of date, and a git submodule update that becomes a ritual.

The meta-repo keeps repos separate and pins nothing. The parent knows which repositories exist and where they go, not which commit they're at. Each child moves at its own pace. The layer on top is just coordination.

That's the middle ground. It's also where tooling tends to get thin.

Why the classic tools struggle at ~200

Script-based meta tools are the right shape. But at this scale a few things need to be first-class, not bolted on:

  • Selecting a subset by intent ("all backend services"), not by listing names.
  • Surviving one broken remote without aborting the other 199 operations.
  • Never hanging on a password prompt in the middle of a batch.
  • Producing something an AI assistant can actually read.

That last one is new. Two years ago nobody asked for it. Today it's the reason I'd reach for a meta-repo at all.

What overrepo does

You describe the fleet once in overrepo.yaml: a path, a URL, some tags, an optional description per repo. Then:

overrepo init                              # build a manifest from repos already on disk
overrepo sync                              # clone what's missing
overrepo status --dirty                    # what has local changes?
overrepo exec --tags backend -- git pull   # run a command in a subset
Enter fullscreen mode Exit fullscreen mode

Selection works the same across commands: --tags, --paths, --projects, or --all. Criteria combine with AND. exec refuses to run without a selection, so you can't accidentally run something everywhere.

Nobody wants to write 195 manifest entries by hand. import reads a JSON catalogue from stdin. With the GitHub CLI and jq, repository topics become tags:

gh repo list acme --limit 1000 --json name,sshUrl,description,repositoryTopics \
  | jq '{projects: map({name: .name, url: .sshUrl, desc: .description,
      tags: ((.repositoryTopics // []) | map(.name))})}' \
  | overrepo import --sync
Enter fullscreen mode Exit fullscreen mode

Re-run it next month, and it picks up the new repositories without touching your edits.

The map: context for AI agents

This is the part that matters most to me.

An agent dropped into one repository sees one repository. Ask it to change an API, and it has no idea which of the other 194 repos call that API. Ask it "where does billing live?" and it starts grepping, burning context on files that don't matter, or worse, it guesses.

A human senior engineer doesn't work that way. They carry a mental map: what exists, roughly what each piece does, who talks to whom. Agents need that map written down.

overrepo context writes it:

  • One Markdown summary per repository, built from the manifest (name, path, remote, tags, description) and the repository itself.
  • An index.md listing the whole fleet, so an agent can scan 195 entries in one read and pick the three that matter.
  • Output next to the manifest in a folder you set with summary.outDir. Version those summaries, and point your assistant's instructions at that folder.

Two details make the map trustworthy:

  • Summaries are read from each clone's origin/HEAD. Your dirty local checkout, half-finished branch or debug hack never leaks into what the agent believes is true.
  • In CI, overrepo context --check fails when a summary is stale. --prune removes summaries of repositories that are gone. The map doesn't rot quietly.

The effect is simple: the agent reads the index first, then opens the right repositories. Less wandering, fewer wrong guesses, smaller context windows. Tags and descriptions you'd write anyway for humans become navigation for machines.

Built for the bad day

At 200 repos, something is always broken. overrepo assumes it:

  • One repository failing never stops the others. The exit code tells you something failed.
  • Git never prompts. SSH runs in batch mode.
  • Timeouts and retries with backoff on transient network errors.
  • Clones land in a ..overrepo-partial directory and only move into place when complete. No half-cloned repos pretending to be real.

overrepo doctor checks git, remote access, the manifest, missing clones, orphan repositories and interrupted clones.

Requirements: Node.js and git on your PATH. That's it.

The takeaway

  • The meta-repo is an old idea: many repos, one manifest, one clone, one place to act.
  • It's not a monorepo (no merged history) and not submodules (no pinned commits). Each repo stays independent.
  • At ~200 repos, the manifest is the product: paths, URLs, tags, descriptions in one plain YAML file.
  • Select by tag or path, not by memory. Assume failure: isolated errors, no prompts, atomic clones.
  • Give your agents a map. Generated from the default branch, versioned, checked in CI.
  • Clone only what the task needs. Crucial when your agents spin up in fresh cloud VMs.

Still in beta, and that's where you come in

npm install -g overrepo
Enter fullscreen mode Exit fullscreen mode

overrepo is still in beta. It's MIT, it works on a real 195-repo fleet, and it has rough edges I haven't hit yet because I haven't run it on your fleet.

That's why feedback matters right now, while the shape is still easy to change. If you juggle dozens or hundreds of repositories, try it and tell me what breaks, what's missing, what the map should contain. Issues, ideas and pull requests are all welcome. Let's build it together on GitHub: github.com/maximeshr/overrepo. The package lives on npm.


I'm Maxime. I used to write code; now I run the kitchen: orchestrating AI agents and scaling agentic workflows across many codebases. More at okq.me.

Follow me here for more on agent-friendly tooling and stories, and tell me in the comments how you handle your multi-repo setup.

Top comments (1)

Collapse
 
max_quimby profile image
Max Quimby •

The "connecting dots needs dots" line lands. The expensive failure mode we kept hitting wasn't agents getting lost — it was every task starting from zero: re-grepping the tree, rediscovering the same service boundaries, spending a third of its context budget just reconstructing a map that was identical to last run's. A written manifest the agent reads first collapses that into one cheap read.

One thing I'd add from running agents in ephemeral VMs: the selector matters as much as the manifest. We found the win isn't just "the agent knows 195 repos exist" — it's that it can resolve a task to the 2-3 it actually needs before cloning anything, so a two-repo fix doesn't drag the whole fleet onto disk. Tag/path selection is the part that turns this from documentation into infra.

Did you hit any drift problems keeping the manifest honest as repos get added? That's the maintenance cost I'd worry about — a stale map is almost worse than none, because the agent trusts it.