DEV Community

Ganesh Joshi
Ganesh Joshi

Posted on Originally published at ganeshjoshi.dev

Turborepo for AI Product Monorepos and Shared Pipelines

This post was created with AI assistance and reviewed for accuracy before publishing.

AI products fragment across runtimes quickly. There is a web app, one or more background workers doing embedding or evaluation, sometimes an edge function, and underneath them shared packages holding prompts, model adapters, and schema definitions. Every one of those consumers needs the shared code, and every change to it should rebuild exactly the things affected and nothing else.

That is what Turborepo is for. Most of the value comes from getting two things right: the task graph and the cache inputs.

Declare the graph, not the order

The ^ prefix means "this task depends on the same task in my dependencies", which is what makes a package build before the app importing it:

{
  "tasks": {
    "build": {
      "dependsOn": ["^build"],
      "outputs": [".next/**", "!.next/cache/**", "dist/**"]
    },
    "test": { "dependsOn": ["^build"] },
    "lint": { "dependsOn": [] },
    "dev": { "cache": false, "persistent": true }
  }
}
Enter fullscreen mode Exit fullscreen mode

Two details are easy to get wrong. lint genuinely has no dependencies, so declaring one serialises work that could run in parallel across every package. And dev must be marked persistent and uncached, otherwise Turborepo waits for a long-running process to exit before starting anything downstream.

The outputs list is what gets restored on a cache hit. Omit a directory and a "cached" build leaves you without files you needed. Excluding .next/cache is deliberate: it is large, it is a cache of a cache, and storing it makes every hit slower to download than the build it replaced.

Environment variables are cache inputs

This is the mistake specific to AI projects, and it produces genuinely confusing failures.

Turborepo hashes inputs to decide whether a task can be restored from cache. Environment variables are not included unless you declare them. So if a build embeds a model name or an API base URL, and you change that variable, Turborepo sees identical inputs and hands back the old artifact.

{
  "tasks": {
    "build": {
      "env": ["MODEL_NAME", "EMBEDDING_MODEL", "NEXT_PUBLIC_*"],
      "dependsOn": ["^build"]
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

The symptom is a deploy that changes nothing. You switch the model, the build succeeds, and production keeps using the old one because the artifact was restored rather than rebuilt.

The reverse mistake matters too. Declaring a variable that changes on every run, such as a build timestamp or a CI job id, means the hash never matches and you never get a cache hit at all. If your hit rate is zero, this is almost always why.

Secrets should not be build inputs in the first place. An API key belongs in the runtime environment, not baked into an artifact, and adding it to env implies it can change the build output, which is a design you do not want.

Remote caching is the point in CI

Local caching helps one developer. Remote caching means a build produced on someone's laptop is reused by CI, and a build from CI is reused by everyone else.

For an AI monorepo where the web app and three workers all depend on the same prompt package, this is the difference between rebuilding four times and once.

Treat the cache token as a credential with write access to your build outputs. A leaked token that can write to the cache lets an attacker poison an artifact that your deploy then trusts. Give CI a token scoped to what it needs, and prefer read-only tokens anywhere that does not need to publish.

Scope commands so you are not building everything

The filter flag is what keeps day-to-day work fast:

turbo build --filter=web                # just the web app and its deps
turbo build --filter=...[origin/main]   # only what changed since main
turbo test --filter=@repo/prompts...    # the package and everything depending on it
Enter fullscreen mode Exit fullscreen mode

The last form is the one worth internalising. A trailing ... means "and everything downstream", so after changing a shared prompt package you can test exactly the consumers that could be affected. That is the correct blast radius for a change to shared AI logic, where a prompt tweak can alter behaviour in three services at once.

Keep model adapters in their own package

A structural recommendation that pays off repeatedly: put the code that talks to model providers in a single package with a stable interface, and let every app depend on that rather than calling providers directly.

It gives you one place to add a new provider, one place to change timeout handling, and one place that fails type-checking when a response shape changes. In a monorepo without this, provider details leak into four codebases and a provider migration becomes four migrations.

It also makes the cache work in your favour. Changing a worker does not invalidate the web app, and changing the adapter invalidates exactly the things that use it.

Check the summary when something feels wrong

When a build is slower than expected or a change did not take effect, ask Turborepo what it decided:

turbo build --summarize
Enter fullscreen mode Exit fullscreen mode

That writes a file listing every task, whether it hit the cache, and the hash inputs used. Comparing two summaries shows precisely which input changed, which is far faster than guessing at why a task rebuilt or, more often, why it did not.

Top comments (0)