DEV Community

Cover image for Go WaitGroup.Go in 1.25: When I Still Reach for errgroup
Qasim Parray
Qasim Parray

Posted on Originally published at abrarqasim.com

Go WaitGroup.Go in 1.25: When I Still Reach for errgroup

Short version for the impatient: since Go 1.25, wg.Go(f) replaces the Add(1), go func(), defer Done() routine, and I now use it everywhere I used to type that routine by hand. If your goroutines can fail, you still want errgroup. The rest of this post is why I draw the line there.

I have written wg.Add(1) a few thousand times. I have also written it in the wrong place more often than I like to admit, usually inside the goroutine, where it races with Wait. The compiler never complains. The program just returns before the work is done, now and then, which is exactly the kind of bug that passes on my laptop and fails on a CI box at 2am. So when the Go 1.25 release notes described WaitGroup.Go as a method that "makes the common pattern of creating and counting goroutines more convenient", I read that as a polite way of saying the old pattern was a footgun.

Below: the before and after, the one place I still avoid the new method, and the rule I use to pick between it and errgroup.

The old pattern and where it bites

Here is the version most of us have in muscle memory. Fetch a list of URLs concurrently, wait for all of them.

var wg sync.WaitGroup
for _, u := range urls {
    wg.Add(1)
    go func() {
        defer wg.Done()
        fetch(u)
    }()
}
wg.Wait()
Enter fullscreen mode Exit fullscreen mode

Several things have to be right at once. Add has to run before the go statement, not inside it. Done has to be deferred so an early return can't skip it. And the closure has to capture the loop variable correctly. That last one used to be the classic trap. Go 1.22 gave each iteration its own copy of the variable, so the code above is safe on a modern toolchain.

There's a catch I keep running into on client repos, though. The per-iteration behavior depends on the language version declared in go.mod. A module that still says go 1.21 keeps the old semantics even when you build it with a new compiler. That is an easy one to miss in a repo nobody has touched for two years, and on the old semantics every goroutine in the loop can end up fetching the last URL.

Now the misplaced Add, which is the bug I actually care about:

for _, u := range urls {
    go func() {
        wg.Add(1) // wrong: Wait may already have returned
        defer wg.Done()
        fetch(u)
    }()
}
wg.Wait()
Enter fullscreen mode Exit fullscreen mode

This compiles. It even works most of the time, because goroutines often get scheduled fast enough. The 1.25 release notes mention that go vet gained a waitgroup analyzer that reports misplaced calls to WaitGroup.Add, which is a nice safety net. But I'd rather not write the bug at all.

What WaitGroup.Go changes

The documentation for the method is short. It calls f in a new goroutine and adds that task to the group. When f returns, the task is removed. Same loop, new version:

var wg sync.WaitGroup
for _, u := range urls {
    wg.Go(func() { fetch(u) })
}
wg.Wait()
Enter fullscreen mode Exit fullscreen mode

Three lines became one. More to the point, there's no Add left to misplace and no Done left to forget. A whole class of mistakes disappears because the method does the counting for you.

I want to be straight about what this isn't. It isn't a new capability. Everything wg.Go does, you could already do. It's a convenience, and the reason I like it is that conveniences which remove a bug class are worth more than features. You also need go 1.25 or later in go.mod, so if you maintain a library that promises support for older toolchains, you can't use it yet. Application code is fair game.

One more practical detail. The function takes no arguments and returns nothing, so anything you want to pass in comes through the closure. If you need results back, don't reach for a shared slice and append from many goroutines. Give each goroutine its own slot:

results := make([]Response, len(urls))
var wg sync.WaitGroup
for i, u := range urls {
    wg.Go(func() { results[i] = fetch(u) })
}
wg.Wait()
Enter fullscreen mode Exit fullscreen mode

Each goroutine writes to a different index, so there's no race on the slice elements, and the results come back in input order for free. Run it under go test -race anyway. I trust the race detector more than my own reading of my own code.

Where it falls over: errors

wg.Go takes a func(). There's no error return, and that's where the convenience stops. The moment one of those fetches can fail and you care, you're back to inventing plumbing: an error channel, a mutex around a "first error" variable, a context you cancel by hand. I've written all three. None of them is hard, and all of them are about fifteen lines I'd rather not own.

This is the case errgroup was built for. It lives in golang.org/x/sync, so it's an extra dependency, but a very boring one.

g, ctx := errgroup.WithContext(ctx)
g.SetLimit(8)
for _, u := range urls {
    g.Go(func() error {
        return fetch(ctx, u)
    })
}
if err := g.Wait(); err != nil {
    return fmt.Errorf("fetching urls: %w", err)
}
Enter fullscreen mode Exit fullscreen mode

Per the package docs, Wait blocks until every function passed to Go has returned, then gives you the first non-nil error. The context from WithContext is canceled the first time one of those functions returns an error, or the first time Wait returns. So a failure in one fetch tells the others to stop, as long as fetch actually honors the context. That condition is the part people forget. A function that ignores ctx will happily run to the end while the group waits for it.

Bounded concurrency is the other half

The line g.SetLimit(8) above deserves its own section, because it solves something WaitGroup never did. Goroutines are cheap. The things they talk to are not. Ten thousand goroutines hitting an API at once means ten thousand sockets, a pile of file descriptors, and a rate limiter that will cheerfully ban you.

With a plain WaitGroup you bolt on a semaphore, usually a buffered channel, and acquire and release around the work. It's a well-known idiom, and it's another place to get the ordering wrong. With errgroup, SetLimit caps the number of active goroutines, and Go blocks until a slot frees up. There's also TryGo, which returns false instead of blocking, handy when you'd rather drop or queue work elsewhere than stall the loop.

One consequence to plan for: because Go blocks, the loop that calls it is itself throttled. If you were counting on the loop to finish quickly and do other things afterward, it won't. Put the dispatch loop in its own goroutine if that matters.

The rule I use now

When the fan-out is small, nothing can fail, and I don't need a cap, I use wg.Go. Think of warming three caches at startup, or running a handful of independent health probes in a script. When any call can return an error, when cancellation matters, or when the input size isn't something I control, I use errgroup with a limit. When the work produces a stream of results that feed another stage, I stop thinking about wait groups at all and design the pipeline with channels, the way the Go team's pipelines article lays out.

Notice that this isn't a ranking. wg.Go doesn't replace errgroup, and errgroup doesn't make WaitGroup obsolete. They answer different questions: "have they all finished?" versus "did any of them fail, and should the rest stop?" Most of the confusion I see in code review comes from using the first tool to answer the second question.

Migrating without breaking things

I don't trust automated rewrites for concurrency code, so I do this by hand. Search for the old shape:

rg -n 'wg\.Add\(1\)' --type go
Enter fullscreen mode Exit fullscreen mode

For each hit, check three things. Is the goroutine body fallible? If yes, it's an errgroup candidate, not a wg.Go one. Does anything else call Add with a number larger than one? Leave those alone, they're doing something cleverer than the standard loop. And is the module's go directive at 1.25 or higher? Then swap the pattern, delete the defer wg.Done(), and run go test -race ./....

I did a similar deletion pass when generic methods landed, and wrote about it in my Go 1.27 post about the helper package I removed. The feeling is the same. The best thing a language upgrade can do is let you delete code you were never proud of. If you want to see the kind of Go services this habit ends up in, my portfolio has the client work I can talk about.

This week, pick one service you own, run that rg command, and convert the single simplest loop. Don't touch the fallible ones yet. Run the race detector, commit it, and see how much quieter the diff feels than you expected.


Originally published at abrarqasim.com. I write there about React, PHP, Rust, Go and the AI tooling around them.

Top comments (0)