DEV Community

Ayi NEDJIMI
Ayi NEDJIMI

Posted on

Implementing graceful shutdown in Go HTTP servers

When you send SIGTERM to a Go HTTP server that's handling 50 concurrent requests, all 50 die instantly. No cleanup, no drain, no responses. In production this means failed API calls, half-written database rows, and clients retrying on errors that shouldn't have happened. Graceful shutdown is the difference between a clean deployment and an incident.

Why Abrupt Shutdown Breaks Things

The default net/http server has no concept of "wait for active connections." When you kill the process, the OS closes all file descriptors, and every in-flight HTTP connection gets a TCP RST. Your clients see connection errors. Your database sees abandoned transactions.

Three scenarios where this matters most:

  • Zero-downtime deploys with rolling restarts (Kubernetes, ECS, Nomad) — the old pod receives a SIGTERM while the load balancer is still routing traffic to it
  • Long-running requests like file uploads, batch processing, or SSE streams
  • Background workers triggered by HTTP handlers that need to finish before the process exits

The Basics: Catching SIGTERM and Draining Connections

Go's net/http package has had Server.Shutdown since Go 1.8. Here's the minimal pattern:

package main

import (
    "context"
    "log"
    "net/http"
    "os"
    "os/signal"
    "syscall"
    "time"
)

func main() {
    mux := http.NewServeMux()
    mux.HandleFunc("/health", func(w http.ResponseWriter, r *http.Request) {
        w.WriteHeader(http.StatusOK)
        w.Write([]byte("ok"))
    })

    srv := &http.Server{
        Addr:         ":8080",
        Handler:      mux,
        ReadTimeout:  10 * time.Second,
        WriteTimeout: 30 * time.Second,
        IdleTimeout:  60 * time.Second,
    }

    // Start server in a goroutine so it doesn't block
    go func() {
        log.Println("Server listening on :8080")
        if err := srv.ListenAndServe(); err != nil && err != http.ErrServerClosed {
            log.Fatalf("ListenAndServe: %v", err)
        }
    }()

    // Block until we receive SIGTERM or SIGINT
    quit := make(chan os.Signal, 1)
    signal.Notify(quit, syscall.SIGTERM, syscall.SIGINT)
    <-quit

    log.Println("Shutdown signal received")

    // Give in-flight requests 30 seconds to complete
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()

    if err := srv.Shutdown(ctx); err != nil {
        log.Fatalf("Server forced to shutdown: %v", err)
    }

    log.Println("Server exited cleanly")
}
Enter fullscreen mode Exit fullscreen mode

srv.Shutdown(ctx) stops accepting new connections, then waits for active connections to finish or for the context to expire — whichever comes first. http.ErrServerClosed is the expected error from ListenAndServe after Shutdown is called; you need to ignore it.

Handling Long-Running Requests With Context Propagation

The simple pattern above has a blind spot: a handler that ignores its context will block until the shutdown timeout expires, even if the work is already futile. Handlers need to respect r.Context() cancellation signals.

Here's a handler that respects cancellation:

func longRunningHandler(w http.ResponseWriter, r *http.Request) {
    workDone := make(chan struct{})

    go func() {
        defer close(workDone)
        // Simulate work: in production this is a DB query, external API call, etc.
        time.Sleep(10 * time.Second)
    }()

    select {
    case <-workDone:
        w.WriteHeader(http.StatusOK)
        w.Write([]byte(`{"status":"done"}`))
    case <-r.Context().Done():
        log.Printf("Request cancelled: %v", r.Context().Err())
        // Don't write — the connection is already closing
        return
    }
}
Enter fullscreen mode Exit fullscreen mode

When Shutdown is called, Go cancels the context of all active requests. Handlers that select on r.Context().Done() can exit early instead of holding up the drain window.

For database calls, pass the request context directly:

func userHandler(db *sql.DB) http.HandlerFunc {
    return func(w http.ResponseWriter, r *http.Request) {
        row := db.QueryRowContext(r.Context(),
            "SELECT name FROM users WHERE id = $1",
            r.URL.Query().Get("id"),
        )
        var name string
        if err := row.Scan(&name); err != nil {
            if r.Context().Err() != nil {
                return // shutdown in progress, don't write error response
            }
            http.Error(w, "not found", http.StatusNotFound)
            return
        }
        w.Write([]byte(name))
    }
}
Enter fullscreen mode Exit fullscreen mode

Coordinating Background Workers

HTTP handlers often enqueue work — sending emails, processing files, updating search indexes. If those workers run as goroutines, they can outlive the HTTP server or get cut off mid-job. The right tool is sync.WaitGroup:

type App struct {
    wg sync.WaitGroup
}

func (a *App) submitJob(ctx context.Context, payload []byte) {
    a.wg.Add(1) // increment before spawning the goroutine
    go func() {
        defer a.wg.Done()
        if err := processPayload(ctx, payload); err != nil {
            log.Printf("job failed: %v", err)
        }
    }()
}

func (a *App) shutdown(srv *http.Server, timeout time.Duration) {
    ctx, cancel := context.WithTimeout(context.Background(), timeout)
    defer cancel()

    // Stop accepting HTTP requests first
    srv.Shutdown(ctx)

    // Then wait for background jobs within the same deadline
    done := make(chan struct{})
    go func() {
        a.wg.Wait()
        close(done)
    }()

    select {
    case <-done:
        log.Println("All background jobs completed")
    case <-ctx.Done():
        log.Println("Shutdown timeout: some jobs may be incomplete")
    }
}
Enter fullscreen mode Exit fullscreen mode

Two details worth calling out: the wg.Add(1) must happen before the goroutine starts — incrementing inside the goroutine is a race condition if wg.Wait() is called concurrently. And sharing the same context deadline between HTTP drain and worker drain keeps the total shutdown window predictable.

Kubernetes and the SIGTERM → SIGKILL Gap

In Kubernetes, pod termination follows this sequence:

  1. Pod gets removed from Endpoints (stops receiving new traffic)
  2. SIGTERM is sent to the container
  3. After terminationGracePeriodSeconds (default: 30s), SIGKILL arrives

The catch: steps 1 and 2 happen in parallel, not sequentially. iptables rules may not have propagated before your server receives SIGTERM, which means requests can still arrive for a few seconds after shutdown starts.

The practical fix is a pre-shutdown sleep:

<-quit
log.Println("SIGTERM received, waiting for load balancer drain")
time.Sleep(5 * time.Second) // wait for iptables/IPVS propagation
srv.Shutdown(ctx)
Enter fullscreen mode Exit fullscreen mode

Five seconds covers most clusters. Pair this with returning 503 from your readiness probe immediately after SIGTERM — that makes the load balancer stop routing to you faster. For a full deployment readiness checklist covering signal handling and zero-downtime releases, the AYI NEDJIMI hardening checklists have a dedicated section.

The Takeaway

Graceful shutdown in Go requires three things working together: catching SIGTERM, calling srv.Shutdown(ctx) with a realistic timeout, and making sure every handler and background worker respects context cancellation. Skip any one of these and you either drop requests or get goroutine leaks at shutdown time.

The patterns above cover the 95% case. The remaining 5% — server-sent events with long-lived connections, Unix domain sockets, streaming responses — follow the same principles but need explicit stream termination logic. Start with this foundation and add edge case handling as your specific workload demands it.


I run AYI NEDJIMI Consultants, a cybersecurity consulting firm. We publish free security hardening checklists — PDF and Excel.

Top comments (0)