DEV Community

Neeraj Singhi
Neeraj Singhi

Posted on Originally published at neerajsinghi.com

Structured Log Correlation Without Trace Context Propagation: The Hidden Gap in Go Microservice Observability

The Problem Nobody Talks About

Most observability stacks are designed assuming trace context propagates cleanly end-to-end. In practice, it often doesn't. A background goroutine fires without inheriting the parent context. A queue consumer reconstructs a request from raw bytes with no W3C traceparent header. A third-party SDK calls your callback with a fresh context.Background(). A timeout kills a span mid-flight and the child trace is orphaned.

In these gaps, structured logs are your only forensic primitive. If they're not designed to correlate independently, incidents become guesswork.

This article addresses the design problem of making structured logs reliably correlatable without assuming trace context is always present—covering log field contracts, propagation fallbacks, the mechanics of log-only correlation in Go, and the failure modes that burn incident response time.


Why Trace-Only Observability Is Fragile

OpenTelemetry's propagation model works when HTTP or gRPC calls carry the traceparent header and every service extracts and re-injects it correctly. The fragile points are:

Async producers/consumers. When a service writes a job to SQS or a Kafka topic, the trace context must be explicitly serialized into message attributes or headers. If the consumer doesn't extract and reconstruct the span context, the distributed trace is severed. You now have two separate trace trees with no causal link.

Goroutine handoff without context threading. In Go, context cancellation and value propagation are explicit. A goroutine launched with go func() { ... }() that doesn't receive the parent context.Context is invisible to the tracer. Any spans created inside it are orphaned.

SDK and library callbacks. Database drivers, HTTP client wrappers, and third-party SDKs frequently invoke callbacks with a detached context. The span hierarchy breaks at that boundary.

Sampling decisions. If the head-based sampler drops a trace, all spans for that request disappear—including error spans. Logs are not dropped by the sampler, which makes them the surviving record for sampled-out errors.

The operational consequence: during an incident, you switch between your tracing backend and your log aggregator, manually copy-pasting IDs to reconstruct causality. This is expensive when MTTR matters.


The Correlation Field Contract

The fix isn't to abandon tracing—it's to define a log field contract that works with or without trace context.

Every log line should carry:

  • trace_id: extracted from the active span if available, otherwise a generated request-scoped ID
  • span_id: the current span ID, or empty string if no span is active
  • request_id: a stable, user-visible correlation handle generated at the edge (API gateway, load balancer, or ingress)
  • service_name, service_version: static fields baked at startup
  • component: the logical subsystem (e.g., auth, payment, worker)

The critical distinction: trace_id and request_id are not the same thing. A trace ID is owned by the tracing backend and may not survive sampling or async boundaries. A request_id is owned by your system and must propagate everywhere—headers, queue message attributes, job records, and log lines.


Go Implementation: A Logger That Bridges Both Worlds

package obslog

import (
    "context"
    "log/slog"
    "os"

    "go.opentelemetry.io/otel/trace"
)

type contextKey string

const requestIDKey contextKey = "request_id"

// WithRequestID attaches a stable request-scoped ID to the context.
// This must be set at the entry point: HTTP handler, queue consumer, cron trigger.
func WithRequestID(ctx context.Context, id string) context.Context {
    return context.WithValue(ctx, requestIDKey, id)
}

// FromContext extracts correlation fields and returns a slog.Logger
// with those fields pre-attached. Falls back gracefully when trace context
// is absent.
func FromContext(ctx context.Context, base *slog.Logger) *slog.Logger {
    attrs := []any{
        slog.String("service", serviceName),
        slog.String("version", serviceVersion),
    }

    // Request ID: always present if callers set it correctly.
    if rid, ok := ctx.Value(requestIDKey).(string); ok && rid != "" {
        attrs = append(attrs, slog.String("request_id", rid))
    }

    // Trace context: present only when a span is active and sampled.
    spanCtx := trace.SpanFromContext(ctx).SpanContext()
    if spanCtx.IsValid() {
        attrs = append(attrs,
            slog.String("trace_id", spanCtx.TraceID().String()),
            slog.String("span_id", spanCtx.SpanID().String()),
            slog.Bool("trace_sampled", spanCtx.IsSampled()),
        )
    }

    return base.With(attrs...)
}

var (
    serviceName    = os.Getenv("SERVICE_NAME")
    serviceVersion = os.Getenv("SERVICE_VERSION")
)
Enter fullscreen mode Exit fullscreen mode

Usage in an HTTP handler:

func (h *Handler) CreateOrder(w http.ResponseWriter, r *http.Request) {
    ctx := r.Context()
    requestID := r.Header.Get("X-Request-ID")
    if requestID == "" {
        requestID = generateID() // UUID or ULID
    }
    ctx = obslog.WithRequestID(ctx, requestID)
    w.Header().Set("X-Request-ID", requestID) // echo back for clients

    log := obslog.FromContext(ctx, h.baseLogger)
    log.Info("create_order.start", slog.String("user_id", userID))

    // Pass ctx downstream—queue publish, DB call, gRPC call all get the same ctx.
    if err := h.orderSvc.Create(ctx, order); err != nil {
        log.Error("create_order.failed", slog.Any("error", err))
        // ...
    }
}
Enter fullscreen mode Exit fullscreen mode

Usage in a queue consumer where no HTTP context exists:

func (c *Consumer) Process(msg *sqs.Message) {
    requestID := aws.StringValue(msg.MessageAttributes["RequestId"].StringValue)
    ctx := obslog.WithRequestID(context.Background(), requestID)
    // Reconstruct OTel span context from message attributes if available.
    // If not, logs still correlate via request_id.
    log := obslog.FromContext(ctx, c.baseLogger)
    log.Info("consumer.process", slog.String("message_id", *msg.MessageId))
}
Enter fullscreen mode Exit fullscreen mode

The structural guarantee: every log line is queryable by request_id regardless of whether the trace context survived the async boundary.


Propagation Across Async Boundaries

For SQS, SNS, or Kafka, treat message attributes as the propagation envelope:

// Publishing: inject both request_id and trace context.
attrs := map[string]*sqs.MessageAttributeValue{
    "RequestId": {
        DataType:    aws.String("String"),
        StringValue: aws.String(requestID),
    },
}
// OTel propagator injects traceparent and tracestate into the map.
otel.GetTextMapPropagator().Inject(ctx, sqsCarrier(attrs))
Enter fullscreen mode Exit fullscreen mode

The consumer extracts both independently. If the traceparent is malformed or absent, the request ID still provides correlation. This is the fallback contract.


Failure Modes and Their Operational Cost

Missing request_id at the entry point. If the edge service doesn't generate or forward the request ID, all downstream logs are uncorrelated. You can query by time range and user ID, but under concurrent load you'll see interleaved log lines from multiple requests. Fix: enforce request ID generation at ingress; treat its absence as a bug.

request_id not threaded through goroutines. A background goroutine launched from a handler without the parent context loses the request ID. This is the Go-specific trap: forgetting that context.Context carries values, not just cancellation signals. Fix: audit any go func() that does I/O; pass context explicitly or use a worker pool that accepts context.

Log level filtering hiding error context. Setting the production log level to WARN suppresses INFO lines that carry the causal chain leading to an error. During incident diagnosis, you need the request flow. Fix: emit structured audit events at INFO level for significant state transitions, separate from verbose debug logs. Consider dynamic log level adjustment via a control endpoint that atomically updates slog.LevelVar.

High-cardinality fields in log aggregators. Adding a request_id to every log line is correct. Adding unbounded fields like raw SQL queries or full request bodies creates ingestion cost and index bloat in aggregators like OpenSearch or Loki. Fix: define a field allowlist; hash or truncate high-cardinality values that are needed for debugging but not for aggregation.


Connecting Logs to Metrics: The Middle Layer

For SLO alerting, metrics are the trigger. For incident diagnosis, traces and logs are the investigation layer. The gap is connecting a firing alert—error_rate > 1% over 5m—to the specific requests that contributed to it.

One pattern: emit a structured log event for every request that breaches its per-request latency budget or returns a 5xx, including the request_id, error category, and latency. A log-based metric (Loki's count_over_time or a CloudWatch metric filter) can then produce the same alert signal while leaving the full log body queryable.

This means a single alert can link directly to the log stream for the degraded request class, rather than requiring a manual join between your metrics dashboard and your log aggregator.


Decision Framework

When evaluating your service's log correlation posture, work through these questions:

  1. Does every entry point generate a stable request_id? HTTP, gRPC, queue consumers, cron triggers, and event-driven callbacks each need an explicit injection point.

  2. Does request_id propagate through async boundaries? SQS/Kafka message attributes, job queues, and outbox records should carry it.

  3. Does your slog (or zap/zerolog) logger extract both request_id and trace context from context.Context at every log call? Not just at the entry point—at every subsystem boundary.

  4. Are sampled-out traces covered? Logs for dropped traces must still be queryable. Confirm your sampler configuration doesn't suppress error logs.

  5. Can you reproduce a request's full log sequence in your aggregator using only request_id, without knowing the time range or service? If not, your propagation has gaps.

Tracing and structured logging are complementary, not redundant. Traces provide the shape of distributed execution; logs provide the content. Designing them to correlate independently—rather than assuming the trace context is always available—is what makes the observability stack trustworthy under the conditions that matter most: production failures in async, sampled, or context-severed environments.

Top comments (0)