Distributed Tracing: Making Microservices Observable
Introduction
Distributed Tracing is the superpower you need to debug microservices. When a user clicks a button and your application fails, where do you look? With monoliths, you check one server's logs. With microservices, that button click might trigger requests across 15 different services in 500ms, and if something breaks, you need to trace exactly where.
Distributed Tracing answers the fundamental question: What happened to this request as it traveled through my system?
It's not logging. It's not metrics. It's the third pillar of observability—the one that connects the dots.
What is Distributed Tracing?
Distributed Tracing is the practice of tracking requests as they flow through your entire system, capturing the journey from the user's browser to the database and back.
Core Concept
A trace is a complete request journey. It's broken into spans, where each span represents a unit of work (an HTTP call, a database query, a message published to a queue).
User Request (HTTP GET /orders/123)
├── Order Service Span (50ms)
│ ├── GetOrder DB Query (15ms)
│ └── Call Payment Service HTTP (30ms)
│ └── Payment Service Span (25ms)
│ ├── Process Payment (20ms)
│ └── Audit Log Write (3ms)
├── Inventory Service Call (35ms)
│ └── Inventory Service Span (35ms)
│ └── Check Stock DB Query (30ms)
└── Send Confirmation Email (12ms)
└── Email Service Span (12ms)
Total Trace Duration: 115ms
Critical Path: Order → Payment → Inventory
Key Concepts
Trace ID: Unique identifier for the entire request journey
Span ID: Unique identifier for a single unit of work within a trace
Parent Span ID: Links child spans to their parent, building the tree structure
Tags: Key-value pairs for metadata (userId=42, status=success)
Logs: Timestamped events within a span (exceptions, checkpoints)
Baggage: Data passed through the entire trace (user ID, feature flags)
Why Distributed Tracing Matters
The Problem
Without Distributed Tracing:
User: "My order failed!"
You: "Let me check the logs..."
Order Service logs... ✅ Status 200
Payment Service logs... ✅ No errors
Inventory Service logs... ❌ Timeout after 30s
But why? When? Which order? How long was it waiting?
You have fragments of the story, not the full narrative.
With Distributed Tracing:
Trace ID: 550e8400-e29b-41d4-a716-446655440000
Payment Service (50ms)
├── Database Query (10ms)
├── External API Call (35ms) ← SLOW!
│ └── POST https://payment-gateway.com/charge
│ Response Time: 35ms (SLA: 5ms)
│ Error: Timeout
└── Retry Circuit Breaker Open
You immediately see: The payment gateway is slow.
Business Impact
- Faster Debugging: Find root cause in minutes, not hours
- Better SLO Management: Know exactly where latency comes from
- Cost Optimization: Identify expensive operations
- Performance Tuning: See actual bottlenecks
- Customer Support: Answer with data, not guesses
- Capacity Planning: Understand service dependencies
Distributed Tracing Architecture
Components
Application Code (Instrumented)
↓
Tracer Library (OpenTelemetry)
↓
Span Exporter (HTTP/gRPC)
↓
Trace Collector (Jaeger, Datadog, Lightstep)
↓
Trace Storage (Jaeger Backend, Elasticsearch)
↓
Trace Visualization (Jaeger UI)
↓
Alerts & Analysis
Standards
OpenTelemetry (OTEL): Industry standard for observability
Jaeger: Open-source distributed tracing backend by Uber
Zipkin: Alternative open-source tracer
Implementing Distributed Tracing: Java/Spring Boot
1. Add Dependencies
<dependency>
<groupId>io.opentelemetry.instrumentation</groupId>
<artifactId>opentelemetry-spring-boot-starter</artifactId>
<version>0.32.0</version>
</dependency>
<dependency>
<groupId>io.opentelemetry.exporter</groupId>
<artifactId>opentelemetry-exporter-jaeger-thrift</artifactId>
<version>1.32.0</version>
</dependency>
2. Configure OpenTelemetry
otel.sdk.disabled=false
otel.traces.exporter=jaeger_thrift
otel.exporter.jaeger.agent.host=localhost
otel.exporter.jaeger.agent.port=6831
otel.service.name=order-service
otel.traces.sampler=always_on
3. Instrument Code
@RestController
@RequestMapping("/api/orders")
public class OrderController {
private static final Tracer tracer = GlobalOpenTelemetry
.getTracer("order-service");
@PostMapping
public ResponseEntity<Order> createOrder(@RequestBody OrderRequest req) {
Span span = tracer.spanBuilder("order-creation")
.setParent(Context.current())
.startSpan();
try (Scope scope = span.makeCurrent()) {
span.setAttribute("customer.id", req.customerId());
span.setAttribute("order.amount", req.amount());
Order order = orderService.createOrder(req);
span.setStatus(StatusCode.OK);
return ResponseEntity.ok(order);
} catch (Exception e) {
span.recordException(e);
span.setStatus(StatusCode.ERROR, e.getMessage());
throw e;
} finally {
span.end();
}
}
}
4. HTTP Client Tracing
@Component
public class ExternalServiceClient {
private final RestTemplate restTemplate;
private final Tracer tracer;
public PaymentGatewayResponse callPaymentGateway(String customerId) {
Span span = tracer.spanBuilder("external-payment-gateway-call")
.startSpan();
try (Scope scope = span.makeCurrent()) {
span.setAttribute("external.service", "payment-gateway");
span.setAttribute("customer.id", customerId);
HttpHeaders headers = new HttpHeaders();
headers.set("X-Trace-ID", span.getSpanContext().getTraceId());
HttpEntity<PaymentRequest> entity = new HttpEntity<>(
new PaymentRequest(customerId),
headers
);
ResponseEntity<PaymentGatewayResponse> response = restTemplate
.exchange("https://payment-gateway/charge",
HttpMethod.POST, entity,
PaymentGatewayResponse.class);
return response.getBody();
} finally {
span.end();
}
}
}
5. Database Query Tracing
@Repository
public class OrderRepository extends JpaRepository<Order, Long> {
private final Tracer tracer;
public Order findByCustomerId(Long customerId) {
Span span = tracer.spanBuilder("db-query-find-orders")
.startSpan();
try (Scope scope = span.makeCurrent()) {
span.setAttribute("db.system", "postgresql");
span.setAttribute("db.operation", "SELECT");
span.setAttribute("customer.id", customerId);
long startTime = System.currentTimeMillis();
Order order = super.findByCustomerId(customerId);
long duration = System.currentTimeMillis() - startTime;
span.setAttribute("db.query.duration_ms", duration);
return order;
} finally {
span.end();
}
}
}
6. Asynchronous Task Tracing
@Service
public class OrderProcessingService {
private final Tracer tracer;
@Async
public void processOrderAsync(Order order) {
Span span = tracer.spanBuilder("async-order-processing")
.startSpan();
try (Scope scope = span.makeCurrent()) {
span.setAttribute("order.id", order.getId());
sendConfirmationEmail(order);
updateInventory(order);
publishAnalyticsEvent(order);
span.setStatus(StatusCode.OK);
} catch (Exception e) {
span.recordException(e);
span.setStatus(StatusCode.ERROR);
throw e;
} finally {
span.end();
}
}
}
Real-World Scenario: E-Commerce Request Tracing
User clicks "Buy Now" on checkout page
Trace ID: 550e8400-e29b-41d4-a716-446655440000
Timeline:
0ms └─ API Gateway Span (115ms total)
5ms └─ Order Service Span (110ms)
10ms ├─ Authentication Check (5ms)
15ms ├─ Validate Cart (3ms)
18ms └─ CreateOrder DB (8ms)
30ms └─ Payment Service Call (80ms)
35ms └─ Payment Service Span (75ms)
40ms ├─ Validate Card (2ms)
45ms ├─ Call Payment Gateway (65ms) ← SLOW!
110ms └─ Update DB (3ms)
115ms └─ Response sent
Critical Issues:
1. Payment Gateway: 65ms (SLA: 10ms) ← VIOLATION
2. Total: 115ms (could be 50ms if gateway fixed)
3. No parallelization: Services called sequentially
Common Pitfalls to Avoid
❌ Pitfall 1: Too Few Spans
Wrong: Only instrument HTTP endpoints
Right: Instrument at multiple levels (HTTP, DB, cache, external APIs)
❌ Pitfall 2: Sampling Everything in Production
Wrong: Enable 100% tracing (massive storage/cost)
Right: Use adaptive sampling (high error = 100%, normal = 10%)
❌ Pitfall 3: No Context Propagation
Wrong: Trace IDs lost when calling different services
Right: Use W3C Trace Context standard headers
❌ Pitfall 4: Noisy Traces
Wrong: Trace every cache hit, every local function
Right: Trace meaningful operations (external calls, slow ops, errors)
❌ Pitfall 5: Insufficient Tag Cardinality
Wrong: action=create (can't drill down)
Right: action=create, resource.id=12345, user.id=789 (can drill down)
Best Practices
1. Use Semantic Span Names
// ❌ Bad
tracer.spanBuilder("process").startSpan();
// ✅ Good
tracer.spanBuilder("inventory.check-stock").startSpan();
tracer.spanBuilder("payment.process-charge").startSpan();
2. Add Rich Context
span.setAttribute("user.id", userId);
span.setAttribute("order.id", orderId);
span.setAttribute("order.amount", amount);
span.setAttribute("currency", "USD");
3. Record Exceptions Properly
try {
risky_operation();
} catch (TimeoutException e) {
span.recordException(e);
span.setAttribute("exception.type", e.getClass().getName());
span.setAttribute("exception.timeout_ms", 5000);
span.setStatus(StatusCode.ERROR, e.getMessage());
}
4. Use Sampling Wisely
otel.traces.sampler=parentbased_traceidratio
otel.traces.sampler.arg=0.1
5. Correlate with Logs
MDC.put("traceId", span.getSpanContext().getTraceId());
logger.info("Processing order");
Tools & Ecosystem
Open Source
- Jaeger: Distributed tracing backend (Uber)
- Zipkin: Distributed tracing (Twitter/Square)
- OpenTelemetry: Instrumentation standard (CNCF)
Managed Services
- Datadog: APM + distributed tracing
- AWS X-Ray: AWS-native tracing
- Google Cloud Trace: GCP-native tracing
- Lightstep: Enterprise tracing
Conclusion
Distributed Tracing transforms debugging from guesswork to science. Instead of "the checkout is slow," you have data: "Payment gateway calls are taking 3x longer than SLA."
Key Takeaways:
- Traces = complete request journey
- Spans = individual units of work
- OpenTelemetry = standard instrumentation
- Jaeger/Zipkin = open-source backends
- Use sampling in production
- Correlate traces, logs, and metrics
Start simple: instrument critical paths. Expand as you grow.
What's your biggest pain point debugging microservices? Share in the comments!
Top comments (0)