DEV Community

Shehroz Ali
Shehroz Ali

Posted on Originally published at shehroztalks.medium.com on

A Complete Guide to Distributed Tracing in Kotlin and Spring Boot with OpenTelemetry and Grafana…

A Complete Guide to Distributed Tracing in Kotlin and Spring Boot with OpenTelemetry and Grafana Loki

In this blog, I’ll teach you how you can build end-to-end distributed tracing for your backend microservices in Kotlin, OpenTelemetry, Spring Boot and Grafana Loki.

By following the guide, you’ll be able to correlate hundreds and thousands of telemetry data (traces & logs) emitted by your backend services and achieve end-to-end observability to reduce your MTTRs for production incidents and make your developers life easier.

To give you a basic understanding of what we’ll be going to achieve, here’s a system-level diagram to help you understand all the components functioning together:


end-to-end observability

Capturing HTTP traffic within Spring Boot:

First, we need to enable HTTP network logging for our Spring Boot app, for that we’re going to use Zalando’s open-source Logbook which automatically logs HTTP requests and responses on your (preferred output writer). It works by intercepting HTTP traffic in your app and writing TRACE log using Slf4j which is a fascade for Logging in Java. In this example, we’re going to use Logback which is a logging library with auto-configuration support that comes already with spring-starter-web dependency.

First let’s configure Logback to capture INFO, ERROR and TRACE logs and outputs on console:

Setting up Logback appenders:

First we added, Console-Info to capture logs at INFO level which will likely be our classes/service logs emitting some business/user insights/actions (which later devs will use to understand code behaviour during debugging).

We also added another Console-Error to capture ERROR logs emitted by our classes to indicate an exception/error message in logs.

Lastly we added Console-Trace to capture TRACE logs. Since Zalando’s Logbook library write TRACE logs for HTTP traffic, we’ve passed the Console-Trace appender to logger org.zalando to use Console-Trace appender to write TRACE logs on System output. Here’s the final logback-logging.xml file:

<?xml version="1.0" encoding="UTF-8"?>
<included>

    <appender name="Console-Info" class="ch.qos.logback.core.ConsoleAppender">
        <!-- Targeting System.out -->
        <target>System.out</target>
        <!-- JSON Log Formatting -->
        <filter class="ch.qos.logback.classic.filter.LevelFilter">
            <level>INFO</level>
            <onMatch>ACCEPT</onMatch>
            <onMismatch>DENY</onMismatch>
        </filter>
        <encoder class="net.logstash.logback.encoder.LogstashEncoder">
            <fieldNames>
                <message>[ignore]</message>
                <levelValue>[ignore]</levelValue>
                <version>[ignore]</version>
            </fieldNames>
            <provider class="net.logstash.logback.composite.loggingevent.LoggingEventPatternJsonProvider">
                <pattern>
                    {
                    "message": "%replace(%.-150message){'${MASKPATTERNS}', ' *****'}"
                    }
                </pattern>
            </provider>
            <customFields>
                {"service": "${logHost}"}
            </customFields> <!-- Optional: Add custom fields -->
        </encoder>
    </appender>

    <appender name="Console-Error" class="ch.qos.logback.core.ConsoleAppender">
        <!-- Targeting System.out -->
        <target>System.err</target>
        <!-- JSON Log Formatting -->
        <filter class="ch.qos.logback.classic.filter.LevelFilter">
            <level>ERROR</level>
            <onMatch>ACCEPT</onMatch>
            <onMismatch>DENY</onMismatch>
        </filter>
        <encoder class="net.logstash.logback.encoder.LogstashEncoder">
            <fieldNames>
                <message>[ignore]</message>
                <levelValue>[ignore]</levelValue>
                <version>[ignore]</version>
            </fieldNames>
            <provider class="net.logstash.logback.composite.loggingevent.LoggingEventPatternJsonProvider">
                <pattern>
                    {
                    "message": "%replace(%.-150message){'${MASKPATTERNS}', ' *****'}",
                    "stack_trace": "%replace(%ex{short}){'${MASKPATTERNS}', ' *****'}"
                    }
                </pattern>
            </provider>
            <customFields>
                {"service": "${logHost}"}
            </customFields> <!-- Optional: Add custom fields -->
        </encoder>
    </appender>

    <appender name="Console-Trace" class="ch.qos.logback.core.ConsoleAppender">
        <!-- Targeting System.out -->
        <target>System.out</target>
        <!-- JSON Log Formatting -->
        <filter class="ch.qos.logback.classic.filter.LevelFilter">
            <level>TRACE</level>
            <onMatch>ACCEPT</onMatch>
            <onMismatch>DENY</onMismatch>
        </filter>
        <encoder class="net.logstash.logback.encoder.LogstashEncoder">
            <fieldNames>
                <message>[ignore]</message>
                <levelValue>[ignore]</levelValue>
                <version>[ignore]</version>
            </fieldNames>
            <provider class="net.logstash.logback.composite.loggingevent.LoggingEventPatternJsonProvider">
                <pattern>
                    {
                    "message": "#asJson{%message}"
                    }
                </pattern>
            </provider>
            <customFields>
                {"service": "${logHost}", "pod": "${podName}"}
            </customFields> <!-- Optional: Add custom fields -->
        </encoder>
    </appender>

    <!-- LOG everything at INFO level -->
    <root level="info">
        <appender-ref ref="Console-Info"/>
    </root>

    <logger name="org.zalando.logbook" level="INFO" additivity="false">
        <appender-ref ref="Console-Trace"/>
    </logger>
Enter fullscreen mode Exit fullscreen mode

In application.yml file:

logging:
  config: 'classpath:logback-logging.xml'
Enter fullscreen mode Exit fullscreen mode

Structured JSON Logging for Grafana Loki:

Here comes the important detail, since we’re going to structure logs on Grafana Loki based on labels and fields for filtering and querying, we’re going to emit logs as JSON output. For that we need to add this dependency for Logback (maven central).

        <dependency>
            <groupId>net.logstash.logback</groupId>
            <artifactId>logstash-logback-encoder</artifactId>
            <version>6.6</version>
        </dependency>
Enter fullscreen mode Exit fullscreen mode

This will allow us to output logs as JSON on system output and later we can relabel fields on Grafana Agent (or Alloy) for Loki.

Writing HTTP logs as JSON:

Now since our Logback supports writing logs in JSON format, we’ve added this line in our Console-Trace to write Zalando’s HTTP log as JSON on output. The %message captures the entire TRACE level log and output on console with label message like this:

{
 message: <zalando's TRACE log including request, response, headers>
}
Enter fullscreen mode Exit fullscreen mode

Later we can extract out individual fields like status, type, body, path and display on Grafana Loki as fields for filtering and querying logs.

Configuring Logbook in Kotlin:

Now, since our Logback is all ready to emit logs (INFO, ERROR, TRACE) on System output, now it’s finally time to configure Logbook for capturing HTTP traffic inside our app.

Below is the code for configuring Logbook in Kotlin for Spring Boot, add dependency as well:

        <dependency>
            <groupId>org.zalando</groupId>
            <artifactId>logbook-spring-boot-starter</artifactId>
            <version>3.10.0</version>
        </dependency>
Enter fullscreen mode Exit fullscreen mode

As it is a spring-starter dependency it already comes with lots of things pre-configured, let’s tweak it according to our needs:

@Configuration
open class LogbookAutoConfiguration(
    private val jsonBodyFilter: JSONBodyFilter,
    private val sinkConfig: SinkConfiguration,
) {

    @Bean
    open fun logbook(filterConfig: MDCFilter?, objectMapper: ObjectMapper): Logbook {
        return Logbook.builder()
            .correlationId(filterConfig)
            .sink(sinkConfig.defaultSink(objectMapper))
            .bodyFilter(BodyFilter.none())
            .condition(exclude(requestTo("/actuator/**")))
            .build()
    }
}
Enter fullscreen mode Exit fullscreen mode

We’ve excluded /actuator to log because this will be our Kubernetes liveness/readiness probe endpoint for health check, so k8s will hit this endpoint every after X mins/secs to keep pod up and healthy, therefore we excluded to avoid de-bloating our Loki storage.

@Component
open class JSONBodyFilter {

    open fun runFilter(): BodyFilter {
        return JsonBodyFilters.replaceJsonStringProperty(
            setOf("token", "refreshToken", "scopes"),
            " ****"
        )
    }
}
Enter fullscreen mode Exit fullscreen mode

Here, we’ve masked these fields in our response/request bodies as it contains JWT tokens of users as a security practice, you might don’t need it depending upon your infra/team security policies.

@Component
open class SinkConfiguration {

    open fun defaultSink(objectMapper: ObjectMapper): Sink {
        val formatter: HttpLogFormatter = JsonHttpLogFormatter(objectMapper)

        val writer: HttpLogWriter = DefaultHttpLogWriter()

        return DefaultSink(formatter, writer)
    }
}
Enter fullscreen mode Exit fullscreen mode

Here we’re using Zalando’s default HTTP log writer to write TRACE logs on output.

@Component
open class MDCFilter : CorrelationId {
    override fun generate(request: HttpRequest): String {
        val correlationId = UUID.randomUUID().toString()

        val httpMethod = request.method
        val httpPath = request.path
        val userId = request.headers["userId"]?.first()
        val clientKey = request.headers["X-Client-Key"]?.first()
        val ipAddress = request.headers["x-forwarded-for"]?.toString()
        val userAgent = request.headers["user-agent"]?.toString()

        MDC.put("httpMethod", httpMethod)
        MDC.put("httpPath", httpPath)
        MDC.put("userId", userId)
        MDC.put("correlationId", correlationId)
        MDC.put("ClientKey", clientKey)
        MDC.put("ipAddress", ipAddress)
        MDC.put("userAgent", userAgent)

        return correlationId
    }
}
Enter fullscreen mode Exit fullscreen mode

This is an important component, here we’ve extended CorrelationId class by Zalando and override generate() method, so whenever internally this method invokes for correlationId generation which will be attached to each HTTP log, we’re also adding the same correlationId in our MDC along with some other fields as well. This means whenever a HTTP request comes inn, here’s what will happen:

  1. Generate unique UUID as correlationId
  2. Attach correlationId in HTTP request/response TRACE log
  3. Also populate same correlationId in MDC
  4. Whenever we log (info, error) at class-level, same correlationId is attached because MDC is shared, allowing us to correlation HTTP logs with class-level logs

And, that’s it! We’re good to go to start writing HTTP logs on console/output propagating HTTP path, ip address, method and your own unique header fields correlating with your class-level logs.

Using Grafana Agent to ingest Kubernetes Pod logs:

Since, we’re running our Spring Boot app in Kubernetes and writing logs on stdout (system output), k8s maintain a file on the node with all logs data at /var/log/pods/__//0.log we can configure Grafana Agent to ingest this log file and perform some parsing to transform raw JSON logs into structure logs and forward it to Loki tenant.

This is a sample config for Grafana Alloy (similar to Grafana Agent as well):

// Discover pods running on this node
discovery.kubernetes "pods" {
  role = "pod"
}

// Discover the actual log file paths for pods
local.file_match "pod_logs" {
  paths = ["/var/log/pods/*/*/*.log"]
}

// Relabel to add useful metadata
discovery.relabel "pod_labels" {
  targets = discovery.kubernetes.pods.targets

  // Extract namespace, pod, and container from the file path
  rule {
    source_labels = [" __path__"]
    regex = "/var/log/pods/([^/]+)/([^_]+)_([^/]+)/(.+)/(.+)\\.log"
    target_label = "namespace"
    replacement = "$1"
  }

  rule {
    source_labels = [" __path__"]
    regex = "/var/log/pods/([^/]+)/([^_]+)_([^/]+)/(.+)/(.+)\\.log"
    target_label = "pod"
    replacement = "$2"
  }

  rule {
    source_labels = [" __path__"]
    regex = "/var/log/pods/([^/]+)/([^_]+)_([^/]+)/(.+)/(.+)\\.log"
    target_label = "container"
    replacement = "$4"
  }
}

// Read logs and forward to Loki
loki.source.file "pod_logs" {
  targets = local.file_match.pod_logs.targets
  forward_to = [loki.write.default.receiver]
}

// Loki write target
loki.write "default" {
  endpoint {
    url = "https://<your-loki-endpoint>/loki/api/v1/push"
  }
  // Optional: authentication
  // basic_auth {
  // username = "<user>"
  // password = "<password>"
  // }
}
Enter fullscreen mode Exit fullscreen mode

Grafana Alloy (or Agent) runs as daemon sets on your k8s nodes, so each daemon set read pod log file on the node, labels it to pod, namespace and container and push it to Loki tenant over HTTPs (make sure your Loki is accessible within cluster) or if outside cluster use VPC PrivateLink to connect to Loki if running in different AWS network/account.


Grafana Alloy to Loki for logs

Multi Tenant Loki Dashboards:

For better developer experience, we’re going to push class-level logs (info, error) to “service-logs” tenant, and for HTTP logs we’re going to push in another tenant “http-logs” so devs can filter/query based on fields.

Each Loki tenant is a “logical” partition on your storage backend (we’re using S3 bucket for our logs storage), so it writes each tenant logs in its own separate chunk on S3. Each tenant is served separately on UI/frontend so engineers can debug faster and helps reduce cognitive load when scanning through logs.


Multi tenant Loki dashboards

Pushing HTTP logs to separate tenant on Loki:

Now since we understand the underneath architecture of tenants in Loki, we’ll configure our Grafana Alloy (or Agent) config to filter out HTTP logs and push it to new tenant on Loki.

#########################
# TEAM A — Only HTTP logs (org.zalando)
#########################
loki.source.file "team_a_logs" {
  targets = [{
    __path__ = "/var/log/myapp/*.log"
  }]

  pipeline {
    # Step 1: Parse JSON logs
    stage.json {
      expressions = {
        level = "level"
        logger_name = "logger_name"
        msg = "message"
      }
    }

    # Step 2: Keep only logs where logger_name == org.zalando
    stage.keep {
      expressions = {
        logger_name = "org.zalando"
      }
    }

    # Step 3: Optional — add labels for Loki queries
    stage.label {
      values = {
        level = "level"
        logger = "logger_name"
        tenant = "team-A"
      }
    }

    # Step 4: Send to Loki tenant A
    stage.output {
      forward_to = [loki.write.http_logs.receiver]
    }
  }
}

loki.write "http_logs" {
  endpoint {
    url = "https://loki.example.com/loki/api/v1/push"
  }
  tenant_id = "team-A"
}

#########################
# TEAM B — Only class-level logs (INFO, DEBUG, ERROR)
#########################
loki.source.file "team_b_logs" {
  targets = [{
    __path__ = "/var/log/myapp/*.log"
  }]

  pipeline {
    # Step 1: Parse JSON
    stage.json {
      expressions = {
        level = "level"
        logger_name = "logger_name"
        msg = "message"
      }
    }

    # Step 2: Keep only certain log levels
    stage.keep {
      expressions = {
        level = "~^(INFO|DEBUG|ERROR)$"
      }
    }

    # Step 3: Add useful labels
    stage.label {
      values = {
        level = "level"
        logger = "logger_name"
        tenant = "team-B"
      }
    }

    # Step 4: Send to Loki tenant B
    stage.output {
      forward_to = [loki.write.service_logs.receiver]
    }
  }
}

loki.write "service_logs" {
  endpoint {
    url = "https://loki.example.com/loki/api/v1/push"
  }
  tenant_id = "team-B"
}
Enter fullscreen mode Exit fullscreen mode

Tenant service_logs: Kept only logs with log level INFO, DEBUG, ERROR dropped TRACE level logs (i.e. HTTP logs).

Tenant http_logs: Kept only logs with logger_name equals to org.zalando i.e. HTTP logs, dropped other class-level logs (info, debug, error).

Distributed Tracing: OpenTelemetry

Okay, now, we’re almost about to wrap up. The last thing we need to add is OpenTelemetry in our Spring Boot app so we can enable distributed tracing throughout our request lifecycle across microservices.

The good thing is OpenTelemetry for JVM offers automatic instrumentation by running a JAR agent inside your docker container. So before running our app JAR in docker, we can download OTel agent JAR and pass as a JVM arg with our app, this will enable automatic instrumentation for tracing inside our app and will also propagate trace_id in MDC which can be correlated with logs as well.

# === Stage 1: Build the application ===
FROM maven:3.9.8-eclipse-temurin-21 AS build

# Set working directory
WORKDIR /app

# Copy source and build
COPY pom.xml .
COPY src ./src
RUN mvn clean package -DskipTests

# === Stage 2: Runtime image ===
FROM eclipse-temurin:21-jre

# Set working directory
WORKDIR /app

# Copy the Spring Boot fat JAR from the build stage
COPY --from=build /app/target/*.jar app.jar

# Download the latest OpenTelemetry Java agent
# (Alternatively, you can include a specific version in your repo for consistency)
ADD https://github.com/open-telemetry/opentelemetry-java-instrumentation/releases/latest/download/opentelemetry-javaagent.jar /app/opentelemetry-javaagent.jar

# Expose the application port
EXPOSE 8080

# Environment variables for OpenTelemetry
ENV OTEL_SERVICE_NAME=my-spring-service \
    OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4317 \
    OTEL_RESOURCE_ATTRIBUTES=deployment.environment=prod,team=backend \
    OTEL_INSTRUMENTATION_LOGBACK_MDC_ENABLE=true

# Start the app with the OTel Java Agent attached
ENTRYPOINT ["java", "-javaagent:/app/opentelemetry-javaagent.jar", "-jar", "/app/app.jar"]
Enter fullscreen mode Exit fullscreen mode

Transporting Otel Telemetry over OTLP protocol to Grafana Tempo:

The Otel agent JAR transports its telemetry data over OTLP protocol. Here we can either transport telemetry data to an Otel Collector acting as a central fascade for various tracing backends like (Tempo or Jaegar) or we can directly transport telemetry over OTLP protocol to Tempo as it supports OTLP for ingestion.


Emitting Otel telemetry over OTLP to Tempo

Testing:

Done! That’s all, now finally we’re done with setting up our Grafana stack (Alloy, Tempo, Loki), we can start capturing and emitting HTTP layer and business-logic layer logs to Grafana Loki correlated with OpenTelemetry traces on Tempo based on trace_id allowing us to achieve end-to-end distributed tracing across entire system.

Here how it looks like:

HTTP log (request):


Grafana Loki (http_logs tenant)

HTTP log (response):


Grafana Loki (http_logs tenant)

Class-level log (business-logic layer):


Grafana Loki (service_logs tenant)

Full request trace on Grafana Tempo using trace_id


Grafana Tempo

Top comments (0)