DEV Community

Cover image for Self-hosted observability for .NET: logs, traces and metrics over plain OpenTelemetry.
amin parsa
amin parsa

Posted on

Self-hosted observability for .NET: logs, traces and metrics over plain OpenTelemetry.

If you run .NET services and want logs, traces and metrics in one place, the usual choices are a hosted vendor that bills by the gigabyte or a stack of separate open-source tools you wire together yourself. Flare is my attempt at a third option: a self-hosted, MIT-licensed observability platform for .NET. Your apps send logs, traces and metrics over standard OpenTelemetry (OTLP), Flare stores them in ClickHouse, and a web dashboard lets you search, correlate and alert on them. There's no Flare agent to install and no Flare SDK in your code. If your app can export OTLP, it can already talk to Flare.

Repo: https://github.com/aminparsa18/Flare.Net

Flare's logs explorer: event volume chart above a searchable list of log events

Try it in two minutes

You need Docker. Clone the repo and start the stack:

git clone https://github.com/aminparsa18/Flare.Net
cd Flare.Net
docker compose up
Enter fullscreen mode Exit fullscreen mode

If you'd rather not clone anything, the install script pulls the published images and starts the same stack (it installs Docker first if it's missing):

curl -fsSL https://raw.githubusercontent.com/aminparsa18/Flare.Net/main/scripts/install.sh | bash
Enter fullscreen mode Exit fullscreen mode

The dashboard comes up at http://localhost:7777. Sign-in is off by default, so you land straight on the Logs page. You can turn on local accounts, Microsoft Entra ID, Active Directory/LDAP, OpenID Connect or reverse-proxy headers later from the /auth page.

Now point an app at it. With plain Microsoft.Extensions.Logging you need two OpenTelemetry packages:

dotnet add package OpenTelemetry.Extensions.Hosting
dotnet add package OpenTelemetry.Exporter.OpenTelemetryProtocol
Enter fullscreen mode Exit fullscreen mode
builder.Logging.AddOpenTelemetry(logging =>
{
    logging.IncludeFormattedMessage = true;
    logging.IncludeScopes = true;
});

builder.Services.AddOpenTelemetry().UseOtlpExporter();
Enter fullscreen mode Exit fullscreen mode

and two environment variables:

export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317
export OTEL_SERVICE_NAME=my-service
Enter fullscreen mode Exit fullscreen mode

Every ILogger call now shows up in Flare. Serilog, NLog and ZLogger work the same way through their existing OpenTelemetry sinks, and the dashboard's Data sources page has copy-paste snippets for Python, Node.js, Java, Go, Kubernetes and Prometheus scraping.

If you use .NET Aspire

Flare has an Aspire hosting integration, so it can live in your AppHost next to everything else:

dotnet add package Flare.Hosting.Aspire
Enter fullscreen mode Exit fullscreen mode
var flare = builder.AddFlare("flare");

builder.AddProject<Projects.MyApi>("myapi")
       .WithReference(flare)
       .WaitForFlare(flare);
Enter fullscreen mode Exit fullscreen mode

In the project itself, the Flare.Aspire client package reads that reference and registers OTLP exporters for logs, traces and metrics:

builder.AddFlareOtlpExporter("flare");
Enter fullscreen mode Exit fullscreen mode

It sits next to whatever OpenTelemetry setup you already have, so the Aspire dashboard keeps working too. Flare starts and stops with the rest of your app like any other resource.

For the opposite case, where you have several unrelated projects on one machine and want a single Flare instance they all share, there's a global tool:

dotnet tool install --global Flare.Cli
flare start
Enter fullscreen mode Exit fullscreen mode

What you get

The logs explorer has full-text and structured search, a facet sidebar, live tail over WebSocket, and pattern grouping, which collapses thousands of similar lines into one template like User <*> logged in from <*>.

Traces get a waterfall and a flame graph, a service map built from span data, and search by trace structure, so you can ask for traces where service A calls service B and B then fails. Logs and traces are linked in both directions through the trace ID.

A trace waterfall with the slowest spans listed on top and the critical path highlighted

Metrics come in as regular OTLP metrics (gauges, sums, histograms and exponential histograms). You can chart them in the explorer or build custom dashboards with variables, formula panels and collapsible rows.

The metrics explorer showing p50, p90 and p99 of a histogram metric

Alert rules can fire on log counts, metric thresholds, exception counts, missing data or anomalies, and they notify Slack, Telegram, email, PagerDuty or a plain webhook. Maintenance windows mute them during planned work.

Creating a log-count alert rule with a threshold, a cooldown and a Slack webhook

There are also pages built on standard OpenTelemetry data for things people usually want next: an exceptions view, hosts from the collector's hostmetrics receiver, Kubernetes nodes and pods, message queues, and outbound calls to external APIs.

How it's put together

your apps --OTLP (gRPC :4317 / HTTP :4318)--> Flare.Ingest
    --> Redis Streams (buffer) --> ClickHouse
                                       ^
                          Flare.Api (search, aggregate, live tail, alerts)
                                       ^
                              SvelteKit dashboard
Enter fullscreen mode Exit fullscreen mode

Here are the decisions behind it that I'd want to know about before trusting a tool with my telemetry.

OTLP is the only way in

I didn't write a Serilog sink or an NLog target. Every logging library already has a maintained OpenTelemetry exporter, so Flare accepts OTLP and nothing else. That keeps the ingest side small, and it means switching away from Flare later is a config change in your app, not a code change.

Ingest writes to Redis Streams before ClickHouse

ClickHouse wants large batched inserts, so something has to hold events between the OTLP request and the flush. An in-memory buffer would lose everything in flight whenever Ingest restarts or redeploys. Redis Streams gives a durable queue with consumer groups: the flush worker reads with XREADGROUP and only acknowledges after ClickHouse has the batch, so delivery is at least once. It isn't zero-loss. The bundled Redis saves RDB snapshots (at most every 30 seconds, not an append-only log), so a Redis crash can lose events that arrived since the last snapshot, and a retry after a failed acknowledgement can produce a duplicate row. I'd rather say that here than have someone find out in production.

The ClickHouse schema follows the queries

The logs table is ordered by (ServiceName, SeverityNumber, Timestamp, TraceId), going from low to high cardinality and leading with the two filters people pick most in the dashboard. That lets ClickHouse skip most of the data for a typical "errors from checkout in the last hour" search. Every query also carries execution limits (time, rows read, result size), because a self-hosted ClickHouse has none by default and one unfiltered search shouldn't be able to take the server down.

Pattern grouping runs once, at flush time

Log templates are computed with the Drain algorithm right before each batch is written, and stored as columns next to the log line. Grouping by pattern later is a plain GROUP BY instead of re-clustering millions of rows on every query.

All of these are written up as architecture decision records in the repo (there are 73 so far), including the alternatives I rejected. If you disagree with one, that's the place to argue with me.

Where it stands

Flare is young and moving fast. The Docker images, the Aspire integration and the CLI are published and usable, and a multi-node ClickHouse setup is available for larger installs. The biggest open item on the roadmap is retention and cold storage to S3-compatible object storage.

If you try it, I'd like to hear what broke, what was missing, or what felt more complicated than it should be. Issues and discussions are open on GitHub:

https://github.com/aminparsa18/Flare.Net

Top comments (0)