<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sudeep Hazra</title>
    <description>The latest articles on DEV Community by Sudeep Hazra (@sudeephazra).</description>
    <link>https://dev.to/sudeephazra</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4036287%2F0857a768-5c60-4a38-8f60-810e25f439d1.jpg</url>
      <title>DEV Community: Sudeep Hazra</title>
      <link>https://dev.to/sudeephazra</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sudeephazra"/>
    <language>en</language>
    <item>
      <title>Your Kafka cluster is healthy. But the event is still missing</title>
      <dc:creator>Sudeep Hazra</dc:creator>
      <pubDate>Sat, 26 Sep 2026 08:42:01 +0000</pubDate>
      <link>https://dev.to/sudeephazra/your-kafka-cluster-is-healthy-but-the-event-is-still-missing-5e18</link>
      <guid>https://dev.to/sudeephazra/your-kafka-cluster-is-healthy-but-the-event-is-still-missing-5e18</guid>
      <description>&lt;p&gt;A recent &lt;a href="https://www.reddit.com/r/dataengineering/comments/1w3pocx/monitoring_tools_available/" rel="noopener noreferrer"&gt;r/dataengineering question&lt;/a&gt; described a familiar setup: many microservices publish JSON events to Kafka, each project has an expected throughput, and the business wants an alert when those events stop arriving.&lt;/p&gt;

&lt;p&gt;The natural response is to look for a Kafka monitoring tool. That is useful, but it solves only part of the problem.&lt;/p&gt;

&lt;p&gt;A Kafka dashboard can tell you that the brokers are available, partitions have leaders, replicas are in sync, and consumers are not falling behind. Every panel can be green while an upstream service has stopped producing the one event the business needs.&lt;/p&gt;

&lt;p&gt;The system is healthy. The data product is not.&lt;/p&gt;

&lt;p&gt;That distinction changes what we need to monitor.&lt;/p&gt;

&lt;h2&gt;
  
  
  Kafka health and business health are different signals
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://kafka.apache.org/43/operations/monitoring/" rel="noopener noreferrer"&gt;Apache Kafka's monitoring guidance&lt;/a&gt; covers the platform well. It recommends watching message and byte rates, request latency, fetch rates, replica state, and consumer lag. Those metrics answer important operational questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Can producers reach the cluster?&lt;/li&gt;
&lt;li&gt;Are brokers accepting requests?&lt;/li&gt;
&lt;li&gt;Are replicas keeping up?&lt;/li&gt;
&lt;li&gt;Are consumers processing records fast enough?&lt;/li&gt;
&lt;li&gt;Is one partition behaving differently from the others?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They do not answer whether the expected &lt;code&gt;invoice.created&lt;/code&gt; event arrived for a particular market during the last fifteen minutes.&lt;/p&gt;

&lt;p&gt;Consumer lag is a good example. A lag of zero sounds healthy. It can also mean that no records were produced. If the source application failed before publishing, the consumer has nothing to read and therefore nothing to lag behind.&lt;/p&gt;

&lt;p&gt;The same problem appears with throughput. An aggregate topic rate may look normal while one tenant, event type, or producer has gone silent. A busy topic can hide a very specific outage.&lt;/p&gt;

&lt;p&gt;I would separate monitoring into four layers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Kafka infrastructure
        ↓
Message flow
        ↓
Event contract
        ↓
Business expectation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each layer needs different evidence and a different owner.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 1: prove the platform can carry messages
&lt;/h2&gt;

&lt;p&gt;The first layer is normal Kafka operations. Watch broker availability, request errors, under-replicated partitions, ISR changes, disk pressure, network saturation, produce latency, fetch latency, and authentication failures.&lt;/p&gt;

&lt;p&gt;This is where JMX exporters, a managed Kafka metrics API, Prometheus, Grafana, or an observability platform fit. The product matters less than collecting the right broker and client signals.&lt;/p&gt;

&lt;p&gt;Consumer lag belongs here too, although I would treat it as a flow signal rather than a complete service-level indicator. Kafka exposes &lt;code&gt;records-lag&lt;/code&gt;, &lt;code&gt;records-lag-max&lt;/code&gt;, fetch rates, commit latency, and time between polls. Those metrics can identify a slow or stuck consumer.&lt;/p&gt;

&lt;p&gt;An alert such as this is useful:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;consumer lag is increasing
AND
records consumed per second is below the normal range
FOR ten minutes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It tells us that records exist and the consumer is not keeping up. It still says nothing about records that were never published.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 2: prove messages move through the expected path
&lt;/h2&gt;

&lt;p&gt;The next layer measures flow at boundaries. For every important producer and consumer, record a small set of counters and timestamps:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;events produced
events accepted by Kafka
events consumed
events processed successfully
events rejected or sent to a dead-letter path
timestamp of the newest event
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The labels need care. &lt;code&gt;service&lt;/code&gt;, &lt;code&gt;event_type&lt;/code&gt;, &lt;code&gt;environment&lt;/code&gt;, and perhaps &lt;code&gt;region&lt;/code&gt; are usually manageable. Raw customer IDs, transaction IDs, and message IDs are not. Putting high-cardinality business identifiers into a metrics system is an efficient way to turn an observability improvement into a cost incident.&lt;/p&gt;

&lt;p&gt;Keep per-message identifiers in logs or traces. Keep metrics dimensions bounded.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://opentelemetry.io/docs/specs/semconv/registry/attributes/messaging/" rel="noopener noreferrer"&gt;OpenTelemetry's messaging conventions&lt;/a&gt; provide common attributes for messaging systems, destinations, operations, consumer groups, and message context. The conventions are still marked as development in several areas, so I would pin the version used by the instrumentation instead of assuming the attribute names will never change.&lt;/p&gt;

&lt;p&gt;Tracing is helpful when a business operation crosses several services. A correlation or conversation ID can connect the API request, producer span, Kafka operation, consumer span, and downstream write. That gives an engineer a route through the failure instead of a collection of unrelated charts.&lt;/p&gt;

&lt;p&gt;Tracing every message may be too expensive at high volume. Sampling is reasonable for diagnostics. Counts and freshness metrics should remain complete because they drive alerts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 3: prove the event is usable
&lt;/h2&gt;

&lt;p&gt;Arrival is not success.&lt;/p&gt;

&lt;p&gt;A producer can publish malformed JSON at the expected rate. A serializer can omit a required field. A schema change can preserve valid syntax while changing the meaning of a value. The throughput chart will look excellent right up to the point where someone opens the downstream report.&lt;/p&gt;

&lt;p&gt;Validate the contract where ownership is clearest. That may be in the producer before publish, in a schema registry, or at the consumer boundary. Record at least:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;schema validation failures&lt;/li&gt;
&lt;li&gt;unsupported schema versions&lt;/li&gt;
&lt;li&gt;deserialization failures&lt;/li&gt;
&lt;li&gt;required-field failures&lt;/li&gt;
&lt;li&gt;duplicate or out-of-order events when those conditions matter&lt;/li&gt;
&lt;li&gt;dead-letter volume and age&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not collapse all of these into &lt;code&gt;processing_error_total&lt;/code&gt;. The response to a broken schema is different from the response to a timeout writing to a database.&lt;/p&gt;

&lt;p&gt;I would also resist putting the full payload into observability events. It creates a second, poorly governed copy of potentially sensitive data. Record the failure category, event type, schema version, and a safe lookup reference. Keep payload access inside the data platform's normal security boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 4: encode the business expectation
&lt;/h2&gt;

&lt;p&gt;This is the part a Kafka dashboard cannot infer.&lt;/p&gt;

&lt;p&gt;"Expected TPS" sounds like a threshold, but a single number is rarely enough. Traffic changes by hour, weekday, region, and business calendar. Some event types are continuous. Others arrive in bursts after a batch closes. A flat minimum can create noise during quiet periods and miss a partial outage during busy ones.&lt;/p&gt;

&lt;p&gt;Represent the expectation explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;event_type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;invoice.created&lt;/span&gt;
&lt;span class="na"&gt;producer&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;billing-service&lt;/span&gt;
&lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production&lt;/span&gt;
&lt;span class="na"&gt;schedule&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;weekdays 08:00-20:00 Europe/London&lt;/span&gt;
&lt;span class="na"&gt;minimum_events&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;300&lt;/span&gt;
&lt;span class="na"&gt;window&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5m&lt;/span&gt;
&lt;span class="na"&gt;maximum_silence&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;10m&lt;/span&gt;
&lt;span class="na"&gt;owner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;billing-platform&lt;/span&gt;
&lt;span class="na"&gt;severity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;high&lt;/span&gt;
&lt;span class="na"&gt;runbook&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://example.internal/runbooks/invoice-events&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is configuration, not dashboard decoration. Put it in version control. Review changes. Give each expectation an owner.&lt;/p&gt;

&lt;p&gt;The evaluation process can be simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;expected event rules
        +
observed event counters and freshness
        ↓
evaluation job
        ↓
pass, warn, fail, or no-data
        ↓
alert with owner and runbook
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;no-data&lt;/code&gt; state matters. Monitoring systems often treat a missing time series differently from a zero value. Your evaluator must decide what absence means for each rule. &lt;a href="https://prometheus.io/docs/prometheus/latest/configuration/alerting_rules/" rel="noopener noreferrer"&gt;Prometheus alerting rules&lt;/a&gt; support duration-based conditions through &lt;code&gt;for&lt;/code&gt;, which helps avoid paging on a brief gap. The business rule still has to define how long a gap is acceptable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reconcile counts instead of trusting one point
&lt;/h2&gt;

&lt;p&gt;For important flows, compare counts across boundaries over the same business window:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;producer accepted:  100,000
Kafka observed:      99,998
consumer processed:  99,990
target committed:    99,987
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those numbers do not prove which eight or thirteen records are missing, but they tell you where to investigate.&lt;/p&gt;

&lt;p&gt;Use event IDs or business keys for periodic reconciliation when exact completeness matters. Counters are operational signals. Reconciliation is evidence.&lt;/p&gt;

&lt;p&gt;The window must account for retries and late arrivals. Comparing two live counters at the current second will generate false differences because each stage observes the event at a different time. Close a window, allow a defined lateness period, then evaluate it. Streaming does not remove the need for accounting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Alert the team that can act
&lt;/h2&gt;

&lt;p&gt;The Kafka platform team owns broker capacity, replication, cluster access, and platform availability. The producer team owns whether it emits the promised event. The consumer team owns processing and downstream delivery. A data product owner may own the business expectation.&lt;/p&gt;

&lt;p&gt;Sending every alert to the Kafka team recreates the old middleware support queue with newer software.&lt;/p&gt;

&lt;p&gt;An alert should include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what expectation failed&lt;/li&gt;
&lt;li&gt;when the last valid event arrived&lt;/li&gt;
&lt;li&gt;expected and observed counts&lt;/li&gt;
&lt;li&gt;the affected producer, topic, event type, and consumer&lt;/li&gt;
&lt;li&gt;whether Kafka itself is healthy&lt;/li&gt;
&lt;li&gt;the owning team&lt;/li&gt;
&lt;li&gt;the first diagnostic query or dashboard&lt;/li&gt;
&lt;li&gt;the runbook&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;"Kafka events missing" is not enough. "&lt;code&gt;invoice.created&lt;/code&gt; from &lt;code&gt;billing-service&lt;/code&gt; has been silent for 12 minutes; brokers are healthy and other producers are active" is actionable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the important flows
&lt;/h2&gt;

&lt;p&gt;I would not begin by buying a broad data-observability platform or instrumenting every event type. Start with the ten flows whose absence causes money, compliance, or customer impact.&lt;/p&gt;

&lt;p&gt;For each one, define:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The event contract.&lt;/li&gt;
&lt;li&gt;The expected schedule and volume.&lt;/li&gt;
&lt;li&gt;The maximum acceptable silence.&lt;/li&gt;
&lt;li&gt;The producer and consumer owners.&lt;/li&gt;
&lt;li&gt;The reconciliation rule.&lt;/li&gt;
&lt;li&gt;The response when the rule fails.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then use the monitoring stack already in place. Add a new product only when the existing stack cannot express, evaluate, or route these rules without unreasonable work.&lt;/p&gt;

&lt;p&gt;The useful dashboard is not the one with the most Kafka metrics. It is the one that can distinguish a broken broker, a slow consumer, an invalid event, and a producer that stopped doing its job.&lt;/p&gt;

&lt;p&gt;Kafka health is necessary. The business still needs proof that the right event arrived.&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>kafka</category>
      <category>microservices</category>
      <category>monitoring</category>
    </item>
    <item>
      <title>Your agent does not need to write an essay to make a decision</title>
      <dc:creator>Sudeep Hazra</dc:creator>
      <pubDate>Sat, 26 Sep 2026 08:15:32 +0000</pubDate>
      <link>https://dev.to/sudeephazra/your-agent-does-not-need-to-write-an-essay-to-make-a-decision-1aen</link>
      <guid>https://dev.to/sudeephazra/your-agent-does-not-need-to-write-an-essay-to-make-a-decision-1aen</guid>
      <description>&lt;p&gt;Something interesting appeared in a recent &lt;a href="https://www.reddit.com/r/LocalLLaMA/comments/1wgcpww/biweekly_megathread_project_showcase/" rel="noopener noreferrer"&gt;LocalLLaMA project showcase&lt;/a&gt;. Several projects were exploring small models that do not generate prose. They take a state, a question, and a fixed set of choices, then return a probability for each choice.&lt;/p&gt;

&lt;p&gt;That sounds less impressive than an autonomous agent narrating its plan. It may also be exactly what many agent loops need.&lt;/p&gt;

&lt;p&gt;Consider the decisions inside a typical tool-using workflow:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which tool should handle this request?&lt;/li&gt;
&lt;li&gt;Is this command safe enough to run automatically?&lt;/li&gt;
&lt;li&gt;Did the previous step succeed?&lt;/li&gt;
&lt;li&gt;Should the workflow retry, stop, or ask a person?&lt;/li&gt;
&lt;li&gt;Which retrieved document should be examined next?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are closed-set decisions. Asking a generative model to produce a paragraph, parsing that paragraph, and hoping it stayed inside the allowed choices is a surprisingly elaborate way to select one value from an enum.&lt;/p&gt;

&lt;p&gt;My current view is simple: use generative models for open-ended reasoning and content. Use constrained decision components when the answer space is already known.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generation and decision are different jobs
&lt;/h2&gt;

&lt;p&gt;A generative model is useful when the output cannot be enumerated in advance. Writing a migration plan, diagnosing an unfamiliar failure, explaining a trade-off, or proposing a patch all need flexible output.&lt;/p&gt;

&lt;p&gt;Routing a ticket to one of twelve queues does not.&lt;/p&gt;

&lt;p&gt;We often use the same large model for both jobs because the API is convenient. The model receives a prompt such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Choose one action: retry, stop, escalate.
Return JSON only.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The application then validates the response, repairs malformed JSON, rejects invented actions, and perhaps asks the same model to try again. We added a language generator, then built a fence around its language generation.&lt;/p&gt;

&lt;p&gt;Structured output improves the interface, but it does not change the underlying job. The system still generates tokens to choose from a fixed set.&lt;/p&gt;

&lt;p&gt;Projects such as &lt;a href="https://github.com/Contrastive-LM/CLM" rel="noopener noreferrer"&gt;Contrastive Language Models&lt;/a&gt; approach the problem differently. CLM encodes the current state and candidate actions, scores their relationship, and returns a distribution over the supplied actions. The repository describes uses such as tool routing, trajectory ranking, and retrieval shortlisting.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Mapika/decider" rel="noopener noreferrer"&gt;Decider&lt;/a&gt; takes another route. It reads a state and typed questions, then produces probabilities for fixed choices, yes-or-no decisions, or described score levels in one pass. Its core promise is deliberately narrow: no free-form decoding and no output outside the choices supplied by the application.&lt;/p&gt;

&lt;p&gt;These projects are early and their benchmark claims need independent validation. The architecture is still worth paying attention to.&lt;/p&gt;

&lt;h2&gt;
  
  
  The application should own the choice set
&lt;/h2&gt;

&lt;p&gt;In a production agent, the model should not decide which actions exist. The application should.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;application state
        +
allowed actions
        ↓
decision component
        ↓
probabilities
        ↓
policy and threshold
        ↓
execute, ask, or stop
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That boundary gives us several useful properties.&lt;/p&gt;

&lt;p&gt;First, impossible actions stay impossible. If &lt;code&gt;delete_customer&lt;/code&gt; is not in the candidate set, the decision component cannot select it. This is stronger than telling a model in a prompt not to mention deletion.&lt;/p&gt;

&lt;p&gt;Second, the result is inspectable. A distribution such as this is easier to evaluate than a persuasive paragraph:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"retry"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.18&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"stop"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.07&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"escalate"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.75&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Third, the application can apply policy after inference. A high score is evidence, not authority.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;retry&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;probability&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.90&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;retry_count&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;retry&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;request_review&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model proposes. Code decides what that proposal is allowed to do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Probability is useful only when it is calibrated
&lt;/h2&gt;

&lt;p&gt;A number that looks like confidence is not automatically a trustworthy probability.&lt;/p&gt;

&lt;p&gt;If a model assigns roughly 0.8 to one hundred decisions, we would like about eighty of them to be correct. That is calibration. A model can have good top-choice accuracy and still be overconfident, which makes threshold-based automation dangerous.&lt;/p&gt;

&lt;p&gt;This is where a typed decision interface improves the engineering conversation. We can measure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;accuracy by decision type&lt;/li&gt;
&lt;li&gt;false-allow and false-deny rates&lt;/li&gt;
&lt;li&gt;calibration error&lt;/li&gt;
&lt;li&gt;coverage at a chosen threshold&lt;/li&gt;
&lt;li&gt;latency and cost per decision&lt;/li&gt;
&lt;li&gt;performance when none of the supplied choices is correct&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The last case is easy to miss. A closed set can be wrong. If the workflow offers &lt;code&gt;retry&lt;/code&gt;, &lt;code&gt;stop&lt;/code&gt;, and &lt;code&gt;escalate&lt;/code&gt;, but the correct action is &lt;code&gt;refresh_credentials&lt;/code&gt;, the model cannot repair the application's incomplete choice set.&lt;/p&gt;

&lt;p&gt;Include an &lt;code&gt;unknown&lt;/code&gt; or &lt;code&gt;none_of_the_above&lt;/code&gt; option where appropriate. More importantly, evaluate whether it is selected when the input falls outside the known cases.&lt;/p&gt;

&lt;p&gt;The threshold also depends on consequence. A 0.75 routing decision may be acceptable for choosing a knowledge-base category. It is not enough evidence to run a destructive infrastructure command.&lt;/p&gt;

&lt;h2&gt;
  
  
  Guardrails belong around the tool, not inside the prompt
&lt;/h2&gt;

&lt;p&gt;The decision model is not a security boundary.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/openai/openai-agents-python/blob/main/docs/guardrails.md" rel="noopener noreferrer"&gt;OpenAI Agents SDK guardrail documentation&lt;/a&gt; makes a useful separation: tool input guardrails run before execution, tool output guardrails run after execution, and either can reject content or halt the run. The SDK also documents important coverage limits for different tool types.&lt;/p&gt;

&lt;p&gt;The exact framework is less important than the placement. Authorization and safety checks must sit on the execution path.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model chooses candidate action
        ↓
schema validation
        ↓
authorization check
        ↓
risk policy
        ↓
human approval if required
        ↓
tool execution
        ↓
result validation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not ask the model, "Is the user allowed to delete this resource?" when the target system can answer that question from its own permission model. Do not treat a 0.99 probability as permission. Use identity, authorization, resource state, and explicit policy.&lt;/p&gt;

&lt;p&gt;A decision component can help classify risk or select a route. It cannot grant authority the caller does not have.&lt;/p&gt;

&lt;h2&gt;
  
  
  Small decisions can make large agents easier to operate
&lt;/h2&gt;

&lt;p&gt;Agent diagrams tend to focus on the large reasoning step. Production failures often happen in the small transitions around it.&lt;/p&gt;

&lt;p&gt;Should the agent retry after a timeout? Was the tool response complete? Does the evidence support the conclusion? Is the next operation reversible? Those decisions determine whether an isolated model error becomes a repeated side effect.&lt;/p&gt;

&lt;p&gt;Breaking them out creates observable control points:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;request
   ↓
route decision
   ↓
reasoning or retrieval
   ↓
action decision
   ↓
policy check
   ↓
tool
   ↓
success decision
   ↓
continue or stop
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each decision can have its own dataset, threshold, fallback, and owner. That is less magical than one agent prompt. It is also much easier to test.&lt;/p&gt;

&lt;p&gt;This direction fits the broader advice in Anthropic's &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;Building effective agents&lt;/a&gt;: start with the simplest workable pattern, prefer predefined workflows when the path is known, and add agent autonomy when flexibility justifies the added cost and risk.&lt;/p&gt;

&lt;p&gt;The interesting part is not replacing every generative call with a classifier. It is noticing which calls were never generative problems in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evaluation needs real workflow data
&lt;/h2&gt;

&lt;p&gt;A benchmark can tell us whether a model distinguishes choices in a prepared dataset. It cannot tell us whether those choices represent the messy states in our system.&lt;/p&gt;

&lt;p&gt;I would build the evaluation set from production-shaped examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;normal cases that should pass automatically&lt;/li&gt;
&lt;li&gt;ambiguous cases that should request review&lt;/li&gt;
&lt;li&gt;rare failures that previously caused incidents&lt;/li&gt;
&lt;li&gt;adversarial input that tries to steer the decision&lt;/li&gt;
&lt;li&gt;stale or incomplete context&lt;/li&gt;
&lt;li&gt;cases where the correct choice is absent&lt;/li&gt;
&lt;li&gt;repeated attempts after a tool failure&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Label them with the decision that the workflow should take, not the explanation we hope the model writes.&lt;/p&gt;

&lt;p&gt;Then compare at least three baselines:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Deterministic rules.&lt;/li&gt;
&lt;li&gt;A general generative model with structured output.&lt;/li&gt;
&lt;li&gt;A specialized or constrained decision model.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Rules may win. For example, &lt;code&gt;retry_count &amp;gt;= 2&lt;/code&gt; does not need inference. A model becomes useful when the state contains language or evidence that code cannot classify reliably with a small, stable rule set.&lt;/p&gt;

&lt;p&gt;Run shadow evaluations before granting automation. Record the model's decision and probability without executing it. Compare that output with the actual operator or workflow decision. This shows where a threshold would automate safely and where it would merely automate confidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do not create a second platform by accident
&lt;/h2&gt;

&lt;p&gt;There is a predictable failure mode here. A team sees specialized decision models, creates a separate serving stack, a feature store, a training pipeline, an evaluation service, and a new control plane before proving that one decision benefits from any of it.&lt;/p&gt;

&lt;p&gt;Start with one bounded choice that occurs frequently enough to matter.&lt;/p&gt;

&lt;p&gt;Tool routing is a reasonable candidate. Define the allowed tools, create a representative test set, compare rules and models, and measure error cost. Serve a smaller model only if it produces a material latency, cost, privacy, or accuracy benefit.&lt;/p&gt;

&lt;p&gt;For low-volume workflows, a general model with strict structured output may remain the simplest operational choice. Fewer services can be worth a little extra inference cost.&lt;/p&gt;

&lt;p&gt;The recommendation changes when decision calls dominate the loop. If an agent makes dozens of small choices per task, generation latency and token cost accumulate. A fast decision component can then reduce both, while giving the application a consistent probability interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the agent for the parts that need an agent
&lt;/h2&gt;

&lt;p&gt;Open-ended work still benefits from capable generative models. A decision model will not investigate an unfamiliar production failure, compare architecture options, or write a useful migration plan.&lt;/p&gt;

&lt;p&gt;It can decide which diagnostic to run next from an approved list. It can estimate whether the evidence supports another step. It can route uncertainty to a person before the workflow turns a weak guess into an action.&lt;/p&gt;

&lt;p&gt;That is enough.&lt;/p&gt;

&lt;p&gt;An agent does not become less intelligent because some decisions move into smaller, constrained components. The system becomes clearer about where intelligence is needed, where policy belongs, and where ordinary code is still the better tool.&lt;/p&gt;

&lt;p&gt;If the answer must be one of five choices, make the system choose one of five choices. Save the essay for the part that needs an explanation.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>Before You Add an Event Bus, Name the Failure</title>
      <dc:creator>Sudeep Hazra</dc:creator>
      <pubDate>Mon, 21 Sep 2026 11:53:00 +0000</pubDate>
      <link>https://dev.to/sudeephazra/before-you-add-an-event-bus-name-the-failure-25k</link>
      <guid>https://dev.to/sudeephazra/before-you-add-an-event-bus-name-the-failure-25k</guid>
      <description>&lt;p&gt;A &lt;a href="https://www.reddit.com/r/softwarearchitecture/comments/1ucjpmv/has_eventdriven_architecture_become_the_new/" rel="noopener noreferrer"&gt;discussion about event-driven architecture&lt;/a&gt; caught my attention because the replies disagreed on something more useful than Kafka versus a database. People were describing different problems with the same word: &lt;em&gt;event&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;One team needs to run a slow task after a transaction. Another needs to tell several independent systems that a business fact has changed. A third needs an audit history it can replay. All three may produce a message, but they have different requirements for ownership, ordering, and recovery.&lt;/p&gt;

&lt;p&gt;I would start by naming the failure that the current design cannot handle. If that failure is a web request waiting on slow work, a database-backed job may be enough. If independent consumers need a durable record of a business change, publishing an event may be worth the extra machinery. That gives the team a reason for the extra component.&lt;/p&gt;

&lt;h2&gt;
  
  
  A slow operation is not automatically an event
&lt;/h2&gt;

&lt;p&gt;Imagine an application that accepts an order and generates a PDF receipt. The user needs to know whether the order was accepted. They do not need to hold the connection open while the PDF renderer runs.&lt;/p&gt;

&lt;p&gt;The first design I would test is simple:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fppda9h5ozscifqa9eg4k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fppda9h5ozscifqa9eg4k.png" alt="order_processing_workflow" width="799" height="245"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The order and job are written in one database transaction. If the transaction rolls back, neither exists. If the renderer fails, the job remains visible for a retry or investigation. The application can show the user an order state and a separate receipt state. There is no broker to operate merely because work happens later.&lt;/p&gt;

&lt;p&gt;That job still needs engineering. Give it a stable ID, an attempt count, a timeout, and a clear terminal failure state. Decide how a worker claims a job and how another worker takes over after a crash. PostgreSQL documents &lt;a href="https://www.postgresql.org/docs/current/sql-select.html" rel="noopener noreferrer"&gt;&lt;code&gt;SKIP LOCKED&lt;/code&gt;&lt;/a&gt; as useful for avoiding contention among consumers of a queue-like table, while warning that it gives an inconsistent view of rows. That is a tool for workers, not a way to make arbitrary reporting queries faster.&lt;/p&gt;

&lt;p&gt;A database queue has limits. Heavy fan-out, many unrelated consumers, a large backlog, or retention far beyond the operational life of a job can make the table and its cleanup awkward. Those are measurable reasons to consider messaging infrastructure. “We want async” is not yet one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Publish facts when other owners need them
&lt;/h2&gt;

&lt;p&gt;Now add a finance service, a customer notification service, and a fulfillment system. Each has its own release schedule and its own idea of what to do when an order is paid. The order service should not need to call all three during the payment request and wait for their health before it can record payment.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;OrderPaid&lt;/code&gt; can be a useful event here. It says something that has happened, with an order ID, event ID, and version of the contract. The consumers decide what that fact means for their own systems. This is a better reason for an event than the mere existence of multiple functions in one application.&lt;/p&gt;

&lt;p&gt;The awkward part is the boundary between the order database and the broker. This code is unsafe:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3rqb6bw7rrqrc5qbjqvt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3rqb6bw7rrqrc5qbjqvt.png" alt="process_order_payment" width="684" height="471"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The process can die between those lines. Reversing the lines creates the opposite problem: a consumer may act on an event for a payment that never committed. &lt;a href="https://docs.aws.amazon.com/prescriptive-guidance/latest/cloud-design-patterns/transactional-outbox.html" rel="noopener noreferrer"&gt;AWS describes the transactional outbox pattern&lt;/a&gt; for this dual-write problem. Store the business update and an outbox record in the same transaction; a separate publisher sends committed records to the broker.&lt;/p&gt;

&lt;p&gt;The outbox improves the handoff. It does not make the rest of the system atomic. The publisher can send an event and crash before marking it delivered. The consumer must cope with that event arriving again. A stable event ID and a consumer-side record of processed IDs are more useful than a diagram that promises “exactly once” without saying where the guarantee ends.&lt;/p&gt;

&lt;p&gt;For example, a notification consumer can store &lt;code&gt;OrderPaid:event-123&lt;/code&gt; before or with the state change that schedules an email. If it sees the same event again, it can skip the duplicate scheduling step. A payment consumer has a higher bar: it must establish whether an external charge already succeeded before retrying. Its idempotency boundary is the payment provider and the local record together, not only the broker.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decide what the user must know now
&lt;/h2&gt;

&lt;p&gt;Asynchrony changes the product contract. Suppose checkout returns “payment complete” immediately, but fulfillment might reject the order ten minutes later because stock ran out. That may be a valid business process, but the UI and support team need to know what “complete” meant.&lt;/p&gt;

&lt;p&gt;For each step, I would write down three states:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Example answer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What is committed before the response?&lt;/td&gt;
&lt;td&gt;Payment authorization and order record.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What can finish later?&lt;/td&gt;
&lt;td&gt;Receipt generation and loyalty points.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What happens if later work fails?&lt;/td&gt;
&lt;td&gt;Retry, visible pending state, then support action.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If the next operation must succeed before the user can proceed, a synchronous API call may be easier to reason about. An event can still be emitted afterward for observers. If the user can wait, an accepted response plus a status endpoint can make the delay explicit. &lt;a href="https://docs.aws.amazon.com/prescriptive-guidance/latest/modernization-integrating-microservices/asynchronous.html" rel="noopener noreferrer"&gt;AWS's asynchronous communication guidance&lt;/a&gt; describes that claim-check approach: acknowledge the request, return an identifier, and let the client retrieve the later result.&lt;/p&gt;

&lt;p&gt;The delay has a user-facing cost: the text on a screen, the retry a user may press, and the question a support engineer has to answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Delivery guarantees do not finish the design
&lt;/h2&gt;

&lt;p&gt;The team needs to decide what duplicates or ordering would do to the business operation before selecting a broker. A standard Amazon SQS queue, for example, delivers messages &lt;a href="https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/standard-queues-at-least-once-delivery.html" rel="noopener noreferrer"&gt;at least once&lt;/a&gt;. The same message may be received again. Its standard queue type also gives &lt;a href="https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-queue-types.html" rel="noopener noreferrer"&gt;best-effort ordering&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Those properties are fine for many tasks. A consumer updating a current projection can ignore an event it has already processed. A consumer applying account balance deltas cannot casually apply the same delta twice. It may also need per-account ordering. Picking a FIFO queue can help with ordering and broker deduplication, but it does not remove the need to reason about retries across external systems and the full business workflow.&lt;/p&gt;

&lt;p&gt;I would ask the following before choosing the transport:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Is this a command for one worker, or a fact several owners may consume?&lt;/li&gt;
&lt;li&gt;What state is committed when the message is created?&lt;/li&gt;
&lt;li&gt;Can the same message be processed twice without damage?&lt;/li&gt;
&lt;li&gt;Must messages be ordered globally, per entity, or not at all?&lt;/li&gt;
&lt;li&gt;How long must a consumer be able to recover after an outage?&lt;/li&gt;
&lt;li&gt;Who notices a stuck consumer, and how is the backlog cleared?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The answers narrow the choice considerably. A one-owner receipt job does not need the same infrastructure as a company-wide change stream. A durable business event deserves more thought than a callback disguised as a topic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Operate the failure path before scaling the happy path
&lt;/h2&gt;

&lt;p&gt;Events make producers and consumers independent in useful ways. They also make a transaction harder to follow across time. An API request may be complete while the business process is still pending in another service. A dashboard that shows only broker throughput cannot tell you whether orders are actually progressing.&lt;/p&gt;

&lt;p&gt;I would give each event a correlation ID, a producer timestamp, a contract version, and an owning service. Then I would measure the delay from publication to the consumer's business outcome, not merely the time until a message is read. A dead-letter queue is a holding area for work that failed repeatedly. It is not a recovery plan until someone owns inspection, correction, and replay.&lt;/p&gt;

&lt;p&gt;Changes to the event contract need similar care. If consumers deploy independently, a producer cannot assume everyone upgrades on the same day. Add fields in a compatible way, document their meaning, and test old consumers against new payloads. When the meaning of a fact changes, a new event version may be clearer than silently reusing an old name.&lt;/p&gt;

&lt;p&gt;This is the cost side of the architecture decision. If a system has one application and one database, a job table may give the team all the separation it needs. If several autonomous systems depend on the same business fact and must recover independently, the broker, outbox, consumer idempotency, and operational tooling earn their place.&lt;/p&gt;

&lt;p&gt;My rule is to write down the failure and the recovery path first. Once those are concrete, the decision between a worker, an API call, and an event is usually much less mysterious.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>softwareengineering</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>What Stays in Secrets Manager After Workload Identity?</title>
      <dc:creator>Sudeep Hazra</dc:creator>
      <pubDate>Mon, 21 Sep 2026 11:41:03 +0000</pubDate>
      <link>https://dev.to/sudeephazra/what-stays-in-secrets-manager-after-workload-identity-3e5a</link>
      <guid>https://dev.to/sudeephazra/what-stays-in-secrets-manager-after-workload-identity-3e5a</guid>
      <description>&lt;p&gt;I liked the question in a recent &lt;a href="https://www.reddit.com/r/devops/comments/1vzyd12/after_moving_to_workload_identity_whats_left_in/" rel="noopener noreferrer"&gt;r/devops discussion&lt;/a&gt;: after moving workloads to identity federation, how many secrets are left?&lt;/p&gt;

&lt;p&gt;The answer depends on where each credential crosses a trust boundary.&lt;/p&gt;

&lt;p&gt;GitHub Actions can exchange an OIDC token for &lt;a href="https://docs.github.com/en/actions/concepts/security/openid-connect" rel="noopener noreferrer"&gt;short lived cloud credentials&lt;/a&gt;. An application on AWS can use an IAM role instead of carrying an access key. AWS even supports &lt;a href="https://docs.aws.amazon.com/rolesanywhere/latest/userguide/introduction.html" rel="noopener noreferrer"&gt;temporary credentials for workloads outside AWS&lt;/a&gt;, provided you operate the certificate infrastructure needed to establish that trust.&lt;/p&gt;

&lt;p&gt;Those changes remove a class of long lived credentials. They do not cause a payment provider, an older database, or a webhook sender to accept your cloud identity. Some secrets remain because the other system controls its own authentication method.&lt;/p&gt;

&lt;p&gt;My recommendation is to migrate the paths that support federation, then give the remaining credentials more deliberate ownership. &lt;strong&gt;The useful measure is how much risk and work remain per credential, not the number of entries in a vault.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Draw the boundary before deleting anything
&lt;/h2&gt;

&lt;p&gt;Consider a service that writes to S3, calls an external billing API, receives signed webhooks, and connects to an older SQL Server.&lt;/p&gt;

&lt;p&gt;The S3 access key is a good candidate for removal. Give the workload an identity and a narrowly scoped role. The billing API key still exists because that API expects a key. The webhook signing secret still exists because the receiver must verify what the sender signed. The database password might remain until that particular server version, driver, and deployment can use another supported authentication path.&lt;/p&gt;

&lt;p&gt;I would draw that as four separate edges:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9rf6gdonbwve2dhtjl6m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9rf6gdonbwve2dhtjl6m.png" alt="service_auth_methods" width="756" height="659"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The last line needs investigation, not an assumption. Some SQL Server deployments can use Microsoft Entra authentication; an older on premises deployment may have different constraints. Check the exact version and connection path before promising a passwordless migration.&lt;/p&gt;

&lt;p&gt;There is another subtle point here. A workload role can authorize a service to read a secret from a manager. It does not remove the secret stored there. It improves how the service obtains that value and who may read it. The provider credential still needs scoping, rotation, revocation, and an owner.&lt;/p&gt;

&lt;p&gt;A report saying “we use identity for all workloads” tells me little about the secrets inventory. I still want to know how each destination authenticates the caller.&lt;/p&gt;

&lt;h2&gt;
  
  
  The leftovers are often harder than the keys you removed
&lt;/h2&gt;

&lt;p&gt;Cloud access keys are painful when they spread through CI settings and configuration files, but cloud platforms give us mature replacements. The remaining secrets may have awkward lifecycles.&lt;/p&gt;

&lt;p&gt;One vendor lets you create a replacement key while the old key stays valid. Another gives you one active credential, so rotation requires a coordinated cutover. A webhook endpoint may accept more than one signing secret during a transition, or it may not. An appliance may require a manual change through a browser. A shared password may have several consumers that nobody documented.&lt;/p&gt;

&lt;p&gt;These differences determine the operating procedure. A scheduled rotation button is useful only when the destination also changes and every consumer picks up the new value. &lt;a href="https://docs.aws.amazon.com/secretsmanager/latest/userguide/rotating-secrets.html" rel="noopener noreferrer"&gt;AWS Secrets Manager's rotation documentation&lt;/a&gt; is explicit about updating both the stored secret and the database or service. Changing the vault entry alone is a reliable way to schedule an outage.&lt;/p&gt;

&lt;p&gt;For each remaining credential, I would record:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Why it matters&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Who issues and revokes it?&lt;/td&gt;
&lt;td&gt;Determines the actual recovery path.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Which workload uses it?&lt;/td&gt;
&lt;td&gt;Limits access and exposes accidental sharing.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Can old and new values overlap?&lt;/td&gt;
&lt;td&gt;Determines whether rotation can avoid downtime.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;How does the workload reload it?&lt;/td&gt;
&lt;td&gt;A new vault value is useless to a process holding the old one.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What happens if it leaks?&lt;/td&gt;
&lt;td&gt;Sets urgency and containment steps.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Who can test the replacement?&lt;/td&gt;
&lt;td&gt;Makes rotation an operation rather than a calendar reminder.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That inventory is more informative than “37 secrets remaining.” Two of those 37 might be broad production credentials with no tested revocation path. They deserve attention before twenty low privilege integration tokens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make migration reversible until the new path is proven
&lt;/h2&gt;

&lt;p&gt;I would migrate one workload edge at a time. Start with a well understood cloud API call. Give the workload a role with only the permissions it needs, then run a real operation through the new identity. Check which principal appears in the audit trail. Check that the workload refreshes credentials before expiry. Run the job long enough to exercise a refresh, not merely its startup path.&lt;/p&gt;

&lt;p&gt;Then remove the old key from the application's configuration and rerun the job. Only after that should you revoke the old key. Finally, search the places where copies tend to survive: CI secrets, deployment templates, local environment files, support scripts, and scheduled jobs. The old key may be unused by the main service but still active elsewhere.&lt;/p&gt;

&lt;p&gt;For GitHub Actions, restrict which repository, branch, or environment can assume the cloud role. OIDC removes a stored cloud key, but a broad trust policy can still authorize the wrong workflow. The token exchange is only as narrow as the trust conditions and permissions on the role.&lt;/p&gt;

&lt;p&gt;I would also test failure behavior. If federation is unavailable, does the job fail with a clear error? Does someone have a documented recovery path? Secretless authentication can still have a dependency on an identity provider, metadata endpoint, certificate authority, or token service. Those dependencies deserve monitoring.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the secrets manager, but give it a smaller job
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/secretsmanager/latest/userguide/intro.html" rel="noopener noreferrer"&gt;AWS describes Secrets Manager&lt;/a&gt; as a place for database credentials, application credentials, OAuth tokens, API keys, and other secrets. After federation, that list may shrink. The service still has a useful job: controlling access to values that cannot be replaced by an identity handshake.&lt;/p&gt;

&lt;p&gt;I would organize the remaining inventory by system boundary and owner rather than by a generic &lt;code&gt;prod/secrets&lt;/code&gt; folder. A billing key should be readable only by the billing workload. A webhook signing secret should be readable only by the verifier and rotation procedure. A database credential should have permissions that match the application's actual queries, with separate credentials for migration or administration.&lt;/p&gt;

&lt;p&gt;Set rotation intervals according to the system's capability and exposure. Do not promise automatic rotation for a vendor that offers no safe API or overlapping credentials. For those cases, write a cutover runbook and rehearse it. A manual procedure that has been tested is better than an “automated” rotation that only changes one side of the connection.&lt;/p&gt;

&lt;p&gt;There is a cost decision too. If a team has a small number of static credentials, it may not need another vault product or a custom broker. Reuse the secrets service already in the platform, provided it gives the team access control, audit history, and a workable recovery path. Additional infrastructure should solve a specific remaining problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would measure
&lt;/h2&gt;

&lt;p&gt;The percentage of credentials replaced is easy to celebrate and easy to misread. I would track four operational results instead:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Which long lived cloud keys have actually been revoked?&lt;/li&gt;
&lt;li&gt;Which remaining secrets have a named owner and known consumers?&lt;/li&gt;
&lt;li&gt;Which credentials can be rotated without an outage, and when was that last tested?&lt;/li&gt;
&lt;li&gt;How quickly can a leaked credential be identified and disabled?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Those answers tell us whether the migration changed production risk. They also expose the next improvement: a vendor integration to replace, a database auth path to modernize, or a shared key to split by workload.&lt;/p&gt;

&lt;p&gt;Workload identity is worth adopting where it fits. It removes credentials that never needed to be stored in the first place. The interesting work begins after that, when the remaining list is short enough to read and inconvenient enough to deserve proper attention.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>aws</category>
      <category>devops</category>
      <category>security</category>
    </item>
    <item>
      <title>Your Orchestrator Is an Operating Model, Not a Feature List</title>
      <dc:creator>Sudeep Hazra</dc:creator>
      <pubDate>Wed, 16 Sep 2026 19:20:11 +0000</pubDate>
      <link>https://dev.to/sudeephazra/your-orchestrator-is-an-operating-model-not-a-feature-list-51mf</link>
      <guid>https://dev.to/sudeephazra/your-orchestrator-is-an-operating-model-not-a-feature-list-51mf</guid>
      <description>&lt;p&gt;Data orchestration became interesting again, although I am not sure it ever stopped being complicated.&lt;/p&gt;

&lt;p&gt;In July, &lt;a href="https://dagster.io/prefect" rel="noopener noreferrer"&gt;Prefect announced that it was acquiring Dagster Labs&lt;/a&gt;. The announcement says Dagster will keep its name, support, and open-source license, with no migration required for customers. In September, &lt;a href="https://kestra.io/blogs/release-2-0" rel="noopener noreferrer"&gt;Kestra released version 2.0&lt;/a&gt; with a rewritten engine, separate control and data planes, remote workers, and flows that agents can call as tools.&lt;/p&gt;

&lt;p&gt;Naturally, the familiar question returned: which orchestrator should we use?&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.reddit.com/r/dataengineering/comments/1uvnu24/what_orchestrator_should_you_use/" rel="noopener noreferrer"&gt;Reddit discussion&lt;/a&gt; includes almost every answer you would expect. Stay with Dagster. Use Airflow because it is established. Choose a managed service. Try something smaller for a lean team. A few vendors also arrived to recommend their own products, because some laws of distributed systems are social.&lt;/p&gt;

&lt;p&gt;I think the product comparison is starting one step too late.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An orchestrator encodes how a team owns work.&lt;/strong&gt; The syntax and user interface matter, but the bigger decision is who defines workflows, where code runs, how failures are recovered, and which boundaries the platform enforces.&lt;/p&gt;

&lt;h2&gt;
  
  
  The DAG is the easy part
&lt;/h2&gt;

&lt;p&gt;Most tools can express this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Extract
   |
   v
Transform
   |
   v
Validate
   |
   v
Publish
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A demo usually ends when those four boxes turn green. Production starts when the second box runs for five hours, the third box discovers bad data after the publish step, and someone needs to backfill Tuesday without rerunning Wednesday.&lt;/p&gt;

&lt;p&gt;That is where orchestrators differ in ways a feature matrix struggles to capture.&lt;/p&gt;

&lt;p&gt;Does the platform think in tasks, assets, events, or deployments? Does it store enough state to reconstruct what happened? Can one team update a workflow without gaining access to every credential used by the worker? Can workloads run inside the network where the data already lives? How are upgrades tested, rolled back, and supported?&lt;/p&gt;

&lt;p&gt;The answers shape team behaviour. They also determine who gets paged.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the ownership boundary
&lt;/h2&gt;

&lt;p&gt;Before comparing Airflow, Dagster, Prefect, Kestra, or a cloud-native scheduler, I would draw the boundary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Platform team owns
  - control plane
  - identity and secrets
  - worker environments
  - deployment path
  - shared monitoring

Data teams own
  - workflow definitions
  - transformation code
  - data-quality rules
  - schedules and dependencies
  - domain runbooks
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is only one model, but it makes the questions concrete. If every data team needs the platform team to install a Python dependency, self-service is mostly a logo. If every workflow author can run arbitrary code with shared production credentials, self-service has gone too far in the other direction.&lt;/p&gt;

&lt;p&gt;The orchestrator should make the desired boundary easier to enforce.&lt;/p&gt;

&lt;p&gt;Kestra 2.0 is interesting in this context because its release notes describe workers connecting to the control plane through an outbound gRPC stream. User code runs in the data plane, and workers can sit in another region, cloud, or outbound-only network. That is more than an engine detail. It supports an operating model where a central team manages orchestration while execution stays near the workload.&lt;/p&gt;

&lt;p&gt;It also comes with migration work. The 2.0 release removes or changes several constructs, requires an upgrade through 1.3.x, and changes the default behaviour for unmatched worker routing from waiting to failing. Those details belong in the decision because an orchestrator is a long-lived operational dependency, not a library we casually swap on Friday afternoon.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decide what the system is organizing
&lt;/h2&gt;

&lt;p&gt;The task-versus-asset distinction is another operating choice.&lt;/p&gt;

&lt;p&gt;A task-oriented model asks, "What should run next?" It maps naturally to jobs, scripts, APIs, and infrastructure operations. Teams with mixed workloads often find this easy to reason about.&lt;/p&gt;

&lt;p&gt;An asset-oriented model asks, "What data should exist, and what does it depend on?" That can make partitions, lineage, freshness, and backfills easier to express for data-heavy platforms.&lt;/p&gt;

&lt;p&gt;Neither model is universally better. The useful question is what engineers spend their time debugging.&lt;/p&gt;

&lt;p&gt;If incidents usually sound like "this job did not run," task state may be the natural center. If they sound like "the customer dimension is missing yesterday's partition," asset state may be more useful. If the same platform also runs infrastructure automation, ML training, and API workflows, a strongly data-specific abstraction can become awkward.&lt;/p&gt;

&lt;p&gt;Pick the model that matches the failure language of the team.&lt;/p&gt;

&lt;h2&gt;
  
  
  An acquisition is a signal, not a migration plan
&lt;/h2&gt;

&lt;p&gt;Prefect's acquisition of Dagster creates legitimate questions about long-term product direction. It does not, by itself, make a running Dagster deployment unsafe.&lt;/p&gt;

&lt;p&gt;The public announcement commits to continued support, the existing name, and the existing open-source license. That is the fact available today. Future convergence is a possibility, not a documented migration requirement.&lt;/p&gt;

&lt;p&gt;I would respond with boring engineering work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;pin the current version and test upgrades in a representative environment;&lt;/li&gt;
&lt;li&gt;document the APIs, metadata, and deployment assumptions that create lock-in;&lt;/li&gt;
&lt;li&gt;keep transformation logic outside orchestration definitions where practical;&lt;/li&gt;
&lt;li&gt;export workflow and run metadata needed for audit or migration;&lt;/li&gt;
&lt;li&gt;define the event that would trigger a reevaluation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last item is useful. "We feel nervous" is difficult to act on. "Security fixes fall outside our required window" or "the supported deployment no longer fits our network boundary" is a decision trigger.&lt;/p&gt;

&lt;p&gt;The same discipline applies to every orchestrator. Project ownership can change. Cloud pricing can change. A managed feature can move tiers. Even a stable open-source project can become difficult for a small team to operate.&lt;/p&gt;

&lt;p&gt;An exit path is part of the architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run a failure-focused proof of concept
&lt;/h2&gt;

&lt;p&gt;Happy-path evaluations flatter every product. I would test a small workflow and then deliberately make it unpleasant.&lt;/p&gt;

&lt;p&gt;Use one hourly ingestion, one daily aggregation, and one data-quality rule. Add a partitioned backfill. Run a task longer than the worker lifetime. Rotate its credential. Remove network access halfway through a run. Deploy an incompatible workflow change and roll it back.&lt;/p&gt;

&lt;p&gt;Then score the products on the work the team will actually own:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Decision area&lt;/th&gt;
&lt;th&gt;Question to answer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Deployment&lt;/td&gt;
&lt;td&gt;Can we upgrade and roll back without improvisation?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Isolation&lt;/td&gt;
&lt;td&gt;Can workloads use separate identities and dependencies?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recovery&lt;/td&gt;
&lt;td&gt;Can an operator restart the correct unit without duplicating data?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backfills&lt;/td&gt;
&lt;td&gt;Can historical work run without starving current schedules?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Observability&lt;/td&gt;
&lt;td&gt;Can we identify the failed data, code version, and owner quickly?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-service&lt;/td&gt;
&lt;td&gt;Can a team ship safely without platform tickets?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;What grows with workflow count, run count, retention, and worker size?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exit&lt;/td&gt;
&lt;td&gt;How much domain logic is trapped in the orchestrator?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two days of deliberate failure will tell you more than two weeks of clicking through feature pages.&lt;/p&gt;

&lt;h2&gt;
  
  
  My recommendation changes with the starting point
&lt;/h2&gt;

&lt;p&gt;For an organization already running Airflow reliably, I would keep it unless there is a measured problem. A newer orchestration model may look cleaner, but migration creates parallel operations, retraining, rewritten workflows, and a new failure surface. Familiarity is not glamorous. It is still an asset.&lt;/p&gt;

&lt;p&gt;For a new data platform, I would compare one task-oriented and one asset-oriented option using a real pipeline. If the team is small, include the cost of operating the control plane, not just the speed of writing the first workflow. A managed service may be cheaper than months of occasional maintenance even when its invoice is higher.&lt;/p&gt;

&lt;p&gt;For workloads spread across restricted networks or multiple clouds, I would test the control-plane and worker boundary first. Kestra 2.0 now makes an explicit architectural claim in that area. The PoC should verify identity, connectivity, failure isolation, and upgrade behaviour rather than accepting the diagram.&lt;/p&gt;

&lt;p&gt;For a team whose warehouse or data platform already provides adequate scheduling, I would start there. Another orchestrator must earn its database, deployment, alerting, and support burden.&lt;/p&gt;

&lt;p&gt;The orchestrator decision is not Airflow versus Dagster versus Prefect versus Kestra. It is centralized versus federated ownership, tasks versus assets, shared execution versus isolated workers, and product convenience versus operational control.&lt;/p&gt;

&lt;p&gt;Choose those boundaries first. The product shortlist usually becomes much smaller after that.&lt;/p&gt;

</description>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Your Data Engineering Roadmap Is Probably Too Long</title>
      <dc:creator>Sudeep Hazra</dc:creator>
      <pubDate>Wed, 16 Sep 2026 19:19:04 +0000</pubDate>
      <link>https://dev.to/sudeephazra/your-data-engineering-roadmap-is-probably-too-long-2eam</link>
      <guid>https://dev.to/sudeephazra/your-data-engineering-roadmap-is-probably-too-long-2eam</guid>
      <description>&lt;p&gt;A data engineering roadmap crossed my feed recently. It is thoughtful, detailed, and 500+ lines long. The associated Reddit discussion had a predictable reaction: useful reference, terrifying learning plan.&lt;/p&gt;

&lt;p&gt;That distinction matters.&lt;/p&gt;

&lt;p&gt;A map can show every road in a country. It does not mean you need to drive every road before you are allowed to leave home.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/ErdemOzgen/Data-Engineering-Roadmap/blob/main/DATA_ENGINEERING_ROADMAP_2026.md" rel="noopener noreferrer"&gt;Data Engineering Roadmap 2026&lt;/a&gt; covers software engineering, Python, SQL, operational databases, warehouses, lakehouse storage, orchestration, observability, security, and cost. As an inventory of the field, that is helpful. As a checklist for becoming employable, it is too easy to read it as an entrance exam that never ends.&lt;/p&gt;

&lt;p&gt;Several engineers in the &lt;a href="https://www.reddit.com/r/dataengineering/comments/1vzoiwh/2026_data_engineering_roadmap/" rel="noopener noreferrer"&gt;Reddit discussion&lt;/a&gt; made the same point in different ways. Learn SQL, Python, modelling, and the fundamentals of transformation. Add tools when the problem asks for them. The disagreement was mostly about how much product knowledge someone needs before starting.&lt;/p&gt;

&lt;p&gt;My view is simpler: &lt;strong&gt;learn one complete data system before collecting ten disconnected technologies.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Tools are the visible part of the job
&lt;/h2&gt;

&lt;p&gt;Job descriptions make data engineering look like a shopping list:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Python + SQL + Spark + Kafka + Airflow
+ dbt + Snowflake + Databricks + AWS
+ whatever was added to the platform last Tuesday
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The list is visible because product names are easy to search for. The harder skills hide underneath it.&lt;/p&gt;

&lt;p&gt;Can you tell whether a source is producing inserts, updates, or replacements? Can you make a load safe to retry? What happens when a column changes type? How do you backfill three months without corrupting today's incremental run? Who gets alerted when the pipeline finishes successfully but loads zero rows?&lt;/p&gt;

&lt;p&gt;Those questions survive tool changes.&lt;/p&gt;

&lt;p&gt;Airflow can become Dagster. Redshift can become Snowflake. Spark can become a warehouse query or a small Python process because the data was never large enough to justify a cluster. The names move around. Ordering, idempotency, schema evolution, reconciliation, and operational ownership remain.&lt;/p&gt;

&lt;p&gt;That is why a roadmap built around products creates a strange learning pattern. A person can finish six courses and still not know how to recover a failed pipeline. They have seen every component, but they have not owned a system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build one pipeline with consequences
&lt;/h2&gt;

&lt;p&gt;I would structure the learning path around one modest project:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PostgreSQL source
       |
       v
Incremental extraction
       |
       v
Object storage as Parquet
       |
       v
Warehouse tables
       |
       v
Quality checks and reconciliation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The business story can be ordinary: orders, payments, support tickets, or device events. Ordinary is good. The learning comes from making the pipeline dependable, not from inventing a futuristic use case.&lt;/p&gt;

&lt;p&gt;Start with a full load. Record the source row count, loaded row count, start time, end time, and status. Run it twice and prove that the second run does not duplicate data.&lt;/p&gt;

&lt;p&gt;Then add incremental processing. Choose a watermark and explain why it is safe. If you use &lt;code&gt;updated_at&lt;/code&gt;, decide what happens when records arrive late or a source clock is wrong. If you use a monotonically increasing ID, decide how updates are detected. There is no magic watermark, only a trade-off you can defend.&lt;/p&gt;

&lt;p&gt;Next, break the schema on purpose. Add a nullable column. Rename another one. Change a numeric field into text. The pipeline should either handle the change or fail with enough evidence for someone to diagnose it. A silent partial load does not count as resilience (it counts as tomorrow's incident).&lt;/p&gt;

&lt;p&gt;Finally, add a backfill that can run alongside the current schedule. Keep historical and incremental state separate. Reconcile totals after the backfill, then document how to restart it.&lt;/p&gt;

&lt;p&gt;This single project teaches more useful engineering than a folder of disconnected tutorials because every new decision affects something already running.&lt;/p&gt;

&lt;h2&gt;
  
  
  Learn in layers, not brands
&lt;/h2&gt;

&lt;p&gt;The order I would use is deliberately boring.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Data correctness
&lt;/h3&gt;

&lt;p&gt;Begin with SQL, data modelling, transactions, indexes, query plans, and the difference between an operational schema and an analytical one. Learn how nulls, duplicates, time zones, and changing business definitions damage results.&lt;/p&gt;

&lt;p&gt;Python matters here, but mainly as software. Use functions, tests, type hints, logging, dependency management, and a command-line entry point. A notebook is useful for exploration. It is a poor substitute for a repeatable job.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Pipeline behaviour
&lt;/h3&gt;

&lt;p&gt;Add scheduling, retries, idempotency, checkpoints, and backfills. Learn the difference between a task completing and the data being correct. Store enough run metadata to explain what happened without reading raw logs for an hour.&lt;/p&gt;

&lt;p&gt;This is also the right time to learn an orchestrator. Pick one. The goal is to understand dependencies, scheduling, state, retries, and failure recovery. You do not need equal fluency in Airflow, Dagster, Prefect, Kestra, and every managed cloud service.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Scale and distribution
&lt;/h3&gt;

&lt;p&gt;Only after the single-node version becomes limiting would I add Spark, Kafka, or a lakehouse table format. Otherwise, it is difficult to separate the complexity of the problem from the complexity of the platform.&lt;/p&gt;

&lt;p&gt;When Spark enters the project, measure why. Is the input too large for memory? Is parallelism shortening a batch window? Does the team already run the platform? "Spark appears in many job descriptions" is a career reason to understand it, but it is not an architecture reason to deploy it.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Production ownership
&lt;/h3&gt;

&lt;p&gt;Security, cost, deployment, observability, and support are not advanced electives. They are what turns a pipeline into a service.&lt;/p&gt;

&lt;p&gt;Use workload identity instead of static credentials. Set a budget alert. Define freshness and correctness checks. Write a runbook for the two failures most likely to happen. Package the project so another engineer can run it.&lt;/p&gt;

&lt;p&gt;That final step changes the portfolio from "I followed a tutorial" to "I can own a data workload."&lt;/p&gt;

&lt;h2&gt;
  
  
  Breadth still matters, just later
&lt;/h2&gt;

&lt;p&gt;There is a reasonable counterargument. Engineers often join environments with a fixed stack, and hiring filters still look for product names. A consultant may need to compare several platforms quickly. Breadth has career value.&lt;/p&gt;

&lt;p&gt;The mistake is treating breadth and depth as competing destinations. They are a sequence.&lt;/p&gt;

&lt;p&gt;Build depth with one stack. Then map nearby tools onto concepts you already understand:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concept&lt;/th&gt;
&lt;th&gt;First implementation&lt;/th&gt;
&lt;th&gt;What to compare later&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Orchestration&lt;/td&gt;
&lt;td&gt;One scheduler&lt;/td&gt;
&lt;td&gt;State model, retries, deployment, access control&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage&lt;/td&gt;
&lt;td&gt;One warehouse and object store&lt;/td&gt;
&lt;td&gt;Cost, isolation, schema evolution, query patterns&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transformation&lt;/td&gt;
&lt;td&gt;SQL and one processing engine&lt;/td&gt;
&lt;td&gt;Pushdown, testing, lineage, scale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ingestion&lt;/td&gt;
&lt;td&gt;Batch first&lt;/td&gt;
&lt;td&gt;CDC semantics, ordering, replay, source impact&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operations&lt;/td&gt;
&lt;td&gt;Logs and run metadata&lt;/td&gt;
&lt;td&gt;Metrics, tracing, alert routing, SLOs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Now a second tool is not another syllabus. It is a comparison against known constraints.&lt;/p&gt;

&lt;h2&gt;
  
  
  A roadmap should produce decisions
&lt;/h2&gt;

&lt;p&gt;The best learning project leaves behind evidence of judgment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;why the pipeline is batch rather than streaming;&lt;/li&gt;
&lt;li&gt;how retries avoid duplicates;&lt;/li&gt;
&lt;li&gt;what happens when the schema changes;&lt;/li&gt;
&lt;li&gt;how historical loads differ from current ingestion;&lt;/li&gt;
&lt;li&gt;which metric tells you the data is late;&lt;/li&gt;
&lt;li&gt;when the current design would stop being enough.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An interviewer can discuss those decisions. A future teammate can review them. More importantly, you can reuse the reasoning when the product names change.&lt;/p&gt;

&lt;p&gt;The long roadmap is still useful. Keep it as an atlas. Use it to notice gaps and choose the next area to explore.&lt;/p&gt;

&lt;p&gt;For the actual journey, pick one source, one destination, and one pipeline with consequences. Make it correct. Make it restartable. Make it understandable by someone else.&lt;/p&gt;

&lt;p&gt;Then add the next tool because the system needs it, not because the roadmap had another box.&lt;/p&gt;

</description>
      <category>dataengineering</category>
    </item>
    <item>
      <title>AI Security Scanning Needs Evidence, Not Just More Agents</title>
      <dc:creator>Sudeep Hazra</dc:creator>
      <pubDate>Sun, 13 Sep 2026 08:47:20 +0000</pubDate>
      <link>https://dev.to/sudeephazra/ai-security-scanning-needs-evidence-not-just-more-agents-2mlb</link>
      <guid>https://dev.to/sudeephazra/ai-security-scanning-needs-evidence-not-just-more-agents-2mlb</guid>
      <description>&lt;p&gt;Google’s Mantis caught my attention because it points to a problem most AI security demos quietly walk around: &lt;code&gt;finding a vulnerability is not the same as proving one exists.&lt;/code&gt;   &lt;/p&gt;

&lt;p&gt;That distinction matters. Security teams already live with noisy scanners, half-useful alerts, and findings that require someone experienced to separate a real exploit path from a theoretical complaint. Adding an LLM can improve that workflow. It can also make the scanner more fluent while being wrong in more elaborate ways.&lt;/p&gt;

&lt;p&gt;That is not progress. That is just a better-written interruption.&lt;/p&gt;

&lt;p&gt;The interesting part of Mantis is not the label “agentic vulnerability scanning.” Everyone is attaching agentic to things now. Apparently software is not allowed to have a normal workflow anymore. The useful part is the structure around grounding: repository context, history, threat models, reviewer stages, critic stages, sandboxed reproduction, and patches that are tied back to evidence.&lt;/p&gt;

&lt;p&gt;That shape makes sense because vulnerability detection is not one job. It is several jobs that often get collapsed into one vague “scan the code” box.&lt;/p&gt;

&lt;p&gt;The reproduction step is the part I would care about most.&lt;/p&gt;

&lt;p&gt;If a system can show a working crash, a failing test, an exploit path, or a concrete data-flow issue, the review conversation changes. The security team is no longer reading model confidence. They are reading evidence. That does not remove human judgment, but it gives the human something useful to judge.&lt;/p&gt;

&lt;p&gt;Most teams do not suffer because they have too few alerts. They suffer because the alerts arrive without enough context. Someone has to open the repository, understand the service boundary, trace input handling, inspect validation, check the framework behavior, look at previous fixes, decide whether the path is reachable, and then argue with the scanner’s output. That is expensive work.&lt;/p&gt;

&lt;p&gt;AI can help with that work, but only if the system is designed as a workflow rather than a magic box.&lt;/p&gt;

&lt;p&gt;For example, a scanner that says:&lt;/p&gt;

&lt;p&gt;Possible SQL injection in UserController.java&lt;br&gt;
is not enough. A better system should be able to say:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;This route accepts user input here.
The value reaches this query builder here.
This sanitizer does not cover this pattern.
This is a minimal reproducer.
This is the failing test.
This is the proposed fix.
This is the residual uncertainty.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line matters. Residual uncertainty is not weakness. It is honesty. A security tool that pretends every finding is equally certain creates bad incentives. Engineers start ignoring it, security teams start tuning it down, and eventually the tool becomes background noise.&lt;/p&gt;

&lt;p&gt;The useful promise of an agentic scanner is that it can break the work into stages and let each stage challenge the previous one. One agent may identify suspicious flows, whereas another may criticize the finding, and another may try to reproduce it. Another may propose a patch. A final stage may check whether the patch changes behavior outside the intended area.&lt;/p&gt;

&lt;p&gt;That is much closer to how a careful human review works.&lt;/p&gt;

&lt;p&gt;It also changes how I would think about model selection. Not every stage needs the biggest model available. Classification, deduplication, clustering similar findings, and summarizing repository structure are not the same as tracing a subtle authorization bypass through five layers of application code. A practical system should spend reasoning budget where reasoning is actually needed.&lt;/p&gt;

&lt;p&gt;That is a very SRE-ish way to think about AI, and I mean that as a compliment. The model is not magic. It is one component in a workflow with cost, latency, permissions, state, logs, and failure modes.&lt;/p&gt;

&lt;p&gt;The cost piece is easy to ignore in a research announcement and painful to ignore in a real engineering organization. If every pull request triggers a deep multi-agent investigation across a large monorepo, the bill will get interesting very quickly. The system needs triage. It needs cheap filters, expensive analysis only when justified, and clear rules for what runs synchronously in CI versus what runs asynchronously in a security pipeline.&lt;/p&gt;

&lt;p&gt;There is also a timing question. Some checks belong directly in the developer loop. They should run fast, fail clearly, and produce a result while the developer still remembers what they changed. Other checks are better as background analysis. A deep investigation across service boundaries may be valuable, but it probably should not block every commit unless the organization is prepared for that operational cost.&lt;/p&gt;

&lt;p&gt;This is where the design should separate developer feedback from security investigation.&lt;/p&gt;

&lt;p&gt;Developer feedback needs speed and clarity. Security investigation needs depth and evidence. Trying to make one workflow satisfy both usually produces a system that is too slow for developers and too shallow for security teams. That is a familiar failure mode. We have seen it with static analysis, dependency scanning, data-quality checks, and policy-as-code.&lt;/p&gt;

&lt;p&gt;AI does not remove that trade-off. It just makes it easier to hide for a while.&lt;/p&gt;

&lt;p&gt;I would probably separate the workflow like this:&lt;/p&gt;

&lt;p&gt;This keeps the expensive reasoning stages closer to the findings that deserve them.&lt;/p&gt;

&lt;p&gt;There is still plenty of operational work hiding behind the nice diagram. The scanner needs sandboxing. It needs restricted network access. It needs deterministic handoffs between stages. It needs audit logs. It needs a way to prevent an aggressive false-positive filter from suppressing weak-looking findings that are actually real. It also needs ownership, because once a tool starts filing security bugs or proposing patches, someone has to decide what “good enough” means.&lt;/p&gt;

&lt;p&gt;This is where security automation often becomes uncomfortable. A tool that only reports findings is easy to ignore. A tool that opens patches is harder to ignore, but also more dangerous. The patch might fix the immediate issue while changing behavior somewhere else. It might silence a test instead of fixing the cause. It might introduce a different vulnerability. It might be correct technically but wrong for the product’s authorization model.&lt;/p&gt;

&lt;p&gt;Authorization bugs are a good example. They rarely live in one obvious line of code. The check may depend on route configuration, middleware behavior, tenant context, cached permissions, database filters, and assumptions in the UI. A scanner that sees only the controller method may miss the real boundary. A model with repository context may do better, but only if the system gives it the right evidence and then forces it to prove the path.&lt;/p&gt;

&lt;p&gt;The same is true for deserialization, SSRF, file access, and dependency confusion. The interesting question is often not “does this function look suspicious?” It is “can untrusted input actually reach this dangerous capability under realistic conditions?” That requires reachability analysis, test construction, environment modeling, and sometimes domain knowledge about how the application is deployed.&lt;/p&gt;

&lt;p&gt;This is where agents can be useful, but also where they can become overconfident. A fluent explanation of a possible exploit is not the same as exploitability. I would rather have a scanner that says “I cannot prove this yet” than one that files a confident but ungrounded critical finding.&lt;/p&gt;

&lt;p&gt;So I would not give an AI scanner unlimited write access to a production codebase. I would start with evidence generation, then move to suggested patches behind review, then allow more automation only for narrow, well-tested classes of changes.&lt;/p&gt;

&lt;p&gt;Something like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Phase 1: identify and explain&lt;/li&gt;
&lt;li&gt;Phase 2: reproduce with tests&lt;/li&gt;
&lt;li&gt;Phase 3: suggest patches&lt;/li&gt;
&lt;li&gt;Phase 4: auto-fix low-risk patterns&lt;/li&gt;
&lt;li&gt;Phase 5: measure escaped defects and false negatives&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important part is that automation earns trust through measured behavior. Not vibes. Not a launch post. Measured behavior.&lt;/p&gt;

&lt;p&gt;The metrics should reflect that. I would track more than “number of vulnerabilities found.” That metric is too easy to game. A useful program would track confirmed true positives, false-positive rate by category, time from finding to reproduction, time from reproduction to patch, developer review burden, escaped vulnerabilities, and whether suggested fixes survived regression testing.&lt;/p&gt;

&lt;p&gt;Those numbers tell you whether the system is improving the security process or merely producing activity.&lt;/p&gt;

&lt;p&gt;There is also a governance concern. If the scanner learns from internal code, writes summaries to disk, executes reproducers in sandboxes, and uses multiple models, the organization needs to know where sensitive data goes. Security tooling often has broad repository access by design. That makes data handling, retention, model-provider boundaries, and access logs part of the architecture, not an appendix.&lt;/p&gt;

&lt;p&gt;I would treat an AI security scanner almost like a privileged internal service:&lt;/p&gt;

&lt;p&gt;That is more work than running a command-line scanner. But if the system is going to reason over sensitive code and propose security patches, the extra discipline is not optional.&lt;/p&gt;

&lt;p&gt;The bigger issue is that AI security tooling can easily become another alert generator. Teams do not need more findings. They need better evidence, better prioritization, and a shorter path from suspicion to validated fix.&lt;/p&gt;

&lt;p&gt;That is where agentic workflows may actually help. Not because an agent can read a lot of files, although that helps. Not because it can write a patch, although that is useful. The real value is in connecting the steps that human reviewers already perform manually:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;understand context&lt;/li&gt;
&lt;li&gt;form a hypothesis&lt;/li&gt;
&lt;li&gt;challenge it&lt;/li&gt;
&lt;li&gt;reproduce it&lt;/li&gt;
&lt;li&gt;and only then act&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is also a cultural side to this. Developers will trust the system faster if it explains itself in the language of the codebase. Security teams will trust it sooner if there is reproducible evidence. Engineering leaders will trust it faster if it reduces mean time to validated fix without flooding teams with noise. Those are different success metrics, and a serious platform needs to satisfy all three.&lt;/p&gt;

&lt;p&gt;My current view is simple: &lt;code&gt;AI belongs in security scanning when it is treated as an evidence-generation system.&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The agent can suggest. The harness has to prove. That is enough for the first serious version.&lt;/p&gt;

&lt;p&gt;Only after that, let us talk about autonomy.&lt;/p&gt;

&lt;p&gt;| References:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Inspired by InfoQ’s coverage of Google Mantis: &lt;a href="https://www.infoq.com/news/2026/09/google-mantis-vulnerability-scan/" rel="noopener noreferrer"&gt;https://www.infoq.com/news/2026/09/google-mantis-vulnerability-scan/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Related trend context: InfoQ’s September 2026 AI, ML, and data engineering coverage highlighted agentic testing, security, and production-readiness themes.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>security</category>
    </item>
  </channel>
</rss>
