<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: praveenlavu</title>
    <description>The latest articles on DEV Community by praveenlavu (@praveenlavu).</description>
    <link>https://dev.to/praveenlavu</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3991566%2Ff7152a58-11e0-4256-b1d8-a564907bd5a1.png</url>
      <title>DEV Community: praveenlavu</title>
      <link>https://dev.to/praveenlavu</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/praveenlavu"/>
    <language>en</language>
    <item>
      <title>Duplicate Claims and the Savepoint Pattern</title>
      <dc:creator>praveenlavu</dc:creator>
      <pubDate>Mon, 14 Sep 2026 14:00:07 +0000</pubDate>
      <link>https://dev.to/praveenlavu/duplicate-claims-and-the-savepoint-pattern-3dl8</link>
      <guid>https://dev.to/praveenlavu/duplicate-claims-and-the-savepoint-pattern-3dl8</guid>
      <description>&lt;h1&gt;
  
  
  When the Claim Arrives Twice
&lt;/h1&gt;

&lt;p&gt;The first time it happened in production, I stared at the logs for twenty minutes before I understood what I was looking at.&lt;/p&gt;

&lt;p&gt;Two identical claims. Same member, same date of service, same procedure code, same everything. Processed twice. Not because anyone submitted the form twice on purpose. Because a network retry fired, the receiving system acknowledged too slowly, and the queue decided "maybe it didn't work" and tried again. Both went through. Both wrote records. Both triggered downstream workflows. The system was perfectly correct about each one individually, and catastrophically wrong in aggregate.&lt;/p&gt;

&lt;p&gt;That moment is the reason I think about idempotency differently now than I did early in my career. Back then, it felt like a theoretical computer science concept. Now it keeps me from being paged at 2am. And it forced me to ask a question I did not expect to be so hard: how do you make a system safe when the same input can arrive more than once?&lt;/p&gt;

&lt;h2&gt;
  
  
  The Deceptive Simplicity of "Just Check First"
&lt;/h2&gt;

&lt;p&gt;The obvious answer is: before you process a claim, check whether it already exists. If it does, skip it. Simple.&lt;/p&gt;

&lt;p&gt;Except it isn't.&lt;/p&gt;

&lt;p&gt;The check and the write are two separate operations. Between the moment you check and the moment you write, another process can sneak in. In a busy production environment, that gap is milliseconds, which sounds like nothing right up until you have a queue consuming messages with eight concurrent workers. Now the same claim can pass the "does this exist" check in two workers simultaneously, both conclude "no, it's new," and both proceed to write. You have not solved the problem. You have given it a race condition.&lt;/p&gt;

&lt;p&gt;You can tighten the gap with locks. A pessimistic lock before the check prevents concurrent reads, which prevents the race. But now you have serialized your entire processing pipeline on a single bottleneck, and your throughput drops as fast as your retry budget climbs. The cure starts to feel like a disease.&lt;/p&gt;

&lt;p&gt;There is another failure mode that does not get talked about enough: partial processing. A claim does not land in one atomic moment. It triggers a chain: validation, benefit lookup, adjudication, writing the determination, notifying downstream systems. If a failure happens at step four, you have three steps' worth of state in your database and a half-processed claim that will come back around and fail your "does this exist" check in confusing ways. The system is neither fully processed nor safely reprocessable. It is stuck.&lt;/p&gt;

&lt;p&gt;The naive approach assumes the failure modes are simple and linear. Production systems are neither.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting the Savepoint Before You Act
&lt;/h2&gt;

&lt;p&gt;The insight that changed my thinking came from asking a different question. Instead of "how do I know if I've already done this?" I started asking "how do I make it safe to do this twice?"&lt;/p&gt;

&lt;p&gt;That shift matters. The first question tries to prevent duplication by detecting it early. The second accepts that duplication will happen and designs the processing to survive it gracefully.&lt;/p&gt;

&lt;p&gt;The pattern begins at ingest, before any processing logic runs. Every claim arrives carrying its complete payload. Hashing that full payload through SHA-256 produces the same 64-character string every time the byte-identical message arrives, regardless of which system sent it, which retry attempt it represents, or when it comes through. The hash is not derived from a selection of identifying fields. It is derived from the complete, unmodified byte content of the payload as received. Any two transmissions that differ by even a single character produce different hashes and are treated as distinct claims. Any two transmissions that are byte-identical produce the same hash and are recognized as the same unit of work.&lt;/p&gt;

&lt;p&gt;That ingest-time identity check is the foundation. With it in place, the processing sequence becomes: set a savepoint in the active transaction, then check whether this content hash already appears in the staging layer.&lt;/p&gt;

&lt;p&gt;If the hash is new, the system writes the staging records fresh and proceeds. This is the common path.&lt;/p&gt;

&lt;p&gt;If the hash already exists, the system does something that initially seems counterintuitive. It cleans up the existing staging records associated with that hash, then rewrites them from scratch using the payload it just received. Cleanup first, then rewrite, all within the boundary of the savepoint. The reason is specific: existing staging records for a given hash may be partial. A previous arrival of this claim may have failed partway through the staging phase, leaving some records written and others not. Returning that partial state downstream would propagate the error. Cleaning it up and rewriting it from the byte-identical payload in hand produces one complete, consistent staging state that downstream processing can act on reliably.&lt;/p&gt;

&lt;p&gt;The savepoint is what makes this recoverable. If anything fails during cleanup or rewrite, the transaction rolls back to the savepoint marker. The database returns to its pre-attempt state. The claim goes back to the queue. The next arrival runs through the same sequence and starts from clean ground.&lt;/p&gt;

&lt;p&gt;The concurrent-duplicate race closes at the database level. Two workers holding the same content hash, both checking at their respective savepoints and finding no existing record, will both proceed to write staging records. They will collide at the unique constraint on the hash column. One write wins and commits. One fails with a constraint violation. The losing worker rolls back to its savepoint and follows the same duplicate path: cleanup, then rewrite from the payload it holds. The caller gets one result. The database holds one consistent set of staging records. Downstream processing runs once.&lt;/p&gt;

&lt;p&gt;What enforces the invariant is not an application-level lock. It is the unique constraint on the hash column, applied atomically by the database itself, paired with the savepoint that contains each worker's recovery within its own transaction boundary. Neither mechanism is sufficient alone. The constraint without the savepoint gives you collision detection with no clean recovery path. The savepoint without the constraint gives you contained transactions with no atomicity guarantee between concurrent writers. Together they make the staging layer idempotent by construction: any number of arrivals of a byte-identical payload will produce exactly one consistent set of staging records, regardless of timing or order.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Second Arrival Is Not a Problem
&lt;/h2&gt;

&lt;p&gt;The thing nobody tells you about idempotency is that it changes how you operate, not just how you build.&lt;/p&gt;

&lt;p&gt;When you know a claim can arrive ten times and produce the same staging state every time, your relationship with retries changes entirely. You stop being afraid of them. You can tune retry policies aggressively because you know they cannot corrupt state. You can add observability tooling that replays events without worrying about side effects. You can restore from a backup, replay the queue, and walk away knowing the staging records will reflect exactly one consistent attempt per unique payload.&lt;/p&gt;

&lt;p&gt;The oncall conversation changes too. An alert about a claim that "ran twice" becomes bounded. You check the content hash. You verify the staging records are in a consistent state. You confirm exactly one downstream workflow fired. You are not reconstructing damage. You are confirming what the system already resolved.&lt;/p&gt;

&lt;p&gt;Go back to where this started. Two identical claims, both processed, both writing records, both triggering workflows. The system was correct about each one individually and catastrophically wrong in aggregate. What was missing was not detection, not better alerting, not more careful operators. It was a structural guarantee at the transaction level: compute the hash at ingest from the complete byte-identical payload, set the savepoint, clean up and rewrite the staging records within its boundary, enforce uniqueness atomically at insert time. The second arrival runs through exactly the same sequence. The staging layer ends in exactly the same state. Downstream processing sees one result, because there is only one result to see.&lt;/p&gt;

&lt;p&gt;The claim can arrive twice. When the pattern holds, the second arrival is indistinguishable from a no-op. That is not a hope or a best-effort promise. It is what the savepoint, the unique constraint, and the cleanup-then-rewrite commit to together, on every arrival, every time.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>production</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Why Your Pipeline Redeploys Unchanged Code</title>
      <dc:creator>praveenlavu</dc:creator>
      <pubDate>Sun, 13 Sep 2026 14:08:36 +0000</pubDate>
      <link>https://dev.to/praveenlavu/why-your-pipeline-redeploys-unchanged-code-3n2o</link>
      <guid>https://dev.to/praveenlavu/why-your-pipeline-redeploys-unchanged-code-3n2o</guid>
      <description>&lt;h1&gt;
  
  
  The Deploy That Didn't Need to Happen
&lt;/h1&gt;

&lt;p&gt;There is a particular kind of waste that looks like productivity. A deployment pipeline spins up, pulls the code, runs the checks, packages the artifact, pushes to the target, and reports success. Minutes of wall-clock time, gone. The logs are clean. The status badge is green. And if you look closely at what just happened, the thing that got deployed is identical to what was already running. Not close. Identical. Same commit, same everything.&lt;/p&gt;

&lt;p&gt;The pipeline reported success. Nothing changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline knows when it runs
&lt;/h2&gt;

&lt;p&gt;The reason it happens is simple once you see it. The pipeline knows when it runs. It does not know what is already deployed. Those are two different questions, and most workflows only ask the first one. Did something trigger me? Yes. Then deploy.&lt;/p&gt;

&lt;p&gt;The smarter question is: has anything actually changed since the last time this succeeded? That question has a real answer. The platform already knows it. The deployment history is sitting there, queryable, with commit identifiers attached to every run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Asking in the wrong place
&lt;/h2&gt;

&lt;p&gt;The naive fix is to check timestamps. If a file was touched recently, deploy. If not, skip. That breaks within a week. Timestamps lie. Restores lie. Anything that touches the filesystem without changing the code triggers a deploy that should not happen, and anything that changes the code without bumping a timestamp skips one that should.&lt;/p&gt;

&lt;p&gt;I tried branch comparisons next. If the branch is ahead of the last tag, deploy. That one holds longer, but tags are manual. The first time someone forgets to tag, the check silently fails open and you are back to deploying nothing for the cost of deploying something.&lt;/p&gt;

&lt;p&gt;Both of those are workarounds for a question I was asking in the wrong place.&lt;/p&gt;

&lt;h2&gt;
  
  
  The answer was in the run history
&lt;/h2&gt;

&lt;p&gt;The answer was not in the code at all. It was in the run history.&lt;/p&gt;

&lt;p&gt;Every successful deployment leaves a record: when it ran, what triggered it, and which exact commit it built from. That record is an API call away. Query the last successful run on the target environment. Read the commit identifier it deployed. Compare it to the one you are about to deploy. If they match, you are done. Not "skip for now and hope it is fine." Done, with proof, logged, traceable, auditable.&lt;/p&gt;

&lt;p&gt;The whole decision fits before the expensive steps start. Check first, then commit to the work.&lt;/p&gt;

&lt;p&gt;But here is what I did not expect. The first time the skip condition fired on a live pipeline, I did not believe it. The run finished in seconds. Nothing built, nothing pushed, no artifact in the queue. I checked the logs twice. Then a third time. The commit identifier in the run history matched exactly. The platform had known this the entire time. Not because of anything I had put there. It had always known. I had just never asked.&lt;/p&gt;

&lt;p&gt;I sat with that for a moment. The system was not broken. I had been interrogating it with the wrong question for months, and every green run it returned was an honest answer to a question I should not have been asking.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed was not just speed
&lt;/h2&gt;

&lt;p&gt;What changed was not just speed. It was honesty.&lt;/p&gt;

&lt;p&gt;The pipeline stopped lying to me about what it had accomplished. Before, a successful run meant "we ran the steps." After, a successful run means "we changed something, or we confirmed nothing needed changing." Those are different things, and I had been conflating them for years without noticing.&lt;/p&gt;

&lt;p&gt;The confidence the fingerprint check gives is real, but bounded. The run history is the ground truth the platform generated when the last real deployment happened, and it answers one question cleanly: was this exact commit already deployed? That covers the common case. It does not cover every case.&lt;/p&gt;

&lt;p&gt;A rollback lands you at a commit the history already shows as deployed, but the environment still needs a run. A change applied outside the pipeline matches the same fingerprint for different reasons. An environment rebuilt from scratch needs a full deploy regardless of what the record says. For all of these, you need a force-deploy path: an explicit override that bypasses the skip check and runs unconditionally. Without it, the fingerprint comparison becomes a blind spot exactly where you cannot afford one. The check and the override are not competing ideas. They describe the same discipline: know what you are doing and be deliberate about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Systems aware of their own state
&lt;/h2&gt;

&lt;p&gt;Systems that are aware of their own state do not need people to compensate for their ignorance.&lt;/p&gt;

&lt;p&gt;Most pipelines are aware of half their state. They know when they run, what they build, and whether the steps pass. Adding state-awareness on the deployment side, knowing what code is actually behind the last green run, closes a loop that was always open.&lt;/p&gt;

&lt;p&gt;If your pipeline redeploys the same code every night because it ran last night and nothing broke, that is not stability. That is a loop with no exit condition. It looks like a working system. It is a system that does not know what "done" means.&lt;/p&gt;

&lt;p&gt;Ask the question the platform already knows how to answer. Query the history. Read the fingerprint. Build the override for when you genuinely need to force a run. The answer costs one API call and saves everything that comes after.&lt;/p&gt;

</description>
      <category>automation</category>
      <category>cicd</category>
      <category>deployment</category>
      <category>devops</category>
    </item>
    <item>
      <title>The Agent Atom: Build AI That Lasts</title>
      <dc:creator>praveenlavu</dc:creator>
      <pubDate>Sat, 12 Sep 2026 14:00:06 +0000</pubDate>
      <link>https://dev.to/praveenlavu/the-agent-atom-build-ai-that-lasts-hb4</link>
      <guid>https://dev.to/praveenlavu/the-agent-atom-build-ai-that-lasts-hb4</guid>
      <description>&lt;h1&gt;
  
  
  The Agent Atom
&lt;/h1&gt;

&lt;p&gt;Three months in, a production agent I had stopped watching started producing outputs that looked right but weren't. A dependency upstream had shifted. No crash. No alert. The agent kept running. By the time anyone noticed, weeks of downstream work had to be unwound.&lt;/p&gt;

&lt;p&gt;That is the specific shape of silent failure. Not something you triage at 2am. A drift you only discover when the trust is already gone.&lt;/p&gt;

&lt;p&gt;I had built maybe twenty agents before that incident. All of them had the same structural problem. Once I could name it, I could not unsee it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem With Scaffolding Wrappers
&lt;/h2&gt;

&lt;p&gt;What I was building were configurations, not agents. The thing I called an agent was a name in a routing table, a system prompt, and a set of assumptions about what the environment around it would provide. Memory lived in the orchestrator. Skills were implicit in the prompt. Knowledge was injected at call time with no guarantee of format or freshness. Identity was a label.&lt;/p&gt;

&lt;p&gt;I started calling these scaffolding wrappers. They look like agents. They respond like agents. But strip the pipeline away and there is nothing left that can stand on its own. The scaffolding is not support for the agent. It is part of the agent. Which means it is fragile in exactly the way the scaffolding is fragile.&lt;/p&gt;

&lt;p&gt;This is not the same failure mode as a bad prompt or the wrong tools. Those break loudly. The scaffolding wrapper breaks quietly, when the environment shifts in ways its author did not anticipate, and nothing internal to the agent can detect or correct the drift. That is the failure I kept building and kept shipping and kept explaining to people months later.&lt;/p&gt;

&lt;p&gt;The question I finally started asking was structural: what would an agent need to carry inside itself so it could function independently, compose with other agents reliably, and improve over time without being patched every time the world shifted?&lt;/p&gt;

&lt;h2&gt;
  
  
  The Turn That Changed Everything
&lt;/h2&gt;

&lt;p&gt;I was staring at the broken production agent when the word came to me: atom.&lt;/p&gt;

&lt;p&gt;Not as a metaphor I had been planning. As a diagnosis.&lt;/p&gt;

&lt;p&gt;In chemistry, an atom is the smallest unit that retains the properties of an element. It does not borrow protons from neighboring molecules. Everything it needs to be what it is lives inside the nucleus. Strip the surrounding structure away and it is still the same element. The environment does not define it. It defines itself.&lt;/p&gt;

&lt;p&gt;I had been building molecules and calling them atoms. That was the root cause. Not bad prompts. Not missing tools. A structural confusion about what belongs inside the agent and what belongs outside it. The dependency that shifted upstream had not broken a robust thing. It had exposed a thing that was never robust in the first place, because it had never been self-contained.&lt;/p&gt;

&lt;p&gt;That realization was not comfortable. It meant every agent I had shipped had the same flaw. But it also meant the fix was structural, not cosmetic. Get the structure right and the failure mode disappears.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an Agent Atom Actually Contains
&lt;/h2&gt;

&lt;p&gt;The structure that emerged is five layers, and all five have to be present.&lt;/p&gt;

&lt;p&gt;Identity comes first. Not a name and a description, but a full declaration of what the agent is, what it is capable of, and how it behaves when it succeeds, fails, or encounters something outside its domain. This is where the twelve self-star properties live: behavioral commitments that describe how the agent heals when it makes a mistake, how it monitors its own outputs, how it governs its own decisions, how it audits the work it produces. These cannot be bolted on from outside. They have to be part of the core definition, because the behavior they describe has to be consistent regardless of what is calling the agent.&lt;/p&gt;

&lt;p&gt;Skills are the second layer, and not just a list of things the agent can do. Each skill is a contract: here is what I accept as input, here is what I return as output, here is what success looks like, here is what failure looks like and what to do about it. Without that contract you cannot compose agents reliably. You are back to hoping the prompt is clear enough.&lt;/p&gt;

&lt;p&gt;Knowledge is third. An agent that has to be handed its domain context at call time is only as good as whatever gets passed to it. The knowledge layer is the agent's permanent relationship with its domain: the core concepts it reasons from, the research it has absorbed, the war stories from past invocations that taught it something real. This layer grows. The agent gets more capable over time without being rewritten.&lt;/p&gt;

&lt;p&gt;Memory is fourth, and this is where I spent the most time getting things wrong before I got them right. Memory has to have structure. Episodic memory holds what happened recently. Semantic memory holds what a curator process decided was worth keeping for the long term. Procedural memory holds patterns that have been explicitly validated and signed off. Without that structure, memory becomes noise. The agent either forgets everything or drowns in its own history.&lt;/p&gt;

&lt;p&gt;Reflection closes the loop. After every invocation the agent asks itself what it did, whether it worked, and what it would do differently. That is not optional introspective habit. It is the mechanism by which semantic memory accumulates signal instead of static. Reflection at per-invocation cadence is what creates the conditions for meaningful improvement over time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Self-Containment Is the First Principle
&lt;/h2&gt;

&lt;p&gt;Here is the real test: could you lift this agent out of your current system, drop it somewhere else, and have it still work?&lt;/p&gt;

&lt;p&gt;A scaffolding wrapper cannot survive that. It collapses without the environment it was built for.&lt;/p&gt;

&lt;p&gt;An atom can. Because everything it needs is inside it.&lt;/p&gt;

&lt;p&gt;This matters for composability. When agents hand work to each other, you need to trust that each agent in the chain brings its own context, its own judgment, and its own memory of what it has learned. If one agent in that chain is a wrapper, the whole chain becomes fragile at that joint. You end up maintaining not just the agent but the environment it depends on, which grows and shifts and eventually fails in a way you did not see coming.&lt;/p&gt;

&lt;p&gt;It matters for improvement too, but here is where most descriptions of this architecture go wrong. The five-layer atom does not improve by itself through some unsupervised process. It improves through a structured internal cycle: reflection after each invocation surfaces candidates, a curator promotes the ones worth keeping into semantic memory, and procedural patterns require explicit sign-off before they solidify. That is a different kind of improvement than waiting for someone to rewrite the prompt after the next breaking change. It scales. The wrapper model does not. But the improvement is earned through the internal process, not granted automatically.&lt;/p&gt;

&lt;p&gt;This is also not the same thing as giving a wrapper more memory or more tools. Adding memory to a wrapper makes a more capable wrapper. The atom is a different architecture. The five layers are inside the agent, not attached from outside. The improvement cycle is part of the identity definition, not a feature of the orchestrator. That is the distinction the chemistry analogy is actually pointing at, and it is why I think it is more than an analogy. Prior approaches to capable agents add capabilities to the outside of a thin core. The atom approach starts from the inside out and asks what the core has to contain before any external wiring is allowed.&lt;/p&gt;

&lt;p&gt;The agents I build now are atoms first. They take longer to specify up front. The five-layer structure is not trivial to work through for each domain. But what you get at the end is an agent that holds its shape when the world around it shifts, composes with other agents without duct tape, and accumulates real capability through structured reflection rather than through prompt revisions someone makes after the next breaking change.&lt;/p&gt;

&lt;p&gt;I still get the dopamine spike when I watch them run. But now I also get the thing I was missing before: they still work a month later. And the month after that.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>architecture</category>
      <category>production</category>
    </item>
    <item>
      <title>Budgeting for the AppExchange Security Review</title>
      <dc:creator>praveenlavu</dc:creator>
      <pubDate>Fri, 11 Sep 2026 14:00:06 +0000</pubDate>
      <link>https://dev.to/praveenlavu/budgeting-for-the-appexchange-security-review-3jmm</link>
      <guid>https://dev.to/praveenlavu/budgeting-for-the-appexchange-security-review-3jmm</guid>
      <description>&lt;h1&gt;
  
  
  Budgeting for the AppExchange Security Review
&lt;/h1&gt;

&lt;p&gt;The email came on a Thursday afternoon. Seventeen findings. One of them was a sharing model violation I had introduced nine weeks earlier, in a sprint that had nothing to do with the feature we were trying to ship. The launch date I had promised our first design partner was already three weeks behind us.&lt;/p&gt;

&lt;p&gt;I thought that was the hard lesson. I was wrong.&lt;/p&gt;

&lt;p&gt;When I resubmitted after remediation, the review came back with three new findings in code I had not touched. Between my original submission and the resubmission, the platform had updated its scanner version. Checks that had passed the first time failed under the new version. The fixes were not complicated, but the calendar moved again, and another external commitment slipped.&lt;/p&gt;

&lt;p&gt;The third surprise came after we were listed. Not from releasing a new version. The listing came back into review on an annual cycle, independent of anything we shipped. The scanner baseline had moved. Some of what the review had approved the first time was being evaluated against a higher standard.&lt;/p&gt;

&lt;p&gt;By the time I understood all three of those patterns, I also understood that I had been thinking about the security review entirely the wrong way.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Clock Starts Earlier Than You Think
&lt;/h2&gt;

&lt;p&gt;The AppExchange Security Review is not a formality. The marketplace takes it seriously because the security of every org that installs a package depends on what that review catches. That is the correct call. But the consequence for partners who treat the review as a box to check at the end of development is almost always a bruising surprise.&lt;/p&gt;

&lt;p&gt;The initial review cycle takes weeks, not days. The exact length depends on your submission's complexity, your security posture, and how many findings come back on the first pass. If you submit and get findings, you fix them and resubmit. Each resubmission restarts the clock on that portion of the review. Multiple rounds of back-and-forth can consume two months before approval.&lt;/p&gt;

&lt;p&gt;The math is simple but brutal: a team that submits once with clean findings ships in one review cycle. A team that submits with twenty findings, fixes them, resubmits, and picks up a few more ships in three or four cycles. Same feature set. Very different timelines.&lt;/p&gt;

&lt;p&gt;This is what the timeline conversation is actually about. It is not really about how long the review process takes. It is about how many cycles you hand them.&lt;/p&gt;

&lt;p&gt;There is a second clock most teams forget: the one that starts the moment you commit to a launch date. Every week in remediation is a week your design partners are waiting, your sales pipeline is stalling, and your engineering team is firefighting instead of building. The review does not exist in a vacuum. It sits in the middle of all your other commitments.&lt;/p&gt;

&lt;p&gt;The partners who plan well treat the review as a fixed cost on the timeline, the same way they treat QA or infrastructure provisioning. They do not schedule it for the end. They build backwards from the review window to set the development cutoff.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Triggers a Re-Review, and Why Your Roadmap Depends on Knowing
&lt;/h2&gt;

&lt;p&gt;Once you are listed, the security review does not disappear. It comes back in at least three ways, and most ISV roadmaps account for none of them.&lt;/p&gt;

&lt;p&gt;The first way is change-triggered. A new review is required when certain kinds of changes go into a managed package version. Changes to your permission model almost always trigger a full review. If a new version requests new object permissions, new field access, or elevated system privileges, plan for a complete review cycle. New external integrations or callouts to endpoints that were not in your original submission typically trigger review as well. Significant architectural shifts draw scrutiny even when they are not formally required to. If you are uncertain whether a planned change is a trigger, submit early and find out. The cost of discovering a trigger after you have already announced a GA date is far higher than the cost of asking the question two sprints ahead of time.&lt;/p&gt;

&lt;p&gt;The second way is time-triggered. Existing listings are subject to periodic review on the platform's schedule, not yours. The scanner baseline moves over time. What passed when you were first approved may be evaluated against current standards the next time a review cycle touches your listing. Planning for this means maintaining a compliance posture after listing, not only during submission. Annual reviews belong on the roadmap as a fixed calendar event, the same way recurring infrastructure maintenance does.&lt;/p&gt;

&lt;p&gt;The third way is scanner-version reclassification, and it is the one that hit me hardest mid-cycle. When the platform ships an updated scanner version, the criteria can shift. Findings that passed in a previous round can fail under the new version, even in code that has not changed. This is not a defect in the process; it reflects a genuinely evolving set of security standards. But for teams who are mid-review, it can extend the timeline through no fault of their most recent changes.&lt;/p&gt;

&lt;p&gt;The strategic implication is straightforward. Every significant feature on your roadmap needs a security-review timeline baked in before that release gets a public commitment. And your post-listing roadmap needs to account for annual reviews and scanner-version updates the same way it accounts for new releases.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why CI/CD Changes the Math
&lt;/h2&gt;

&lt;p&gt;The teams that move through review faster are not writing perfect code on the first attempt. What they have in common is that by the time they submit, they have been running security analysis against their builds for months.&lt;/p&gt;

&lt;p&gt;The platform offers a CLI-based static analysis tool that partners can run locally and in automated pipelines. This is a distinct instrument from the portal-submitted scanner the formal review uses; the underlying criteria overlap significantly, but the results are not guaranteed to be identical. The concrete shift is integrating this CLI analysis into your build pipeline as a merge condition, running it against every merge rather than treating a pre-submission scan as the primary safety net.&lt;/p&gt;

&lt;p&gt;A finding caught during a sprint costs a few hours to address. The same finding caught during formal review costs weeks of calendar time and team morale. A criteria update in the CLI tool that happens while you are mid-development shows up in your pipeline immediately, giving you advance notice of potential gaps before you commit to a review window. The practice is specific: add the CLI analysis to your CI configuration, set a threshold for acceptable findings, fail the build when that threshold is crossed, and treat new findings the same way you treat failing tests.&lt;/p&gt;

&lt;p&gt;The shift this requires is cultural more than technical. The analysis cannot live at the end of the pipeline as a pre-submission gate. It has to live in the middle, as a condition of merging. Once that norm is in place, the formal review stops feeling like a gamble. It becomes a confirmation of what you already know about your package, because potential gaps have been tracked continuously rather than inspected at a single point in time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Budget for the Review Like Any Other Real Cost
&lt;/h2&gt;

&lt;p&gt;The reason ISV timelines blow up at the security review is almost always that partners did not budget for it the way they budget for development work.&lt;/p&gt;

&lt;p&gt;I did not. My first review cost six weeks I had not allocated. My second cost another two weeks because a scanner update moved the target between rounds. My third was an annual review I had not put on the roadmap at all. By the time those three lessons landed, the cumulative cost was one missed design partner commitment, two slipped sales conversations, and an engineering team that had learned to dread the review rather than treat it as a normal part of shipping.&lt;/p&gt;

&lt;p&gt;The turn came when I stopped asking how long the review would take and started asking how many findings we would hand them. That single reframe changed how we planned every release afterward. The first decision was to add the CLI-based static analysis to every sprint's definition of done, the same way a failing test blocks a merge. The second was to write the review window into every roadmap before any external commitment went out the door. The third was to put the annual review date on the release calendar next to major version cuts, treating it as recurring maintenance rather than a surprise.&lt;/p&gt;

&lt;p&gt;The next submission we ran through that discipline came back clean on the first pass. Not because the code was flawless, but because the findings had already been found and fixed in development, weeks before anyone hit submit. The review confirmed what the pipeline already knew.&lt;/p&gt;

&lt;p&gt;That is the payoff. Not a clever workaround, and not a secret about how the review works. The teams that move through review quickly submitted something that was already compliant when they hit send. That outcome is the result of decisions made months before submission, and a roadmap that included the review as a recurring line item rather than a one-time gate.&lt;/p&gt;

&lt;p&gt;Build compliant. Submit compliant. Plan for the review to come back.&lt;/p&gt;

</description>
      <category>security</category>
      <category>software</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>When Claims Volume Breaks Salesforce Storage</title>
      <dc:creator>praveenlavu</dc:creator>
      <pubDate>Thu, 10 Sep 2026 14:00:03 +0000</pubDate>
      <link>https://dev.to/praveenlavu/when-claims-volume-breaks-salesforce-storage-53k6</link>
      <guid>https://dev.to/praveenlavu/when-claims-volume-breaks-salesforce-storage-53k6</guid>
      <description>&lt;h1&gt;
  
  
  When Salesforce Storage Limits Meet 50,000 Claims a Day
&lt;/h1&gt;

&lt;p&gt;The notification landed on a Tuesday afternoon in month four. I was the architect running a Salesforce-based claims platform that processed fifty thousand claims every day, and the storage dashboard I had calibrated carefully three months earlier had crossed the line I had marked as safe. The estimates were not wrong. The problem was that fifty thousand records a day does not stay abstract for long. At that volume, every number you were comfortable with eventually becomes a wall, and walls have a way of appearing faster than planned.&lt;/p&gt;

&lt;p&gt;I did not know then that I was three months away from building something I would use to explain distributed data architecture to every team I worked with afterward. I knew I had a problem with no obvious exit.&lt;/p&gt;

&lt;h2&gt;
  
  
  How 50,000 Claims a Day Breaks Your Assumptions
&lt;/h2&gt;

&lt;p&gt;Let me run through the math the way I ran through it at the time.&lt;/p&gt;

&lt;p&gt;Fifty thousand claims a day sounds like a lot. At the individual record level, it is manageable. A claim header in Salesforce is not enormous. Healthcare claims do not travel alone. Each claim brings claim lines, service codes, member information, payer adjudication details, and in many cases supporting documentation. When you count the full object graph, a single claim touches anywhere from five to ten records across the org. At fifty thousand claims a day, that is potentially half a million new records every twenty-four hours.&lt;/p&gt;

&lt;p&gt;Salesforce data storage is generous by many standards. But it is finite. And retention obligations mean you cannot delete claims. State law and payer contracts set the floors, commonly six to ten years, and some contracts run longer. So every record written today lives on the org for years, and the storage trajectory is a line that only goes one direction.&lt;/p&gt;

&lt;p&gt;Storage is the visible constraint. The less obvious one surfaces later. Salesforce enforces a hard governor limit of fifty thousand rows returned per SOQL query transaction. When historical claims data accumulates across years of operation, queries that need to look across the full retention window start hitting that ceiling. You cannot run a meaningful audit of adjudication patterns across eighteen months of data without breaking the query into fragments and reassembling results in application code. The platform that was excellent at processing individual claims in real time starts to strain when asked to think across the full history of those claims. Both constraints, the storage trajectory and the query ceiling, pointed toward the same underlying mismatch. They just announced themselves at different times.&lt;/p&gt;

&lt;p&gt;By month four, the math was clear. The question was what to do about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Wrong Turns Before the Right One
&lt;/h2&gt;

&lt;p&gt;The obvious answers came quickly and fell apart just as fast.&lt;/p&gt;

&lt;p&gt;Buying more storage buys time, not a solution. Salesforce storage is expensive per gigabyte relative to cloud alternatives. At fifty thousand claims a day, you cannot buy your way out of this. The spend compounds with retention. I ran the three-year projection and the numbers were not close to defensible. You are paying a premium to defer a problem, not solve it.&lt;/p&gt;

&lt;p&gt;The platform's native archival storage layer looked promising. It was designed for scale and long retention. But when I went deep on the constraints, the picture changed. Standard reports cannot see the data. Built-in automation tools cannot read from it in any practical sense. You can write data in, but getting it back out in a usable form requires significant custom development. The moment you need to correlate a historical claim with a current member record, you are writing code that fights the platform. For a claims operation, historical correlation is not an edge case. It is a daily workflow.&lt;/p&gt;

&lt;p&gt;Third-party archival tools introduced new operational dependencies, new contracts, and pricing structures that compound with volume. Two evaluations, two passes.&lt;/p&gt;

&lt;p&gt;There was a week in month five where I had eliminated every option I had originally considered and had nothing credible to replace them with. The storage bars kept climbing. The team was watching the dashboard the way you watch a patient's vitals when the numbers are moving in the wrong direction. I had a platform that was excellent at its job and a volume problem that seemed to have no clean answer inside the tools I had been given. The SOQL ceiling meant that even if I solved storage, I would eventually be unable to query the data I was keeping. Fixing one side of the constraint did not fix the other.&lt;/p&gt;

&lt;p&gt;That discomfort was, in retrospect, productive. It forced me out of the frame I had been working in.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Reframing That Became the Architecture
&lt;/h2&gt;

&lt;p&gt;I had been asking the wrong question. The question I kept asking was how to keep more claims in Salesforce. The question I should have been asking was what Salesforce is actually for.&lt;/p&gt;

&lt;p&gt;That shift happened on a Friday evening when I was reading the platform documentation for the third time, looking for something I had missed. I had not missed anything in the documentation. What I had missed was the frame. The moment I stopped asking how to store more and started asking what Salesforce was designed to do, the architectural answer became almost obvious.&lt;/p&gt;

&lt;p&gt;Salesforce is an operational system. It is excellent at managing the active state of a business process: a claim in flight, a case under review, a member record being updated. What it is not is a long-term historical data store designed for large-scale retention at low cost per byte. That is a different shape of problem. Different shapes call for different tools.&lt;/p&gt;

&lt;p&gt;Once I saw that, the architecture followed. Not one tier. Two. And the boundary between them would be the design.&lt;/p&gt;

&lt;p&gt;The first tier stays in Salesforce. It holds the active claims window, the data that people are actually touching, that workflows run against, that automations read and write. We settled on a ninety-day rolling window for claims in active status. Everything needed for daily operations lives here, on the platform that does daily operations well. At that window size, the fifty-thousand-row SOQL ceiling is not a practical obstacle. The queries you need to answer about active claims are queries the platform can answer cleanly.&lt;/p&gt;

&lt;p&gt;The second tier lives in cloud storage and a relational layer built for archival and analytical workloads. When claims age past the active window, they move here. The full fidelity of the record is preserved, everything a compliance audit or retrospective analysis would need, but the data no longer consumes Salesforce storage because it no longer lives there. This tier scales at cloud-object-storage prices, not per-gigabyte platform prices, and carries no row-count ceiling on the queries you can run across it.&lt;/p&gt;

&lt;p&gt;Between the two tiers sits a lightweight sync process. On a schedule, it identifies claims that have crossed the age threshold, writes them to the cold tier with full field mapping, and removes them from the hot tier after confirming the write succeeded. The removal from Salesforce is what frees storage. The write to cold storage is what preserves compliance.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Seam That Makes It an ODS, Not an Archive
&lt;/h2&gt;

&lt;p&gt;The detail that separates this architecture from a simple archive is the read layer that sits in front of both tiers. Without it, what you have is Salesforce for active claims and a compliance dump for everything older. That is a defensible setup if historical data exists only to satisfy auditors. It is not defensible if your team needs to correlate a current claim with its history, or if any downstream process needs a consistent view of claims regardless of age.&lt;/p&gt;

&lt;p&gt;With the unified read layer, downstream consumers do not need to know which tier a given claim lives in. They request a claim by identifier. The layer routes the request to the right store and returns the data in the same schema regardless of source. From the outside, there is one claims store. On the inside, there are two, each doing the job it was built for.&lt;/p&gt;

&lt;p&gt;That distinction matters operationally. An archive is a place you go to retrieve compliance evidence. An operational data store is a place your workflows query in real time. The read seam is the difference between the two. Without it, the cold tier is a write-only compliance artifact. With it, the cold tier is a live part of the system, as queryable as Salesforce, just optimized for a different cost and scale profile. The architecture is not a hot database plus an archive. It is one ODS built from two tiers with different optimization targets, surfaced to consumers as a single interface.&lt;/p&gt;

&lt;p&gt;That is what makes it worth the overhead.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Boundary Taught
&lt;/h2&gt;

&lt;p&gt;The storage problem went away. The query ceiling problem went away with it, because the data that had been accumulating against that ceiling was no longer in Salesforce. Not because we found more room, but because we stopped asking Salesforce to do something it was never designed for.&lt;/p&gt;

&lt;p&gt;Every cloud-native tool has a shape. The shape comes from what the tool was built to do well and what starts to cost more or work worse as you push past that design intent. Salesforce's shape is operational: collaborative, process-oriented, running at the speed of human decisions. Cloud object storage's shape is massive, cheap, durable, built for bulk reads and long retention.&lt;/p&gt;

&lt;p&gt;When those shapes match the problem, everything moves easily. When they do not, you get symptoms. Storage bars climbing faster than expected. Query timeouts at the row ceiling. Pricing that scales nonlinearly with volume. Those are not bugs in the platform. They are the platform communicating something about the shape of what you are trying to do.&lt;/p&gt;

&lt;p&gt;The right response is not to fight the platform. It is to draw a deliberate boundary between what the platform is for and what something else handles better, then make that boundary explicit, tested, and operationally simple.&lt;/p&gt;

&lt;p&gt;The overhead is real. There is a sync process to monitor, a mapping layer to maintain, a cold-tier schema that has to stay in step with the Salesforce schema as both evolve. It is not free. At this volume it is still cheaper than the alternatives. And it is far more stable than asking one system to stretch past what it was shaped for.&lt;/p&gt;

&lt;p&gt;Fifty thousand claims a day taught me that scaling problems dressed up as resource constraints are usually design problems in disguise. Find the right boundary, and the resource constraint stops being the problem. The design becomes the answer.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>database</category>
      <category>performance</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>Treat Your Blog Like an Agent Pipeline</title>
      <dc:creator>praveenlavu</dc:creator>
      <pubDate>Wed, 09 Sep 2026 14:00:06 +0000</pubDate>
      <link>https://dev.to/praveenlavu/treat-your-blog-like-an-agent-pipeline-9oe</link>
      <guid>https://dev.to/praveenlavu/treat-your-blog-like-an-agent-pipeline-9oe</guid>
      <description>&lt;h1&gt;
  
  
  Treating Blog Posts Like Agent Pipelines
&lt;/h1&gt;

&lt;p&gt;The night I shipped my sixth article that week, I knew something had changed. Not in the content itself, in the process. Each piece had moved through six distinct transformations before it reached a reader. Different constraints at every stage. Different quality bars. Clean handoffs between each one. Around midnight the frame snapped into place: I hadn't been writing blog posts. I'd been running a pipeline.&lt;/p&gt;

&lt;p&gt;This is not a post about productivity. It is about a structural realization that changed how I think about content production at a fundamental level. About what happens to output quality when you stop treating a blog post as a thing you write and start treating it as something that gets built, stage by stage, with each stage doing exactly one job and handing off a specific artifact to the next.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mess that made the pattern visible
&lt;/h2&gt;

&lt;p&gt;Before this clicked, my content process looked like most people's. An idea would show up. I'd sit down, write until I ran out of steam, revise once or twice, and ship when it felt done. Sometimes the output was good. Sometimes it was not. I could never reliably predict which kind I'd get.&lt;/p&gt;

&lt;p&gt;The variance was the problem. Not the average quality, the swing. Some pieces came out sharp and original. Others read like they were produced by someone who vaguely understood the topic but could not feel it. Same author, same topics, wildly different results depending on what time I sat down and how much sleep I had gotten.&lt;/p&gt;

&lt;p&gt;I tried the usual fixes. Outlines. Writing at consistent times. Reading good writing to absorb the rhythm. All of that helped at the margins. But none of it addressed the root cause: I was a single-threaded process trying to hold too many concerns at once. Research precision and narrative flow and voice consistency and technical accuracy and sentence-level rhythm all competing for the same cognitive slot at the same time.&lt;/p&gt;

&lt;p&gt;When you try to optimize everything simultaneously, you optimize nothing. That is as true for blog posts as it is for distributed systems. The failure I was heading toward was not dramatic. It was just consistent mediocrity with no way to trace where it came from, and no lever to pull to make it better.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the structure revealed itself
&lt;/h2&gt;

&lt;p&gt;The pivot came from a direction I did not expect. I had been building agent-based workflows for other problems: routing, classification, document processing. Each agent in those systems had one job. One input type, one output type, one quality bar. Agents do not multitask. They transform.&lt;/p&gt;

&lt;p&gt;One afternoon I was sketching a content piece on paper. I noticed I had naturally divided my notes into chunks: the technical claim in its raw form, the same claim translated for a practitioner audience, then the headline-sized summary of what a reader walking away should remember. Three different framings of the same core idea. Three different registers.&lt;/p&gt;

&lt;p&gt;That is when I saw it. I hadn't been sketching an article. I'd been sketching a transformation sequence. And when I traced it carefully, it had exactly six stages, each with a single job and a specific artifact it handed to the next.&lt;/p&gt;

&lt;p&gt;Stage one is claim verification. This stage has one job: extract the core claim and pressure-test it before anything else happens. The output is the claim in its raw technical form, with every load-bearing assumption labeled. It does not read well. It is not supposed to.&lt;/p&gt;

&lt;p&gt;Stage two is structural sequencing. This stage takes the verified claim and maps the dependencies: what does a reader need to accept before they can accept the main claim? What follows from what? The output is a sequence of propositions ordered by dependency, not yet prose.&lt;/p&gt;

&lt;p&gt;Stage three is the technical draft. This stage writes the first complete version, rigorous and dry, taking the sequence from stage two and doing nothing except honoring it. Accuracy is the only quality bar. The output is technically complete prose that has not been optimized for anything else.&lt;/p&gt;

&lt;p&gt;Stage four is the practitioner translation. This is where the technically accurate draft gets made followable by someone who does not live in the domain. The constraint is strict: this stage cannot soften the claim. It can find a better illustration or a more concrete framing. It cannot change what is being claimed. The output is readable without being imprecise.&lt;/p&gt;

&lt;p&gt;Stage five is the voice pass. This stage reads the translated draft and makes it sound like the person who wrote it. Rhythm. Repetition patterns. The specific words someone reaches for when they are not performing neutrality. The output is a voiced draft, but voice work can introduce imprecision that is easy to miss, which is why this stage is not the last.&lt;/p&gt;

&lt;p&gt;Stage six is the surface edit. This stage closes the loop. It reads the voiced draft for accuracy first, then cuts everything that does not carry weight, then tightens rhythm at the sentence level. It catches what the voice pass accidentally softened and what the translation pass left ambiguous. The output is final.&lt;/p&gt;

&lt;p&gt;The key property of this sequence is that each stage inherits its constraints from every stage that came before. Stage four cannot redefine the claim because stage one already fixed it. Stage six cannot introduce new structure because stage two already set it. Downstream stages are tightly coupled to upstream decisions. That coupling is not a limitation. It is what makes the system stable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline as discipline, not shortcut
&lt;/h2&gt;

&lt;p&gt;This is where the analogy can go wrong. Treating your content as a pipeline is not about automating the writing. It is about separating concerns that should never have been bundled together in the first place.&lt;/p&gt;

&lt;p&gt;A claim verification stage has one job: get the claim right. Not make it accessible, not make it compelling. Right. Verification mode and persuasion mode are cognitively incompatible. When you try to do both at once, you end up with claims that sound good but have not been pressure-tested, or claims that are technically defensible but impossible to follow.&lt;/p&gt;

&lt;p&gt;What I found, once I actually separated these stages in practice, was that each individual stage got harder in a productive way. The research stage had nowhere to hide behind "but it reads well." The voice pass had nowhere to hide behind "but the facts are right." Every stage had to do its one job well, because there was nothing upstream or downstream to compensate for it failing.&lt;/p&gt;

&lt;p&gt;That is the discipline. The pipeline does not make the work easier. It makes the failure modes visible, and it makes them one at a time instead of all at once.&lt;/p&gt;

&lt;p&gt;Working under those constraints changes how the work feels. When a piece comes out well, you know which stage it came from. You can go back and strengthen that stage. When a piece comes out poorly, you can usually trace it to a specific handoff that broke down. That is debuggable. The old approach was not. When everything is one thing, nothing is traceable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What holds this together
&lt;/h2&gt;

&lt;p&gt;The principle underneath all of this is simpler than it sounds: complex things get better when you decompose them into stages with clear jobs. Engineers have known this for decades. But there is a reason it took me longer to apply it to writing than to code. Writing feels personal in a way that makes structure feel like a threat to authenticity.&lt;/p&gt;

&lt;p&gt;It is not. Structure does not flatten voice. What flattens voice is writing under cognitive overload, trying to be technically accurate and narratively compelling and emotionally honest all at the same time, in the same session, with no separation between those concerns.&lt;/p&gt;

&lt;p&gt;Six stages gives each concern its own space. The technical stage does not have to apologize for being dry. The voice pass does not have to apologize for caring more about rhythm than precision. Each stage gets to be fully itself because it is not competing for the same moment.&lt;/p&gt;

&lt;p&gt;The first time I read a piece cold the morning after it had gone through this full sequence, sharp, accurate, my voice, nothing padded, no hedging, I felt something I had not felt about my writing in a while. It felt engineered. I mean that in the best possible way. Not assembled. Designed. The difference you feel when something was built with intention at each step, rather than improvised through from beginning to end.&lt;/p&gt;

&lt;p&gt;That is the payoff. Not speed. Not volume. A consistent quality bar, and a process that tells you exactly where to look when the bar gets missed.&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>writing</category>
    </item>
    <item>
      <title>Your Local LLM Has a Hidden Context Limit</title>
      <dc:creator>praveenlavu</dc:creator>
      <pubDate>Sat, 05 Sep 2026 14:00:06 +0000</pubDate>
      <link>https://dev.to/praveenlavu/your-local-llm-has-a-hidden-context-limit-26jf</link>
      <guid>https://dev.to/praveenlavu/your-local-llm-has-a-hidden-context-limit-26jf</guid>
      <description>&lt;h1&gt;
  
  
  The Number Your Local LLM Won't Tell You
&lt;/h1&gt;

&lt;p&gt;There's a particular kind of wrong that feels deeply personal when you're debugging alone late at night. The kind where the model is supposed to be smart enough, the hardware is supposed to be fast enough, and the setup took you a weekend to get right. And then the output is just... wrong. Not wrong in a way you can point at. Wrong in a way that makes you question your own reasoning before you question the machine.&lt;/p&gt;

&lt;p&gt;That's where I was when I first discovered that my local LLM server had a threshold it never told me about.&lt;/p&gt;

&lt;h2&gt;
  
  
  The symptom that pointed nowhere
&lt;/h2&gt;

&lt;p&gt;The pipeline was simple enough on paper. Feed a document in, get a structured response out. I'd done this hundreds of times with the same model. But as the documents grew longer, something subtle started happening. The responses drifted. Answers that should have pulled from the middle of a long input started pulling from the top instead. Summaries felt like they were written by someone who only skimmed the first third.&lt;/p&gt;

&lt;p&gt;My first instinct was to blame the prompt. That's always the first instinct. So I rewrote it. Made the instruction clearer. Added more explicit direction. The responses got worse.&lt;/p&gt;

&lt;p&gt;Then I blamed the temperature setting. Cranked it down to nearly deterministic. Same drift.&lt;/p&gt;

&lt;p&gt;Then the model itself. Maybe I needed a different variant, a different quantization. I downloaded alternatives and spent hours comparing outputs. All of them showed the same pattern: impressive on short inputs, strangely shallow on long ones.&lt;/p&gt;

&lt;p&gt;What I didn't check, for an embarrassingly long time, was the server.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gap between what a model can do and what your server will let it do
&lt;/h2&gt;

&lt;p&gt;Here is the thing about running models locally that took me too long to fully internalize: the model's advertised context window and the server's actual working limit are two entirely different numbers. The model might be capable of attending to a very large window. But the server configuration, the runtime defaults, the available memory allocation at startup, all of these impose their own ceiling. And that ceiling can sit well below what the model card promises.&lt;/p&gt;

&lt;p&gt;The server doesn't refuse you when you cross that line. It doesn't throw an error you can search for. It just quietly handles as much as it can and lets the rest fall off. The output looks plausible. It might even look good. But it's responding to a truncated version of your input, and it has no way to tell you that.&lt;/p&gt;

&lt;p&gt;This is the hidden threshold. Not hidden in the sense that it's a secret, but hidden in the sense that nobody points at it. The documentation talks about model capabilities. The README talks about hardware requirements. The threshold you actually need to know, the one that determines where your pipeline breaks without warning, lives in a configuration file that most people never revisit after the initial setup.&lt;/p&gt;

&lt;p&gt;Finding it is an exercise in methodical elimination. You start with a large input, something you know should work, and you start contracting it. You watch for the point where the responses shift from shallow to accurate. That shift marks the boundary, roughly. Then you probe around that boundary to sharpen it. It is slow, manual, slightly tedious work. It is also the only reliable way to know what you're actually working with.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters more than it sounds
&lt;/h2&gt;

&lt;p&gt;I know how this can sound from the outside. An obscure configuration detail, a one-time debugging session, something you fix and forget. But I've watched this specific failure mode burn hours across multiple pipelines. Smart people, all of them, all making the same assumption I did: that the server would surface meaningful errors when something important was wrong.&lt;/p&gt;

&lt;p&gt;The deeper principle is about the gap between specification and behavior. A model's context window is a specification. What the server actually does with a given input under real memory pressure and real configuration defaults is behavior. These are related but not identical. In every system I've built where they differed, the specification was the thing I remembered and the behavior was the thing that bit me.&lt;/p&gt;

&lt;p&gt;Finding the actual threshold changes how you build. Once you know where the real ceiling is, you can design around it. You can split inputs intelligently, at logical boundaries, rather than discovering at runtime that something important was cut. You can add verification that the piece of information you needed was actually inside the window the server used. You can stop tuning prompts for a problem that was never a prompt problem.&lt;/p&gt;

&lt;p&gt;That late-night debugging session eventually ended with a number. Just a number. A threshold the server would work with reliably, below which everything was crisp and above which things went quietly sideways. Writing it down felt anticlimactic. But then I ran the pipeline again, this time respecting that number, and the outputs were immediately, obviously better. That moment still gives me the particular kind of satisfaction that no amount of prompt engineering ever has.&lt;/p&gt;

&lt;p&gt;The model hadn't gotten smarter. I had just stopped asking it to work with inputs it was never actually seeing.&lt;/p&gt;

&lt;p&gt;The step I now treat as mandatory, before any serious use of a locally-served model, is the threshold audit. Not a formal benchmark, nothing elaborate. Just a structured pass where I feed in inputs of increasing size and watch where the response quality changes. It takes maybe thirty minutes the first time and almost nothing on subsequent setups because I know what I'm looking for.&lt;/p&gt;

&lt;p&gt;The number you find from that pass becomes a hard constraint in whatever you build next. Not a suggestion, a constraint. Every input pipeline that feeds the server gets a budget check against that number before the request goes out.&lt;/p&gt;

&lt;p&gt;This is boring operational discipline. It is also the difference between a pipeline that degrades mysteriously over time and one that fails loudly at a boundary you control. Given the choice, I will take the loud, controlled failure every time. Silent degradation at scale is the thing that keeps you up long after the late-night debugging session should have ended.&lt;/p&gt;

&lt;p&gt;Your local LLM server has a threshold it hasn't told you about. Go find it before it finds you.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>debugging</category>
      <category>llm</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>What AI Deployment Gets Wrong About Security</title>
      <dc:creator>praveenlavu</dc:creator>
      <pubDate>Wed, 02 Sep 2026 14:00:07 +0000</pubDate>
      <link>https://dev.to/praveenlavu/what-ai-deployment-gets-wrong-about-security-2mao</link>
      <guid>https://dev.to/praveenlavu/what-ai-deployment-gets-wrong-about-security-2mao</guid>
      <description>&lt;h1&gt;
  
  
  The Network Security Principle AI Deployment Gets Wrong
&lt;/h1&gt;

&lt;p&gt;I spent a long time thinking about AI agent security the wrong way.&lt;/p&gt;

&lt;p&gt;Every conversation in the field seemed to orbit the same questions. Prompt injection. Jailbreaks. Hallucinated outputs. How do you stop a model from saying something it shouldn't? That's a real question and worth serious attention. But it kept pulling focus away from something that turned out to matter just as much: not what an agent says, but where it reaches.&lt;/p&gt;

&lt;p&gt;The moment I noticed the gap, I was debugging a production issue. One of the agents had made an outbound call to an endpoint that had no business being reachable from its execution environment. The call wasn't malicious. It was just there. An API the agent had encountered in its context window, available because nothing in the infrastructure prevented it. The request went out. The response came back. No alarm fired. From a network perspective, everything worked exactly as designed.&lt;/p&gt;

&lt;p&gt;That's what unsettled me. Everything worked exactly as designed.&lt;/p&gt;




&lt;p&gt;The problem with reasoning about AI security purely through content is that it treats the model as the perimeter. If the model produces safe outputs, you're safe. But modern AI agents don't just generate text. They call tools. They retrieve context from external sources. They write to storage and invoke downstream services. The surface area isn't the model's output window; it's every endpoint the agent can reach at runtime.&lt;/p&gt;

&lt;p&gt;Network engineers figured this out for software systems decades ago. The insight wasn't "trust the software to make good decisions about what it accesses." It was "define what the software is allowed to access in the first place." Zone-based architectures came from that thinking. Instead of asking whether a particular request is safe, you ask whether the traffic pattern fits the relationship between two defined zones. You draw the zones, you define the permitted flows, and everything outside those flows gets dropped before the safety question even comes up.&lt;/p&gt;

&lt;p&gt;That reframing is exactly what AI agent egress is missing.&lt;/p&gt;




&lt;p&gt;Most AI agent deployments handle outbound traffic the same way early web development handled database queries: with good intentions and a hope that nothing goes wrong. The agent has a list of tools. The tools call APIs. There's usually some rate limiting. Maybe an allowlist of domains that's half-maintained and never pressure-tested.&lt;/p&gt;

&lt;p&gt;What's missing is the structural guarantee.&lt;/p&gt;

&lt;p&gt;An allowlist living in application code can be bypassed. A tool that fetches web content can be directed at internal infrastructure if the prompt is crafted to point it there. A model trained to be helpful will try to fulfill requests, and if a request subtly asks it to retrieve a document from a service it was never supposed to reach, nothing in the content layer stops it. The model doesn't know it's doing something wrong, because from its perspective it isn't. It's just following through on what the context asked.&lt;/p&gt;

&lt;p&gt;Zone-based egress addresses this at the infrastructure layer, not the application layer. You define zones based on trust and sensitivity. The agent lives in its execution zone. Internal services live in a separate, protected zone. The open internet is its own zone. Traffic flows between zones follow explicit policy: an agent can call services in the approved-integrations zone, but it cannot initiate traffic toward internal infrastructure. That policy is enforced at the network level, not in the model's decision-making, not in the application code, not in a system prompt that instructs the agent to "only call approved endpoints." Policy in a prompt is behavioral. Zone enforcement is structural.&lt;/p&gt;

&lt;p&gt;This distinction matters more than it sounds. Trust must be structural, not behavioral. You don't secure a system by teaching it to behave correctly. You secure it by constraining what it can do regardless of how it behaves.&lt;/p&gt;




&lt;p&gt;When I started applying zone-based thinking to agent deployments, a few things shifted. The attack surface shrank. An agent that literally cannot reach internal infrastructure doesn't need perfect defenses against every variant of prompt injection targeting that class of threat. The zone boundary does the work before the threat even gets to be a threat. Debugging got cleaner too. When something unexpected happened with egress, the zone policy gave me an audit trail: what was attempted, what was permitted or denied, which zone the traffic originated from. I could answer the question quickly, rather than reconstructing it from scattered logs.&lt;/p&gt;

&lt;p&gt;The biggest shift, though, was in how I reasoned about agent capabilities before deployment. Instead of asking "what can this agent do?", I started asking "what zones can this agent reach, and from which zones can it be reached?" That's a much more answerable question. It translates directly into infrastructure policy, and it can be audited without reading model weights or testing behavioral edge cases.&lt;/p&gt;




&lt;p&gt;None of this is novel engineering. Zone-based architectures are standard network security practice. The idea of separating trust zones and enforcing directional flows is decades old. The reason it's not standard practice in AI deployment is that most attention is concentrated at the model layer. That's where the novelty lives. That's where the interesting research is happening. The infrastructure layer feels mundane by comparison.&lt;/p&gt;

&lt;p&gt;But mundane is usually where the serious vulnerabilities are waiting.&lt;/p&gt;

&lt;p&gt;If you're thinking about how to secure an AI agent, don't only ask what it can say. Ask what it can reach. Draw the zones. Define the permitted flows. Enforce them below the application layer. The security properties that actually hold under adversarial conditions come from structural constraints, not from prompts or training or hoped-for behaviors.&lt;/p&gt;

&lt;p&gt;Zone by zone. That's how you build something you can actually trust.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>cybersecurity</category>
      <category>security</category>
    </item>
    <item>
      <title>The LWC CSP Wall Nobody Warned You About</title>
      <dc:creator>praveenlavu</dc:creator>
      <pubDate>Tue, 01 Sep 2026 14:00:12 +0000</pubDate>
      <link>https://dev.to/praveenlavu/the-lwc-csp-wall-nobody-warned-you-about-41cb</link>
      <guid>https://dev.to/praveenlavu/the-lwc-csp-wall-nobody-warned-you-about-41cb</guid>
      <description>&lt;h1&gt;
  
  
  The CSP Wall Nobody Warned You About
&lt;/h1&gt;

&lt;p&gt;The component worked. It passed review, it passed testing, it passed QA on every sandbox I had access to. Then it shipped and a portion of the customer base started reporting that a specific interaction just stopped working. The other customers had no issues. Same component. Same package version. Different orgs.&lt;/p&gt;

&lt;p&gt;I spent what felt like a long time trying to reproduce it. Swapping API versions, checking feature flags, comparing org configurations side by side. The logs were almost useless. The behavior wasn't a clean crash with a stack trace; it was a silent failure, the kind where the UI just stopped responding with no indication of why.&lt;/p&gt;

&lt;p&gt;What I was actually looking at was a runtime problem. Not a bug in my code, exactly. A collision between two fundamentally different security models that the platform runs simultaneously, and the place where those models diverge happened to be exactly where my code did what it did.&lt;/p&gt;

&lt;p&gt;That is the CSP wall. And almost nobody writes about what it actually feels like to hit it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two Runtimes, One Platform
&lt;/h2&gt;

&lt;p&gt;Salesforce runs your Lightning Web Component JavaScript inside a security container. The purpose is isolation: preventing components from reaching outside their sandbox to do things the platform should not allow. The container enforces Content Security Policy, restricts certain DOM operations, and controls how your JavaScript interacts with the browser's native APIs.&lt;/p&gt;

&lt;p&gt;The catch is that there are two of these containers, with meaningfully different behaviors. The older one, Locker Service, was the first serious attempt at this kind of isolation. It wraps DOM elements in proxy objects, so when your code accesses a native element, it's actually talking to a secured wrapper that filters what you're allowed to do. This worked, it shipped to production orgs everywhere, and a generation of ISV developers wrote code assuming this model was the whole story.&lt;/p&gt;

&lt;p&gt;Then Lightning Web Security arrived as a next-generation approach, built on a different architectural philosophy. Instead of proxy wrappers, it uses the browser's own capabilities for sandboxing, a more principled design that aligns better with how modern JavaScript engines actually work. It's stricter in some places, more permissive in others, and it handles CSP enforcement differently at a fundamental level.&lt;/p&gt;

&lt;p&gt;The problem for anyone shipping a managed package to a broad customer base is that both are real. Depending on an org's settings, API version, and feature enablement, your component runs in one or the other. You don't get to choose. The org chooses.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Shape of the Wall
&lt;/h2&gt;

&lt;p&gt;What this means in practice is that code patterns which pass cleanly under one runtime can hit hard stops under the other. The wall does not announce itself. Your component does not throw an error that says "CSP violation detected under Lightning Web Security." Instead you get behavior differences that look like flaky tests, like environment quirks, like someone has a corrupted browser cache or a plugin conflict.&lt;/p&gt;

&lt;p&gt;The patterns that most reliably create this problem are the ones that touch JavaScript in ways that aren't purely declarative. Dynamic evaluation, runtime script construction, certain ways of reaching into the internals of third-party libraries, certain approaches to DOM manipulation that assume a specific relationship between your JavaScript context and the browser's native layer. These are exactly the patterns that experienced front-end engineers reach for naturally, because they're the ones that have historically given you power and flexibility on the open web.&lt;/p&gt;

&lt;p&gt;Under the stricter runtime, that power is precisely the problem. CSP exists to prevent the kind of dynamic behavior that enables certain attack vectors, and the security container is not designed to distinguish between "this dynamic thing is fine, I promise" and "this dynamic thing is a potential injection surface." It enforces the policy because that is what it is for.&lt;/p&gt;

&lt;p&gt;The most frustrating version of this is when you're integrating a third-party library, a visualization component, a charting tool, anything that was designed for the unrestricted browser environment and has its own ideas about initialization. Libraries like this often do things internally that are entirely normal in an open context but that land hard against a CSP wall. You didn't write the problematic patterns. You're just importing a dependency. And now you're the one debugging why it silently fails in a subset of customer orgs that you can't easily access or replicate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Gets You Through
&lt;/h2&gt;

&lt;p&gt;The shift that finally made this tractable for me was not finding a clever way to detect which runtime is active and branch accordingly. That path leads somewhere bad: a proliferation of conditional logic, an ever-growing matrix of runtime-specific workarounds, a codebase nobody can reason about six months later. It also fails eventually, because the detection surface is not stable across platform releases.&lt;/p&gt;

&lt;p&gt;The shift was treating the strictest interpretation of CSP as the only target. Not "will this pass Locker Service?" Not "will this pass Lightning Web Security?" Instead: does this do what I need using only capabilities that a maximally strict policy would allow? When you frame it that way, the question gets simpler. You're not navigating between two sets of rules. You're writing to a single bar that both runtimes must pass.&lt;/p&gt;

&lt;p&gt;In practice that means no dynamic evaluation in any form. No patterns that construct and execute JavaScript at runtime. No assumptions about what the global scope looks like or what native APIs are directly reachable. No importing external libraries without verifying that they were designed with strict CSP in mind, or without wrapping them in an abstraction that keeps the unsafe internals isolated from your component's execution context.&lt;/p&gt;

&lt;p&gt;It also means: declarative over imperative, everywhere you can achieve it. Build behavior into reactive properties and event handlers rather than reaching out at runtime to touch things imperatively. The reactive model that Lightning Web Components is built on exists partly because it is fundamentally more compatible with security isolation than imperative DOM manipulation. Working with it rather than around it is not a constraint you're accepting reluctantly. It's the design expressing its intent.&lt;/p&gt;

&lt;p&gt;The payoff surprised me. When you get this discipline right, the component becomes more maintainable, not less. The strictest-common-denominator approach forces you to reason clearly about what the component actually needs to do versus what is incidental complexity that accumulated from years of writing in environments without these restrictions. When the platform enforces a discipline, you stop accumulating that debt.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Insight That Doesn't Show Up in the Docs
&lt;/h2&gt;

&lt;p&gt;What nobody told me explicitly, but what I had to work out through a series of painful episodes across orgs I had no local access to, is that the dual-runtime situation is not a transitional state on the way to one unified model. It is the operational reality for any developer shipping broadly on this platform. The variance is not going away. Different customers move at different paces, have different features enabled, have made different configuration decisions at the org level. You are always shipping to a heterogeneous environment, and you are always the last one to find out what that environment looks like.&lt;/p&gt;

&lt;p&gt;This is a harder version of the general problem of writing code that works across environments. It's harder because the environments are not fully documented, because the behavioral differences are subtle enough to miss in standard testing, and because the failure modes are often silent rather than loud. The org that's breaking your component is not generating an obvious error you can act on. It's just not doing the thing you expected.&lt;/p&gt;

&lt;p&gt;The answer is not to master both runtimes deeply enough to code specifically to each. The answer is to code to the constraints that both enforce, which means stopping the patterns that neither should allow. That discipline, once it's internalized, is what makes your components actually portable. Not portable in theory, not portable on the sandboxes you control, but portable in the sense that they pass in every org you actually ship to, including the ones you haven't encountered yet.&lt;/p&gt;

&lt;p&gt;That's the payoff. It's boring to describe. It is genuinely powerful to have.&lt;/p&gt;

</description>
      <category>debugging</category>
      <category>frontend</category>
      <category>javascript</category>
      <category>security</category>
    </item>
    <item>
      <title>Stop Duplicate Healthcare Claims at Intake</title>
      <dc:creator>praveenlavu</dc:creator>
      <pubDate>Sat, 29 Aug 2026 14:00:02 +0000</pubDate>
      <link>https://dev.to/praveenlavu/stop-duplicate-healthcare-claims-at-intake-36kb</link>
      <guid>https://dev.to/praveenlavu/stop-duplicate-healthcare-claims-at-intake-36kb</guid>
      <description>&lt;h1&gt;
  
  
  The Composite Key That Stops Duplicate Healthcare Claims at the Door
&lt;/h1&gt;

&lt;p&gt;There is a class of problem in healthcare claims processing that does not crash anything. No alarm fires. No log goes red. The system keeps running, claims keep flowing, and somewhere inside that volume, the same claim gets paid twice.&lt;/p&gt;

&lt;p&gt;Duplicate claims are one of the quietest and most expensive problems in the space. They hide inside normal-looking transaction counts. They arrive days apart, sometimes weeks apart, in slightly different shapes. Each one, inspected individually, looks valid. The processor ingests them both. The damage surfaces in reconciliation reports, in audits, in conversations nobody wants to initiate.&lt;/p&gt;

&lt;p&gt;When I first started working seriously with claims pipelines, I assumed deduplication was a solved problem. It is not. It is a design problem, and most systems solve it at the wrong layer, at the wrong time, with the wrong tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Catching It Late Costs More Than It Saves
&lt;/h2&gt;

&lt;p&gt;The conventional approach to duplicate claims is remediation. Claims arrive, get staged, run through adjudication, and then somewhere downstream a matching process tries to identify what slipped through. Catch it late, reverse it, recover the overpayment, close the loop.&lt;/p&gt;

&lt;p&gt;This works. Sometimes. But it is expensive in every direction. By the time the duplicate is identified, it has already consumed compute at intake. It occupied a slot in the staging pipeline. It ran through whatever rules engines sit in adjudication. It may have already triggered downstream processes, notifications, or payment disbursements. Reversals are their own workflows, with their own failure modes and their own operational overhead.&lt;/p&gt;

&lt;p&gt;The latency problem compounds this. A claim submitted on Monday might not have its duplicate caught until Friday. By then the payment may already be in transit. Recovery is slower and messier than prevention ever would have been, and the submitter's experience is worse too: they submitted something, it was processed, and now they are getting a reversal notice three days later with no clear signal about what happened.&lt;/p&gt;

&lt;p&gt;I spent a meaningful stretch of time looking at this problem from the wrong end. I was optimizing the recovery path, making downstream matching faster and smarter and more resilient. It helped at the margins. But the fundamental issue kept reasserting itself: I was cleaning up a mess that did not have to be made.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fingerprint You Already Have
&lt;/h2&gt;

&lt;p&gt;The insight, when it finally landed, was almost embarrassing in how obvious it looked in retrospect.&lt;/p&gt;

&lt;p&gt;A healthcare claim is not random data. It describes a specific event: a specific member, receiving a specific service, from a specific rendering provider, on a specific date. That event happened once. The claim that represents it should be unique. Which means if you can fingerprint that event deterministically from the fields already present on the claim, you have everything you need to deduplicate at the moment of intake.&lt;/p&gt;

&lt;p&gt;The composite key is assembled from the fields that define the clinical event. Member identifier. Date of service. Procedure code. Rendering provider. Sometimes the service facility. The exact combination depends on the payer's rules and the claim type, but the underlying principle holds: those fields together describe something that happened once. Two claims that share all of those values are, with very high confidence, the same claim submitted twice.&lt;/p&gt;

&lt;p&gt;Generate that key at intake. Before staging. Before the claim enters any processing queue. Check it against a record of keys you have already seen. If it matches, stop the claim at the door.&lt;/p&gt;

&lt;p&gt;That is the shift. From catching problems on the way out to catching them on the way in. Once you frame it that way, the prior approach starts to look less like a feature and more like a structural design choice that got frozen in place before the cost of it was fully understood.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Changes When You Move the Gate
&lt;/h2&gt;

&lt;p&gt;The downstream effects of stopping duplicates at intake compound quickly.&lt;/p&gt;

&lt;p&gt;Compute cost drops, because every claim stopped at the door is a claim that never touches the adjudication engine, never occupies a staging slot, never triggers downstream workflows. In pipelines handling meaningful claim volume, that arithmetic accumulates fast.&lt;/p&gt;

&lt;p&gt;The rest of the pipeline gets cleaner, because downstream systems can make assumptions they could not make before. If a claim reached adjudication, it already cleared the fingerprint check. You do not need to build duplicate-awareness into every downstream step to catch what the intake layer missed.&lt;/p&gt;

&lt;p&gt;The operational conversation changes too, and this one surprised me more than the performance numbers did. When a duplicate is stopped at intake, there is a clear boundary event. The claim arrived. It was checked against the key store. It was rejected at the door with a reason the submitter can act on immediately. That is a materially better outcome than processing the claim fully and then unwinding it three days later. A clean rejection at intake is actually a service to the submitter, not a failure.&lt;/p&gt;

&lt;p&gt;The design also forces a question that sounds obvious but is genuinely hard to answer without it: what does "duplicate" actually mean for this pipeline? Is a claim a duplicate if the procedure code differs by one position? What if the date of service matches but the billed amount does not? The composite key definition is where you encode those answers. It becomes the authoritative statement of uniqueness for that pipeline, and that clarity has value completely independent of the deduplication logic itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the Complexity Actually Lives
&lt;/h2&gt;

&lt;p&gt;The pattern is simple. The implementation earns its complexity.&lt;/p&gt;

&lt;p&gt;Key generation has to be deterministic across submission formats. Healthcare claims arrive through multiple channels in multiple formats. A claim submitted as a standard transaction file and the same claim submitted through a web portal may express the same clinical data in structurally different ways. If the key generation logic does not normalize those representations carefully before building the fingerprint, you will produce different keys for what is functionally the same claim, and the duplicate walks through unchallenged. Format normalization is where most of the real engineering work lives.&lt;/p&gt;

&lt;p&gt;The key store has to be fast and durable. Every incoming claim goes through a read-then-write at intake speed, which is a low-latency requirement on a store that is also being written to continuously under real load. Durability is strict: if you lose a key you have already seen, you have created a gap. If you return a false positive, you have rejected a legitimate claim. The failure modes in both directions have real consequences for providers and payers alike.&lt;/p&gt;

&lt;p&gt;Concurrency edge cases are real. Two identical claims arriving within a narrow window, before either key has been committed to the store, can both pass through if the intake layer is not designed to handle concurrent arrivals safely. In high-volume periods, this is not a theoretical risk. It is a genuine threat to the deduplication guarantee.&lt;/p&gt;

&lt;p&gt;None of these are unsolvable. They are engineering problems with engineering solutions. But they are worth naming clearly, because the elegance of the pattern can make the implementation look simpler than it is, and the places where it gets hard are exactly the places where a rushed implementation tends to create gaps.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prevention Is a Different Architecture
&lt;/h2&gt;

&lt;p&gt;The shift from downstream matching to intake-layer fingerprinting is not just a technical optimization. It is a change in how you think about data quality across the whole system.&lt;/p&gt;

&lt;p&gt;Remediation architecture assumes a percentage of bad data will get through and designs for recovery. Prevention architecture asks what you actually know at the boundary, and uses that knowledge to stop bad data from entering in the first place.&lt;/p&gt;

&lt;p&gt;Healthcare claims carry enough semantic structure that you can fingerprint them at intake with high confidence. The clinical event is described in the claim itself, in fields that exist for exactly that purpose. That is a property worth using. When you use it, everything downstream gets simpler, because the hard problem was solved at the front door instead of being forwarded to every system in the chain.&lt;/p&gt;

&lt;p&gt;The composite key is not a novel algorithm. It is an application of a principle that holds across almost any domain where uniqueness can be defined: state the definition precisely, enforce it early, and let everything downstream benefit from the guarantee.&lt;/p&gt;

&lt;p&gt;Once you see the pipeline through that lens, it is hard to unsee it.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>softwareengineering</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Why Your Team Formation Floor Is Lying to You</title>
      <dc:creator>praveenlavu</dc:creator>
      <pubDate>Wed, 26 Aug 2026 14:00:07 +0000</pubDate>
      <link>https://dev.to/praveenlavu/why-your-team-formation-floor-is-lying-to-you-3pkk</link>
      <guid>https://dev.to/praveenlavu/why-your-team-formation-floor-is-lying-to-you-3pkk</guid>
      <description>&lt;h1&gt;
  
  
  The Floor You Set
&lt;/h1&gt;

&lt;p&gt;There was a moment, somewhere around the third failed routing attempt in a row, when I realized the problem wasn't the agent. It was me. More specifically, it was a rule I had written with the best of intentions and then quietly forgotten about.&lt;/p&gt;

&lt;p&gt;The rule said: a team needs at least two agents.&lt;/p&gt;

&lt;p&gt;Sounds reasonable, right? "Team" implies multiple people. You wouldn't call a solo run a team effort. So when I built the team formation layer, I put in a floor. Minimum two. If a task only needed one specialist, the system would route it anyway, pad the formation with a second agent who had no real role, and call it a team. Problem solved. Except it wasn't.&lt;/p&gt;




&lt;p&gt;The thing about minimum floors is that they seem conservative. They feel like safety. When you write them, you're thinking: I want to avoid degenerate cases. One-agent "teams" seem wrong, so I'll prevent them. The logic is intuitive. The consequences aren't.&lt;/p&gt;

&lt;p&gt;What actually happens is that the system starts lying to itself. A task arrives that needs exactly one specialist. The router looks at it, sees the floor requirement, and silently pulls in a second agent, not because the task needs two perspectives, but because the rule demands it. That second agent is now consuming resources, participating in a coordination handshake, and occasionally injecting its own framing into a decision it has no business touching. The output drifts. The overhead scales. And the failure mode is quiet, which is the worst kind.&lt;/p&gt;

&lt;p&gt;I spent a while looking for bugs in the wrong places. The agents themselves were fine. The routing logic was correct. The task descriptions were clear. Everything checked out, and the system was still behaving oddly on tasks that should have been simple single-dispatch jobs. It took pulling up the formation logic and staring at it for longer than I'd like to admit before the minimum floor line jumped out at me.&lt;/p&gt;




&lt;p&gt;The fix looked almost too small to be the answer.&lt;/p&gt;

&lt;p&gt;Remove the floor. Let a team of one be a valid team. Stop treating solo dispatch as a degenerate case that needs correcting.&lt;/p&gt;

&lt;p&gt;That's it. No new architecture. No new agents. No refactor of the routing layer. Just the removal of an artificial constraint I had installed in a moment of design conservatism.&lt;/p&gt;

&lt;p&gt;But the moment it was gone, something interesting happened. The system got quieter. Tasks that had been generating unnecessary coordination overhead resolved cleanly. The routing didn't need to pad formations anymore, so it stopped wasting compute on phantom collaborators. And the outputs on single-specialist tasks got sharper, because there was no second voice muddying them.&lt;/p&gt;




&lt;p&gt;The principle that came out of this is one I keep returning to: a minimum floor on a formation constraint doesn't protect you from degenerate cases. It creates them.&lt;/p&gt;

&lt;p&gt;When you say "at least two," you're not describing what teams actually need. You're describing what you imagined teams would look like before you watched them work. The real question isn't "how many agents should be minimum?" The real question is "what does this task actually require?" Sometimes that's five agents in a structured cascade. Sometimes it's one. The formation should follow the work, not the other way around.&lt;/p&gt;

&lt;p&gt;Dynamic team formation, real dynamic formation, means the system sizes the team to the task, not to a predetermined notion of what a team looks like. No floor. No ceiling. Just an honest answer to the question the task is actually asking.&lt;/p&gt;

&lt;p&gt;There's a broader instinct in that for me. We add minimums when we're afraid of what happens without them. But a lot of what we're afraid of is hypothetical. The real failure modes tend to live in the opposite direction, in the overhead you didn't see, the drift you didn't notice, the quiet inefficiency that accumulates across thousands of tasks because one constraint was never revisited.&lt;/p&gt;

&lt;p&gt;The system didn't need the floor. It needed the honesty to work without one.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>OAuth Client Credentials for EDI Pipelines</title>
      <dc:creator>praveenlavu</dc:creator>
      <pubDate>Fri, 21 Aug 2026 07:08:24 +0000</pubDate>
      <link>https://dev.to/praveenlavu/oauth-client-credentials-for-edi-pipelines-2p97</link>
      <guid>https://dev.to/praveenlavu/oauth-client-credentials-for-edi-pipelines-2p97</guid>
      <description>&lt;h1&gt;
  
  
  Securing the Invisible Pipeline: OAuth 2.0 Client Credentials and the EDI Authentication Problem
&lt;/h1&gt;

&lt;p&gt;Healthcare data moves through a hidden layer most people never think about. Between the system that submits a claim and the payer that adjudicates it sits an intermediary that speaks a language older than the modern web. EDI transactions, the X12-formatted messages for claims, eligibility checks, and remittances, have been the backbone of healthcare administration for decades. And for most of that time, the question of how the systems sending those transactions proved their identity to clearinghouses was answered with something embarrassingly simple: a username and a password.&lt;/p&gt;

&lt;p&gt;I spent real time staring at that problem. Not in the abstract, but in the actual mechanics of it. When you are building an integration in Salesforce that reaches out to a clearinghouse to submit claims or check eligibility in real time, the authentication question is not optional. You have to answer it. And the answer that most legacy systems had settled on, a shared credential stored somewhere, accessed by something, renewed when someone remembers, felt wrong the moment I really looked at it.&lt;/p&gt;

&lt;p&gt;The moment you look closely enough at something you have accepted as normal, you cannot unsee it. That is where this started.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Static Credential Problem
&lt;/h2&gt;

&lt;p&gt;There is a category of security risk that hides behind familiarity. Static credentials for machine-to-machine connections fall squarely into that category. You receive a username and password from a clearinghouse. You store it somewhere. Your integration uses it every time it makes a connection. It does not expire unless someone decides to rotate it. It carries no scope at all; it is either valid or it is not. If it leaks, nothing in the system alerts you that someone else is now using it alongside you.&lt;/p&gt;

&lt;p&gt;The EDI context makes this worse, not better. The systems involved are not consumer-facing. There is no human in the authentication loop. Nobody logs in, nobody checks a notification on their phone. The connection is purely machine to machine, running on a schedule or on demand, often overnight, often processing in bulk. A static credential in that context is a standing invitation: anyone with access to the configuration file, the environment variable, or the managed package setting has full, unscoped, time-unlimited access to whatever the clearinghouse permits on the account.&lt;/p&gt;

&lt;p&gt;In Salesforce, this materialized in a specific and uncomfortable way. Named Credentials are the standard mechanism for storing endpoint authentication. For clearinghouses that issued traditional username-and-password credentials, those credentials lived inside a Named Credential. They did not automatically rotate. The scope of what they could do was not encoded in the credential itself; it was implicit, determined by what the clearinghouse permitted at the account level. You had to trust that the setup was right and that nobody had silently changed it.&lt;/p&gt;

&lt;p&gt;That is not a security posture. It is a held breath. And I had been holding it long enough that it had started to feel like normal.&lt;/p&gt;

&lt;p&gt;The failure mode I kept thinking about was not the dramatic one. It was the quiet one. Not a breach announcement, not a forensic investigation. Just a credential that had been in place for years, that had changed hands as teams changed, that existed in three places nobody had fully mapped, and that nobody had rotated because the last person who knew how to rotate it had left. The threat model is not always an adversary. Sometimes it is just entropy.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Client Credentials Flow Actually Changes
&lt;/h2&gt;

&lt;p&gt;OAuth 2.0's client credentials grant was designed for exactly this scenario. Two machines need to communicate. Neither is acting on behalf of a human. There is no authorization code flow, no redirect URI, no user consent screen. There is a client identifier, a client secret, and a token endpoint. The machine that wants to act requests an access token, receives one with a defined expiration window, uses it for the duration of that window, and then requests another.&lt;/p&gt;

&lt;p&gt;What changes in practice is more significant than the mechanics first suggest.&lt;/p&gt;

&lt;p&gt;The token is short-lived by design. When you rotate credentials in the static model, you are replacing a key that has been sitting in a lock for months or years. In the client credentials model, the access token is already expiring constantly. Its existence is inherently temporary. The client secret that generates it still requires protection, but the blast radius of a leaked token is bounded by its remaining lifetime. A secret that expires in minutes causes a different kind of incident than one that is valid indefinitely.&lt;/p&gt;

&lt;p&gt;The token carries scope. When a clearinghouse issues a token in response to a client credentials request, they can encode what that token is permitted to do. An integration that checks eligibility does not need a token with claims-submission permissions. Scope constraints mean that a compromised token in one part of the pipeline does not automatically endanger another. The principle of least privilege stops being a policy statement and becomes a runtime property of the credential itself.&lt;/p&gt;

&lt;p&gt;The credential negotiation happens at the boundary, and only there. In Salesforce, when you configure a Named Credential backed by an OAuth 2.0 client credentials flow, the platform handles the token lifecycle. It requests, receives, caches, and refreshes tokens. Your integration calls the Named Credential; it does not manage tokens directly. The client secret is not passed inline on every API call. Authentication is a separate, managed concern, invisible to the business logic layer.&lt;/p&gt;

&lt;p&gt;That separation matters more than it sounds. The places where credentials tend to surface unexpectedly are the places where they are being actively used: in log output, in error messages, in request captures during debugging. A static password that travels on every API call has repeated opportunities to appear somewhere it should not. A token that was requested once and cached travels on the call, but the secret that generated it stayed at the token endpoint. The exposure surface is structurally smaller.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Integration Reality
&lt;/h2&gt;

&lt;p&gt;Making this work in a Salesforce environment, connecting to clearinghouses that have historically operated on older authentication models, involves navigating a real transition. Not every clearinghouse offers OAuth 2.0 endpoints. Some that do have implemented them in ways that require precise configuration. The token endpoint, the scope parameter, the grant type declaration: these have to match what the clearinghouse actually issues, not what their documentation suggests they issue. Documentation and implementation are not always the same thing. I have learned that more than once.&lt;/p&gt;

&lt;p&gt;There is also the question of failure behavior at runtime. The static credential model has a certain blunt reliability: either the credential works or it does not, and the failure mode is usually obvious. Token acquisition failures in OAuth can be more nuanced. A misconfigured scope, an expired client secret, a rate limit on the token endpoint: these fail in ways that require understanding the distinction between an authentication failure and an authorization failure. Building in the right error handling and the right observability means knowing where in the flow things can go wrong before they go wrong in production.&lt;/p&gt;

&lt;p&gt;What I found, after working through the configuration and the edge cases, was that the complexity is front-loaded. Getting the Named Credential configured correctly, getting the token endpoint and parameters to match the clearinghouse's specific OAuth implementation, understanding the subtleties of their scoping model: that is where the real work lives. Once it is in place, the runtime behavior is more reliable and more transparent than the static credential pattern it replaced. The platform handles renewal. The integration handles the business logic. The concerns stay separated.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Principle at the Boundary
&lt;/h2&gt;

&lt;p&gt;Every integration has a trust boundary. The boundary between a Salesforce org and a clearinghouse is one of those places where two systems, operating under different ownership and different control, have to agree on identity. How they negotiate that agreement determines the security properties of everything that flows across it.&lt;/p&gt;

&lt;p&gt;Static credentials are an informal agreement. We both know the password, so we trust each other. That informality has real costs. When something goes wrong, when a credential leaks, when access needs to be revoked, when you need to reconstruct who did what and when, the informal agreement provides no tools. You are left auditing logs and hoping.&lt;/p&gt;

&lt;p&gt;OAuth 2.0 client credentials are a formal protocol. The agreement is structured. The tokens are bounded. The scopes are declared. The trust is established dynamically, not assumed statically. Every token request is a moment of explicit verification, not an ongoing assumption.&lt;/p&gt;

&lt;p&gt;The EDI pipeline is invisible to most of the people whose data flows through it. The authentication protecting that pipeline does not have to be invisible to the engineers building it. It can be deliberate, auditable, and constrained to the minimum required. That is the shift worth internalizing: not from simple to complex, but from implicit to explicit. From assumed to verified. From a held breath to actual confidence in what you built.&lt;/p&gt;

&lt;p&gt;That confidence is worth taking the time to earn.&lt;/p&gt;

</description>
      <category>api</category>
      <category>architecture</category>
      <category>backend</category>
      <category>security</category>
    </item>
  </channel>
</rss>
