<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: The Unmeshed Team</title>
    <description>The latest articles on DEV Community by The Unmeshed Team (@unmeshed).</description>
    <link>https://dev.to/unmeshed</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3953781%2F5af20b0f-7da7-4490-9183-a2d28dab3978.png</url>
      <title>DEV Community: The Unmeshed Team</title>
      <link>https://dev.to/unmeshed</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/unmeshed"/>
    <language>en</language>
    <item>
      <title>Your API Key Shouldn't Show Up in Your Logs</title>
      <dc:creator>The Unmeshed Team</dc:creator>
      <pubDate>Fri, 09 Oct 2026 06:10:03 +0000</pubDate>
      <link>https://dev.to/unmeshed/your-api-key-shouldnt-show-up-in-your-logs-3hc5</link>
      <guid>https://dev.to/unmeshed/your-api-key-shouldnt-show-up-in-your-logs-3hc5</guid>
      <description>&lt;p&gt;A token goes into a webhook request like any other field. A week later, someone's digging through logs for a totally different bug, and there it is: a plaintext API key, just sitting in the input trace.&lt;/p&gt;

&lt;p&gt;Nobody did that on purpose. It's just what happens when a secret and a regular input get treated the same way.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual problem
&lt;/h2&gt;

&lt;p&gt;A webhook call to trigger a workflow usually looks like any other POST request: a JSON body with whatever inputs the process needs. If one of those inputs happens to be a credential or an access token, it rides along in the same body as everything else, and by default, everything in that body can end up visible in logs, process summaries, or API responses. None of that's malicious. It's just what "log the input" means if nothing tells the system to treat one field differently.&lt;/p&gt;

&lt;p&gt;The fix isn't "be more careful with logging." It's giving secrets their own lane from the start.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping secrets out of the trail
&lt;/h2&gt;

&lt;p&gt;A webhook call can separate regular inputs from secret ones using a dedicated field, something like &lt;code&gt;_secretStatePut&lt;/code&gt;, alongside the normal payload:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"https://your-instance/webhook/trigger-id"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "key1": "value1",
    "_secretStatePut": {
      "apiToken": "sk-live-xxxxxxxxxxxx"
    }
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Anything inside that object gets handled differently from the rest of the payload:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It doesn't show up in logs or process summaries, unless a step explicitly outputs it (which is on the workflow author, not the platform).&lt;/li&gt;
&lt;li&gt;It's not echoed back in API responses.&lt;/li&gt;
&lt;li&gt;It's only reachable inside the running workflow's own execution context, not from the outside.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Regular inputs still work exactly like before. The only thing that changes is where the sensitive values go.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this actually matters
&lt;/h2&gt;

&lt;p&gt;Anywhere a workflow needs to reach something it shouldn't have to show its work on. Triggering a deployment pipeline with environment-specific credentials. Kicking off a data ingestion job that needs to hit a protected API. Any event-driven automation where "what credential did this run use" is a question that should have an answer, just not one sitting in a log file for whoever happens to be grepping through it.&lt;/p&gt;

&lt;p&gt;The real value isn't that secrets are hidden. It's that they stop being an afterthought, something a developer has to remember to scrub out of logs after the fact, and become something the platform just handles correctly by default.&lt;/p&gt;

&lt;p&gt;If a credential has ever ended up somewhere it shouldn't in your workflow logs, &lt;a href="https://unmeshed.io" rel="noopener noreferrer"&gt;Unmeshed&lt;/a&gt;'s secret input handling is built to make that the thing that doesn't happen in the first place.&lt;/p&gt;

</description>
      <category>security</category>
      <category>webdev</category>
      <category>backend</category>
      <category>api</category>
    </item>
    <item>
      <title>Your Nightly Batch Job Doesn't Need to Run Like It's 2005</title>
      <dc:creator>The Unmeshed Team</dc:creator>
      <pubDate>Thu, 08 Oct 2026 06:01:59 +0000</pubDate>
      <link>https://dev.to/unmeshed/your-nightly-batch-job-doesnt-need-to-run-like-its-2005-2k78</link>
      <guid>https://dev.to/unmeshed/your-nightly-batch-job-doesnt-need-to-run-like-its-2005-2k78</guid>
      <description>&lt;p&gt;Picture a job that has to pull data from 25 different APIs every night, process the responses, and store the results for reporting. Nothing about it is hard. It's just a lot of small things that all have to happen, on time, without anyone watching.&lt;/p&gt;

&lt;p&gt;Most teams build it the obvious way: a script that calls API one, then API two, then API three, all the way to 25. It works. It's also the slowest possible way to do it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Most of the work doesn't depend on the rest
&lt;/h2&gt;

&lt;p&gt;Those 25 API calls don't need each other. Call 14 doesn't care what call 3 returned. So there's no reason to make them wait in line.&lt;/p&gt;

&lt;p&gt;Run independent calls in parallel and the total time drops from ""the sum of every call"" to roughly ""the slowest single call."" In our test, we modeled the job as five groups of five parallel HTTP steps, 25 calls in total. The whole thing finished in under 100 milliseconds.&lt;/p&gt;

&lt;p&gt;That's a test against fast endpoints, so don't expect those numbers against slow third-party APIs. The point is the shape: if steps are independent, run them side by side.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scheduling is where jobs quietly go wrong
&lt;/h2&gt;

&lt;p&gt;Once the job works, you need it to run on its own. That's usually a cron expression, and the standard five-field syntax covers most needs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0 * * * *      # every hour, on the hour
0 2 * * 1      # every Monday at 2 AM
0 0 1 * *      # first day of every month at midnight
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cron is the easy part. The part people forget is what happens when a run takes longer than expected.&lt;/p&gt;

&lt;h2&gt;
  
  
  The overlap problem
&lt;/h2&gt;

&lt;p&gt;Say your job normally takes ten minutes, but one night an API is slow and it takes ninety. If the schedule fires again in the meantime, you now have two copies running at once. Two copies can write conflicting data, double-count records, or hammer the same APIs and make things slower.&lt;/p&gt;

&lt;p&gt;The fix is an overlap policy. Setting the schedule to not allow overlap means a new run won't start while the previous one is still going. It's one setting, and it prevents a very annoying class of bugs.&lt;/p&gt;

&lt;h2&gt;
  
  
  You should be able to see what ran
&lt;/h2&gt;

&lt;p&gt;A batch job that runs at 2 AM has one big weakness: nobody is awake to notice when something breaks. So it helps a lot to be able to open a list of past executions and see which ones ran, how long they took, and whether every step finished.&lt;/p&gt;

&lt;p&gt;In our demo, the scheduled run triggered on time, ran all 25 steps, and completed in 55 milliseconds. More useful than the speed was that we could see all of it afterward.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;For any batch job that fans out to many sources:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Run independent steps in parallel instead of one after another&lt;/li&gt;
&lt;li&gt;Schedule with cron, and decide up front what should happen when a run overlaps the last one&lt;/li&gt;
&lt;li&gt;Make sure you can look back and see what actually happened&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Get those three right and a nightly job becomes boring, which is exactly what you want from a nightly job."&lt;/p&gt;

</description>
      <category>automation</category>
      <category>api</category>
      <category>devops</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Why SharePoint Automation Usually Stalls Right When You Need It Most</title>
      <dc:creator>The Unmeshed Team</dc:creator>
      <pubDate>Wed, 07 Oct 2026 06:32:02 +0000</pubDate>
      <link>https://dev.to/unmeshed/why-sharepoint-automation-usually-stalls-right-when-you-need-it-most-204j</link>
      <guid>https://dev.to/unmeshed/why-sharepoint-automation-usually-stalls-right-when-you-need-it-most-204j</guid>
      <description>&lt;p&gt;SharePoint has a weird relationship with files. It's happy to store them forever, organize them into seventeen nested folders, and let everyone fight over who has edit access. What it's a lot less happy to do is actually &lt;em&gt;do&lt;/em&gt; something with them the moment they land.&lt;/p&gt;

&lt;p&gt;And that's usually fine, until someone uploads a form and three people are now manually copying numbers out of it into a spreadsheet, which somewhat defeats the purpose of having a document management system in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gap between "stores files" and "triggers workflows"
&lt;/h2&gt;

&lt;p&gt;SharePoint was built as a document and collaboration platform, not a workflow engine. Its native automation tools can react to simple events, a file uploaded, a value changed, but they struggle the moment a process needs more than one step: pulling data out of a file, validating it, routing it somewhere, and tracking whether each of those steps actually succeeded.&lt;/p&gt;

&lt;p&gt;That gap is usually invisible until a team tries to build something slightly more ambitious than "send a notification when a file changes." Then it becomes the whole problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  What real SharePoint automation actually requires
&lt;/h2&gt;

&lt;p&gt;A handful of genuinely common use cases expose exactly where native tooling runs out of road:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Processing uploaded files, not just noticing them.&lt;/strong&gt; A PDF form gets uploaded, now something needs to extract the data from it and store it somewhere. That's not an event trigger anymore, it's a pipeline with steps that can fail independently.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reacting to spreadsheet changes, reliably.&lt;/strong&gt; Someone updates a shared Excel file, and the new values need to sync to a database or data warehouse. Simple in concept, but it needs to handle partial failures without silently dropping data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auditing changes over time, not just the last one.&lt;/strong&gt; A daily digest of file changes across a SharePoint site, for security or compliance, needs something that runs on a schedule and keeps a record, not just a one-off trigger.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pushing generated output back in.&lt;/strong&gt; Reports or exports created elsewhere need to land in the right SharePoint folder automatically, which means the automation has to work in both directions, not just react to SharePoint.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Letting people browse SharePoint through a different interface.&lt;/strong&gt; Sometimes the actual need isn't automation at all, it's a custom way to surface SharePoint data without forcing everyone into SharePoint's own UI.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these are exotic requests. They're the kind of thing almost any team with SharePoint ends up wanting eventually.&lt;/p&gt;

&lt;h2&gt;
  
  
  The common thread
&lt;/h2&gt;

&lt;p&gt;Every one of these needs the same underlying things: a trigger, one or more processing steps, error handling that doesn't fail silently, and a record of what happened. That's not a SharePoint feature gap you patch with another native automation. It's a workflow orchestration problem, and it needs to be treated as one, connecting SharePoint to a system built to handle multi-step processes, retries, and visibility into what ran and what didn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Worth checking before you build around it
&lt;/h2&gt;

&lt;p&gt;If a SharePoint automation need stops at "when X happens, do one simple thing," native tooling is probably fine. The moment it involves extracting data, branching logic, or multiple systems downstream, it's worth asking whether the automation layer was built for that, or whether it's about to become another fragile workaround.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>automation</category>
      <category>sharepoint</category>
      <category>workflow</category>
    </item>
    <item>
      <title>Looking Beyond LangGraph? 6 AI Agent Frameworks and What Each One Actually Fixes</title>
      <dc:creator>The Unmeshed Team</dc:creator>
      <pubDate>Tue, 06 Oct 2026 05:44:27 +0000</pubDate>
      <link>https://dev.to/unmeshed/looking-beyond-langgraph-6-ai-agent-frameworks-and-what-each-one-actually-fixes-5878</link>
      <guid>https://dev.to/unmeshed/looking-beyond-langgraph-6-ai-agent-frameworks-and-what-each-one-actually-fixes-5878</guid>
      <description>&lt;h2&gt;
  
  
  TLDR
&lt;/h2&gt;

&lt;p&gt;Most "LangGraph alternatives" content right now still recommends AutoGen. Microsoft put it in maintenance mode in October 2025.&lt;/p&gt;

&lt;p&gt;The real current pick is Microsoft Agent Framework, GA since April 2026, merging AutoGen and Semantic Kernel into one SDK.&lt;/p&gt;

&lt;p&gt;Five other picks, each solving a different specific reason people leave LangGraph: CrewAI (verbosity), OpenAI Agents SDK (OpenAI-native), Mastra (TypeScript), Temporal (durability), Unmeshed (governance layer).&lt;/p&gt;

&lt;p&gt;Temporal and Unmeshed aren't agent frameworks at all; they wrap whatever framework you pick rather than compete with it.&lt;/p&gt;

&lt;p&gt;Pick based on your actual constraint, not whichever name shows up most in search results.&lt;/p&gt;

&lt;p&gt;Search "LangGraph alternatives" right now, and most of what comes back still puts AutoGen near the top of the list. That's outdated advice.&lt;/p&gt;

&lt;p&gt;Microsoft moved AutoGen into maintenance mode on October 2, 2025. No new features, no enhancements, just bug fixes and security patches going forward.&lt;/p&gt;

&lt;p&gt;The actual successor, Microsoft Agent Framework, reached general availability in April 2026. It merges AutoGen's orchestration patterns with Semantic Kernel's enterprise tooling into one SDK.&lt;/p&gt;

&lt;p&gt;Even guides published after that date have skipped it, usually because it was still too new to trust. Five months later, it isn't.&lt;/p&gt;

&lt;p&gt;So this list starts by fixing that, then walks the other five LangGraph alternatives by the actual reason people leave LangGraph in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Why Teams Look Past LangGraph
&lt;/h2&gt;

&lt;p&gt;LangGraph earned its adoption. Explicit state, real branching, and built-in checkpointing are exactly what a graph-shaped agent workflow needs. It's one of several agent orchestration frameworks built around that idea, but it isn't the only shape that works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftj87owmybu4z1vam72r8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftj87owmybu4z1vam72r8.png" alt="LangGraph alternatives evaluation highlighting workflow verbosity, paid observability pricing, ecosystem coupling, TypeScript limitations, and breaking changes that lead teams to compare CrewAI, Microsoft Agent Framework, Mastra, Temporal, and OpenAI Agents SDK." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The reasons teams look at LangGraph competitors are specific, not vague dissatisfaction:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Verbosity for small workloads&lt;/strong&gt;: A two-step agent needs real graph declaration, nodes, edges, a compiler. Past a certain simplicity threshold, the graph is paperwork.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LangSmith as the paid observability path&lt;/strong&gt;: Pricing starts at $39 per seat on the Plus tier and climbs fast with trace volume. Self-hosted alternatives exist but are less polished.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ecosystem coupling&lt;/strong&gt;: LangGraph ships from the same team as LangChain and shares primitives, so importing both compounds the exact upgrade tax teams left LangChain to escape.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Python-first&lt;/strong&gt;: The JavaScript port trails Python by months on features and documentation, leaving TypeScript teams in a sidecar pattern.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frequent breaking changes&lt;/strong&gt;: The interrupt API has shifted more than once, and checkpoint formats have broken programs across minor versions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this makes LangGraph the wrong choice for a genuinely graph-shaped workflow. It means a real range of LangGraph alternatives exists, depending on what actually bites for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. What Should You Look for in a LangGraph Alternative?
&lt;/h2&gt;

&lt;p&gt;Before comparing tools, it helps to know what you're actually optimizing for. Not every LangGraph alternative solves the same problem, so match the tool to the constraint that's actually yours.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Language: Python-first, or does your team need real TypeScript support without a sidecar process?&lt;/li&gt;
&lt;li&gt;Durability: Does the agent need to survive a crash and resume mid-task, or is it short-lived and stateless between runs?&lt;/li&gt;
&lt;li&gt;Cost shape: A free open-source core with usage-based cloud pricing, or a flat enterprise license?&lt;/li&gt;
&lt;li&gt;Governance: Do you need audit logging, approvals, and access control built in, or is that being solved somewhere else in your stack?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Answer these four honestly, and the right pick from this list gets a lot more obvious.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The 6 Best LangGraph Alternatives
&lt;/h2&gt;

&lt;p&gt;This is the LangGraph alternatives shortlist that actually holds up in late 2026, not the one still repeating 2025 assumptions about who's actively maintained.&lt;/p&gt;

&lt;p&gt;Here's the LangGraph vs alternatives breakdown, each one picked for a different specific reason to leave, not ranked by popularity:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Framework&lt;/th&gt;
&lt;th&gt;What It Fixes&lt;/th&gt;
&lt;th&gt;License / Pricing&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CrewAI&lt;/td&gt;
&lt;td&gt;Verbosity, role-based crews instead of graph declarations&lt;/td&gt;
&lt;td&gt;Free (50 executions/mo); Enterprise custom&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Microsoft Agent Framework&lt;/td&gt;
&lt;td&gt;Replaces maintenance-mode AutoGen with the current supported path&lt;/td&gt;
&lt;td&gt;Open source, MIT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI Agents SDK&lt;/td&gt;
&lt;td&gt;OpenAI-native production runtime, minimal ceremony&lt;/td&gt;
&lt;td&gt;Open source; usage-priced via OpenAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mastra&lt;/td&gt;
&lt;td&gt;TypeScript-first agents, no Python sidecar&lt;/td&gt;
&lt;td&gt;Apache 2.0 core; Enterprise add-ons&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Temporal&lt;/td&gt;
&lt;td&gt;Durability for long-running, crash-surviving agents&lt;/td&gt;
&lt;td&gt;Self-hosted free; Cloud from $50/million actions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unmeshed&lt;/td&gt;
&lt;td&gt;Orchestration, durability, and governance around any framework&lt;/td&gt;
&lt;td&gt;Free forever; Premium $20/mo&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  1. CrewAI
&lt;/h3&gt;

&lt;p&gt;CrewAI shows up on nearly every langgraph alternatives list for good reason. Its model is roles and tasks, not nodes and edges. Agent, Task, Crew, and Process cover most fixed-sequence pipelines in a third of the code a state graph needs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9spbnae43lskunydres3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9spbnae43lskunydres3.png" alt="CrewAI agent framework for multi-agent workflows, role-based AI agents, enterprise agent orchestration, autonomous agents, and LangGraph alternative evaluations." width="799" height="353"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Readable role-based syntax: researcher, writer, reviewer, mapped directly to the code&lt;/li&gt;
&lt;li&gt;Used by 65% of the Fortune 500, per CrewAI's own pricing page&lt;/li&gt;
&lt;li&gt;Free tier covers 50 workflow executions a month; Enterprise adds SSO, RBAC, and dedicated VPC or on-prem deployment&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Real tradeoff:&lt;/strong&gt; workflows that are genuinely graph-shaped, with real branching and persisted state, fight the roles-and-tasks abstraction instead of fitting it.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Microsoft Agent Framework
&lt;/h3&gt;

&lt;p&gt;Most langgraph alternatives roundups either skip this pick entirely or still lead with AutoGen. This is the corrected pick on this list. AutoGen moved to maintenance mode on October 2, 2025, confirmed directly on Microsoft's own AutoGen repository. No new features; community-managed going forward.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgayuhpvddbnc4v8gfrxw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgayuhpvddbnc4v8gfrxw.png" alt="Microsoft Agent Framework open-source SDK for AI agent orchestration, agent workflows, AutoGen successor, Semantic Kernel integration, and enterprise LangGraph alternatives." width="800" height="355"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Microsoft Agent Framework reached general availability on April 3, 2026, merging AutoGen's orchestration patterns with Semantic Kernel's enterprise tooling into one open-source SDK.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Graph-based workflows: sequential, concurrent, handoff, and group collaboration patterns&lt;/li&gt;
&lt;li&gt;Checkpointing, streaming, human-in-the-loop, and time-travel debugging built in&lt;/li&gt;
&lt;li&gt;Built-in OpenTelemetry observability, no separate paid add-on required&lt;/li&gt;
&lt;li&gt;MIT-licensed, with direct migration guides from both AutoGen and Semantic Kernel in the repo&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Real tradeoff:&lt;/strong&gt; a shorter production track record than LangGraph itself, even though it inherits Semantic Kernel's enterprise maturity underneath.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. OpenAI Agents SDK
&lt;/h3&gt;

&lt;p&gt;A common pick among LangGraph alternatives for teams already all-in on OpenAI models. It's the production-ready successor to OpenAI's earlier Swarm experiment. Three primitives: Agents, Handoffs, and Guardrails, with built-in tracing out of the box.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3u3nk5hv2weay0k502bx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3u3nk5hv2weay0k502bx.png" alt="OpenAI Agents SDK framework for agent orchestration, agent handoffs, guardrails, tracing, and production AI applications compared with LangGraph alternatives." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;First-class in both Python and TypeScript, no second-class port&lt;/li&gt;
&lt;li&gt;A handoff() call transfers control between agents, simpler than a graph for specialist routing&lt;/li&gt;
&lt;li&gt;Open source SDK; cost comes from OpenAI usage, not a separate framework fee&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Real tradeoff:&lt;/strong&gt; OpenAI-aligned by design. Non-OpenAI providers work through adapter layers with a thinner surface than the native path.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Mastra
&lt;/h3&gt;

&lt;p&gt;Among TypeScript-first langgraph alternatives, Mastra is the most complete. It isn't a port of a Python framework. Agents, workflows, memory, and observability ship as part of the same package.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnu13ligcg6pvzaco6l05.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnu13ligcg6pvzaco6l05.png" alt="Mastra AI agent framework for TypeScript developers, agent orchestration, workflow automation, agent memory, observability, and modern LangGraph alternative evaluations." width="800" height="389"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Deployers built in for Vercel, Netlify, Cloudflare, or a standalone Hono server&lt;/li&gt;
&lt;li&gt;Real named customers on Mastra's own site: Salesforce runs a 100k-developer internal harness on it, MongoDB built an internal agent platform called Sage for 20+ engineering teams, and Range runs it for a $17B AUM investment advisory product&lt;/li&gt;
&lt;li&gt;Core framework is Apache 2.0; Enterprise features sit under a separate source-available license&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Real tradeoff:&lt;/strong&gt; a younger ecosystem than LangGraph or CrewAI, and no benefit at all for Python-heavy teams.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Temporal
&lt;/h3&gt;

&lt;p&gt;Temporal rarely shows up on LangGraph alternatives lists because it isn't an agent framework at all. It's a durable execution engine that wraps any agent code, including LangGraph itself, and records full event history so a workflow can crash and resume exactly where it left off.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo2qhgahiwygj4m3l7uhm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo2qhgahiwygj4m3l7uhm.png" alt="Temporal durable execution engine for AI agents, workflow orchestration, fault tolerance, long-running workflows, and production-ready alternatives to LangGraph persistence." width="800" height="243"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Explicitly supports running the OpenAI Agents SDK and Google ADK as durable Activities, per Temporal's own pricing page&lt;/li&gt;
&lt;li&gt;Self-hosted is free and MIT-licensed; Temporal Cloud starts at $50 per million actions, with volume discounts down to $25 per million&lt;/li&gt;
&lt;li&gt;Backed by real momentum: a $12.55B Series E, with Temporal's own newsroom citing AI-driven demand for durable execution as the driver&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Real tradeoff:&lt;/strong&gt; adds a second runtime and operational surface. Duration alone doesn't justify it; the workflow needs a genuine durability requirement, like surviving crashes or running for days.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Unmeshed
&lt;/h3&gt;

&lt;p&gt;Not a reasoning framework, and it doesn't compete with LangGraph, CrewAI, or Microsoft Agent Framework on agent logic. That's not the comparison to make. It sits underneath any of the five frameworks above, not instead of them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feht0qqwaawrb50ucg6hj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feht0qqwaawrb50ucg6hj.png" alt="Unmeshed AI workflow orchestration platform for agentic AI, durable execution, workflow governance, human-in-the-loop approvals, API orchestration, and production AI automation." width="800" height="419"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The gap it fills
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;CrewAI, Microsoft Agent Framework, and the OpenAI Agents SDK all handle how an agent reasons, nothing more&lt;/li&gt;
&lt;li&gt;Mastra adds workflows and memory, but still no durable execution, approvals, or audit trail by default&lt;/li&gt;
&lt;li&gt;Temporal solves durability, but it's a separate runtime you have to learn and operate on top of whichever framework you already picked&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What Unmeshed solves, on the free plan
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Durable execution, so a failed step resumes instead of restarting the whole run&lt;/li&gt;
&lt;li&gt;Human-in-the-loop approvals and a decision engine for branching logic, no separate approval tool bolted on&lt;/li&gt;
&lt;li&gt;Audit logging on every step as part of built-in AI agent governance, so there's a real, queryable answer to what an agent actually did and when&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The practical difference
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;The other five change how you write the agent; Unmeshed changes whether that agent survives contact with production&lt;/li&gt;
&lt;li&gt;No second vendor and no second learning curve just to get durability&lt;/li&gt;
&lt;li&gt;Free forever, Premium at $20 a month, for a team weighing LangGraph alternatives, that's the actual cost comparison&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the framework choice is settled and the actual gap is durability, approvals, or governance, that's a different layer entirely, covered in more depth in what is AI agent infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flt6wvmurpvy5rzifl867.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flt6wvmurpvy5rzifl867.png" alt="AI agent infrastructure diagram showing Unmeshed adding durable execution, governance, approvals, and audit logging underneath agent frameworks such as LangGraph, CrewAI, OpenAI Agents SDK, Microsoft Agent Framework, and Mastra." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;h3&gt;
  
  
  Your Framework Picks the Agent. Something Else Runs It.
&lt;/h3&gt;

&lt;p&gt;Whichever of these six you land on, see what happens when durability, approvals, and audit logging run underneath it instead of getting bolted on later.&lt;br&gt;
&lt;a href="https://unmeshed.io/signup?utm_source=blog&amp;amp;utm_medium=organic&amp;amp;utm_campaign=langgraph_alternatives&amp;amp;utm_content=accent_cta" rel="noopener noreferrer"&gt;See The Layer Underneath&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. How Do LangGraph Alternatives Handle Long-Running Agents?
&lt;/h2&gt;

&lt;p&gt;Most don't, and that's the gap worth knowing about before you pick one. LangGraph's own persistence is real but lightweight, and the same is true for CrewAI, the OpenAI Agents SDK, Microsoft Agent Framework, and Mastra: none of them are built to survive a process crash mid-run or resume an agent that's been paused for hours or days.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy18h8o7d6u5seiprmahz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy18h8o7d6u5seiprmahz.png" alt="Comparison of AI agent frameworks including CrewAI, Microsoft Agent Framework, OpenAI Agents SDK, and Mastra, showing Temporal and Unmeshed as infrastructure layers that wrap agent frameworks rather than replace them." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two of the six LangGraph alternatives on this list actually solve that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Temporal records full event history and replays deterministically, so a workflow resumes exactly where it left off after a crash. It wraps existing agent code, including the OpenAI Agents SDK and Google ADK, but that means running and operating a second system alongside whichever framework you picked.&lt;/li&gt;
&lt;li&gt;Unmeshed builds durable execution directly into the same orchestration layer that also handles human-in-the-loop approvals, decisioning, and audit logging. A failed step resumes automatically, on the free plan, with no separate durability engine to license, deploy, or operate.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your agents run for minutes and finish cleanly, this isn't your problem. If they run for hours, wait on human input, or need to survive a bad deploy, it's the first thing to check before picking a framework.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Are There Open Source LangGraph Alternatives?
&lt;/h2&gt;

&lt;p&gt;Yes, most of them. Licensing varies more than people expect, though, so it's worth checking before you build on top of one.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CrewAI&lt;/strong&gt;: free tier (50 executions/month), Enterprise is custom and closed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Microsoft Agent Framework&lt;/strong&gt;: fully open source, MIT licensed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI Agents SDK&lt;/strong&gt;: open source SDK, though cost comes from OpenAI usage, not the framework itself&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mastra&lt;/strong&gt;: core framework is Apache 2.0, Enterprise features sit under a separate source-available license&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Temporal&lt;/strong&gt;: self-hosted core is free and MIT-licensed; Temporal Cloud is usage-priced&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unmeshed&lt;/strong&gt;: free forever on the base plan; Premium is a flat $20/month&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If open source with no vendor lock-in is the priority, Microsoft Agent Framework and Temporal's self-hosted option are the cleanest picks on this list, both fully MIT with no feature gate behind a paid tier.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Which One Actually Fits
&lt;/h2&gt;

&lt;p&gt;Skip the ranked list. This isn't a popularity-contest kind of AI agent framework comparison; it's matched to the specific reason you're actually looking:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Your Situation&lt;/th&gt;
&lt;th&gt;Pick&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fixed-sequence pipeline of specialists, want less code than a graph&lt;/td&gt;
&lt;td&gt;CrewAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Building on Microsoft Foundry or Azure, want the current supported path&lt;/td&gt;
&lt;td&gt;Microsoft Agent Framework&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mostly OpenAI models, want tracing and guardrails out of the box&lt;/td&gt;
&lt;td&gt;OpenAI Agents SDK&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TypeScript team, deploying to Vercel or Cloudflare&lt;/td&gt;
&lt;td&gt;Mastra&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agents need to survive crashes and run for days&lt;/td&gt;
&lt;td&gt;Temporal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Framework is fine; the pain is durability, approvals, or audit underneath it&lt;/td&gt;
&lt;td&gt;Unmeshed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;The single most common mistake in LangGraph alternative advice right now is recommending AutoGen without mentioning it's in maintenance mode.&lt;/p&gt;

&lt;p&gt;Past that correction, the rest of these LangGraph alternatives are genuinely differentiated, not just renamed versions of each other. The choice comes down to which specific limitation is yours: verbosity, cost, language, durability, or the governance layer underneath the framework entirely.&lt;/p&gt;

&lt;p&gt;Pick based on the actual reason you're looking, not the most popular name in the search results.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>langgraph</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Cron Doesn't Know It's a Holiday</title>
      <dc:creator>The Unmeshed Team</dc:creator>
      <pubDate>Mon, 05 Oct 2026 05:55:15 +0000</pubDate>
      <link>https://dev.to/unmeshed/cron-doesnt-know-its-a-holiday-6d1</link>
      <guid>https://dev.to/unmeshed/cron-doesnt-know-its-a-holiday-6d1</guid>
      <description>&lt;p&gt;We had a job scheduled to run every weekday at 9 AM. Cron handled that part fine, right up until Thanksgiving, when it ran anyway, into an empty office, and nobody noticed the output was garbage until two days later.&lt;/p&gt;

&lt;p&gt;Cron is great at one thing: fixed intervals. Every hour, every day at 5 AM, every Monday. What it has no concept of is a holiday, a blackout window, or "the first working day after a long weekend." Those aren't interval problems. They're calendar problems, and cron was never built to answer them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a calendar actually buys you
&lt;/h2&gt;

&lt;p&gt;The fix isn't a smarter cron expression, there isn't one. It's separating "when does this job normally run" from "is today actually a valid day to run it," and letting the second question filter the first. That's what calendar scheduling does: a calendar sits on top of a schedule and decides, date by date, whether the job fires.&lt;/p&gt;

&lt;p&gt;There are a few shapes this takes, depending on who's setting the rule and how complex it needs to be.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pick the dates by hand.&lt;/strong&gt; For a monthly review or a quarterly task, someone just clicks the valid days on a visual calendar, or applies a quick pattern like "every weekday" or "every Monday." No code, no logic, just a list of dates that's easy for a non-technical person to manage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write the logic.&lt;/strong&gt; "Every third Friday of the month" or "the first working day after a public holiday" isn't something you pick by hand, it's something you compute. A code-based calendar lets a developer define that rule once instead of manually updating a date list every year.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Combine calendars.&lt;/strong&gt; Sometimes the rule you actually want is the relationship between two calendars, not either one alone. Working days minus public holidays. The overlap between two teams' on-call windows. Set operations, union, intersection, difference, handle this without hand-maintaining a third list that has to stay in sync with the other two.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reschedule instead of skip.&lt;/strong&gt; Some jobs can't just not run on a holiday, they need to run later. A reschedule calendar skips the blackout days and pushes the job out by a defined number of days instead of dropping it entirely, useful for anything that absolutely has to happen, just not today.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters more than it sounds like it should
&lt;/h2&gt;

&lt;p&gt;The real cost of the Thanksgiving-cron story isn't the one bad run. It's that nobody could say, with confidence, which other jobs had the same blind spot, because the holiday logic, if it existed at all, was scattered across whatever script each team happened to write. A calendar makes that rule explicit and reusable: define it once, attach it to any schedule that needs it, and get a preview of every date the job will actually run before it goes live.&lt;/p&gt;

&lt;p&gt;That last part matters more than it seems. Scheduling logic is the kind of thing that's nearly impossible to verify by reading code, you want to see the actual dates, not trust that the expression is right. A snapshot preview turns "this should be correct" into "this is visibly correct," which is a meaningfully different level of confidence before something goes live in production.&lt;/p&gt;

&lt;p&gt;If cron has ever run a job on a day it really shouldn't have, &lt;a href="https://unmeshed.io" rel="noopener noreferrer"&gt;Unmeshed&lt;/a&gt;'s calendar scheduling is built to be the layer that catches it before it happens, not after.&lt;/p&gt;

</description>
      <category>backend</category>
      <category>automation</category>
      <category>scheduling</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Your Wait Step Works Once. Then It Stops Waiting</title>
      <dc:creator>The Unmeshed Team</dc:creator>
      <pubDate>Thu, 01 Oct 2026 05:40:43 +0000</pubDate>
      <link>https://dev.to/unmeshed/your-wait-step-works-once-then-it-stops-waiting-19o</link>
      <guid>https://dev.to/unmeshed/your-wait-step-works-once-then-it-stops-waiting-19o</guid>
      <description>&lt;p&gt;We set a Wait step to pause three seconds inside a loop. First time through, perfect. Second time through, it just stops waiting. Like it forgot its one job.&lt;/p&gt;

&lt;p&gt;Here's the twist: nothing's broken. The engine's doing exactly what we told it to do. We just didn't realize what we were actually telling it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it happens
&lt;/h2&gt;

&lt;p&gt;A Wait step takes a function that returns a timestamp, &lt;code&gt;waitUntil&lt;/code&gt;, and most people write it as an offset from when the step started:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;steps&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;startDate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;steps&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;__self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;start&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;waitUntil&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;startDate&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getTime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="nx"&gt;_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// wait 3 seconds&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;start&lt;/code&gt; value is set once, the first time the step runs, and it never updates. On the first pass through the loop, that's fine, the timestamp is three seconds in the future and the engine waits as expected.&lt;/p&gt;

&lt;p&gt;On the second pass, the While loop re-runs the same Wait step instance. It still reads the original &lt;code&gt;start&lt;/code&gt; value, which is now well in the past. As far as the engine's concerned, &lt;code&gt;waitUntil&lt;/code&gt; has already been satisfied, so it continues immediately. It's not skipping the wait. It's correctly honoring a timestamp that's already expired.&lt;/p&gt;

&lt;p&gt;This is the part that trips people up: a Wait step inside a loop doesn't behave like &lt;code&gt;sleep()&lt;/code&gt; in regular code. The engine re-evaluates the same node on every pass, it doesn't restart it with fresh state.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: track a timestamp per iteration
&lt;/h2&gt;

&lt;p&gt;Instead of computing &lt;code&gt;waitUntil&lt;/code&gt; once from a fixed start time, store a separate target timestamp for every loop iteration, and only calculate a new one the first time that iteration runs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;steps&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;iterations&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;steps&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;__self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;iterations&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;currentIter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;steps&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;loop&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;iteration&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;iterations&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;currentIter&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;iterations&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="nx"&gt;_000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// wait 3 seconds from *now*&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;waitUntil&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;iterations&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;currentIter&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="nx"&gt;iterations&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;loop&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;currentIter&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This holds up under retries and restarts, since the timestamp for a given iteration is stored and reused rather than recalculated, and each loop cycle gets its own delay measured from when that cycle actually began, not from whenever the Wait step was first created.&lt;/p&gt;

&lt;h2&gt;
  
  
  Other ways to handle it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Store the timestamp in process context&lt;/strong&gt; if more than one step needs to reference the same schedule.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Push the timing logic into the While condition itself&lt;/strong&gt;, works well for simple polling loops where ""stop once 3 seconds have passed"" can be expressed directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Move the Wait outside the loop entirely&lt;/strong&gt;, if what's actually needed is one delay before the loop starts, not a delay on every pass.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;Treat a Wait step like any other piece of state: make it idempotent, compute targets relative to now instead of a fixed start time when it's inside a loop, and log the iteration index so timing bugs are easy to spot later. None of this is a limitation. It's just not &lt;code&gt;sleep()&lt;/code&gt;, and it behaves differently the moment you understand what it's actually doing on each pass."&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>javascript</category>
      <category>automation</category>
      <category>debugging</category>
    </item>
    <item>
      <title>Your Fraud Pipeline Is Just Vibes and Cron Jobs</title>
      <dc:creator>The Unmeshed Team</dc:creator>
      <pubDate>Wed, 30 Sep 2026 07:11:12 +0000</pubDate>
      <link>https://dev.to/unmeshed/your-fraud-pipeline-is-just-vibes-and-cron-jobs-47bo</link>
      <guid>https://dev.to/unmeshed/your-fraud-pipeline-is-just-vibes-and-cron-jobs-47bo</guid>
      <description>&lt;p&gt;Ask a fraud analyst what's actually killing their day, and it's never the dramatic stuff. It's the boring stuff, copy this number, paste it there, check if it's above 600 or below it, repeat forever. Basically a very serious game of telephone between two computers that refuse to speak to each other directly.&lt;/p&gt;

&lt;p&gt;And here's the part that doesn't show up in the big scary market stats. Fraud detection is a $52.82 billion industry, on its way to $246 billion by 2032. Most of that money isn't buying smarter fraud detection. It's paying humans to be the USB cable between two systems that should've just been connected in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  The checks were never the hard part
&lt;/h2&gt;

&lt;p&gt;Take something as ordinary as opening a bank account online. Verify identity. Score the risk. Decide what happens next. Three steps. Sounds simple.&lt;/p&gt;

&lt;p&gt;It isn't, because each step usually lives in a different system. Someone has to call the identity check, wait for a result, pass that result to the risk engine, take the score it returns, and decide what happens based on that score. Approve it. Block it. Or send it to a person, because the rules said "needs review" and nothing smarter exists yet.&lt;/p&gt;

&lt;p&gt;None of these steps is hard on its own. Each one is just an API call and a response. What's hard is remembering that one step depends on another, and making sure a failure in step two doesn't quietly break step four.&lt;/p&gt;

&lt;h2&gt;
  
  
  This is why teams end up with a pile of scripts
&lt;/h2&gt;

&lt;p&gt;Without something coordinating this, teams write scripts to hold it together. One script calls the identity check. Another feeds that result into scoring. A third checks the score and decides what to do. It works, for a while.&lt;/p&gt;

&lt;p&gt;Then the identity vendor changes their response format, and nobody updates the script that reads it. Or someone needs a new rule, and has to dig through five scripts to find the right one. Or a script fails quietly overnight, and nobody notices until a customer complains. The scripts weren't badly written. There were just too many of them, each one on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the savings actually show up
&lt;/h2&gt;

&lt;p&gt;Put the sequence, the branching, and the failure handling in one place, and three things happen. Manual reviews drop, because clear score thresholds approve or block automatically, and only unclear cases go to a person. Maintenance drops, because there's one pipeline to update instead of a handful of scripts nobody fully remembers. And cost stops climbing with volume, since processing ten million transactions doesn't need ten times the manual effort that ten thousand did.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual point
&lt;/h2&gt;

&lt;p&gt;Fraud detection isn't hard because identity checks or risk scoring are hard problems. They're not. It's hard because those pieces live in separate systems, and someone has to connect them: catch the failures, pass the data along, decide what happens next. Get that part right, and the individual checks mostly take care of themselves.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>backend</category>
      <category>fintech</category>
      <category>automation</category>
    </item>
    <item>
      <title>The Logistics Bottleneck Nobody Puts On the Roadmap</title>
      <dc:creator>The Unmeshed Team</dc:creator>
      <pubDate>Tue, 29 Sep 2026 06:11:35 +0000</pubDate>
      <link>https://dev.to/unmeshed/the-logistics-bottleneck-nobody-puts-on-the-roadmap-4gpa</link>
      <guid>https://dev.to/unmeshed/the-logistics-bottleneck-nobody-puts-on-the-roadmap-4gpa</guid>
      <description>&lt;p&gt;Ask anyone in freight ops what their day actually looks like, and it's rarely about the shipment. It's about the tabs.&lt;/p&gt;

&lt;p&gt;Quoting one shipment means juggling four of them: pricing here, tracking there, billing somewhere else, and customer comms in whatever tool never got replaced.&lt;/p&gt;

&lt;p&gt;Nobody designed it that way. It just piled up until someone's job became copy, paste, and hope.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the hours actually go
&lt;/h2&gt;

&lt;p&gt;Most logistics stacks don't start out fragmented. They get that way one integration at a time. A new carrier system comes online. A partner needs a custom data feed. A script gets written to move data from A to B, then another to move it from B to C, and eventually there's a small army of brittle scripts holding the whole thing together.&lt;/p&gt;

&lt;p&gt;None of those scripts are wrong, exactly. But nobody's job is to know all of them, and when a shipment update needs to touch pricing, tracking, and billing in the right order, the actual coordination work falls on a person instead of the system. Multiply that across freight quoting, invoice reconciliation, and delivery exception handling, and you've got a team spending real hours on work that's mechanical, repetitive, and invisible on any roadmap.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes when the workflow is the unit of work
&lt;/h2&gt;

&lt;p&gt;The fix isn't a bigger integration layer between every pair of systems. It's treating the process itself, quote a shipment, reconcile an invoice, flag a delayed delivery, as one thing with a defined shape, instead of a chain of scripts that happen to run one after another.&lt;/p&gt;

&lt;p&gt;Once a process like freight quoting is built this way, it stops being a one-off. The same module that handles quoting can plug into a different workflow that needs pricing data, without anyone rewriting the logic. Version control and testing apply to the workflow itself, not just the code around it, so a change to how invoices get reconciled doesn't quietly break delivery exception handling six months later.&lt;/p&gt;

&lt;p&gt;That reuse is where the time actually comes back. Teams stop rebuilding the same coordination logic for every new process and start assembling new ones from pieces that already work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers behind it
&lt;/h2&gt;

&lt;p&gt;A 2024 study in the International Journal for Multidisciplinary Research, looking at enterprises that adopted structured orchestration, found meaningful gains on exactly this kind of fragmentation: fewer integration failures, better use of engineering time, and noticeably faster deployment cycles, all without adding infrastructure complexity. The specific percentages will vary by organization, but the direction is consistent with what shows up anywhere manual coordination gets replaced with a defined process: fewer people doing the work of gluing systems together, more of them doing the work only they can do.&lt;/p&gt;

&lt;p&gt;For a digital freight platform juggling carriers, partners, and customer-facing tracking at once, that's not a nice-to-have. It's the difference between scaling the business and scaling the number of people needed to hold it together.&lt;/p&gt;

&lt;p&gt;If your team is stitching together pricing, tracking, billing, and communication by hand, &lt;a href="https://unmeshed.io" rel="noopener noreferrer"&gt;Unmeshed&lt;/a&gt; is built to be the layer that turns that coordination into something reusable instead of something rebuilt every time."&lt;/p&gt;

</description>
      <category>logistics</category>
      <category>automation</category>
      <category>softwareengineering</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>AI Agent Infrastructure Explained: The 5 Layers You Need in Production</title>
      <dc:creator>The Unmeshed Team</dc:creator>
      <pubDate>Mon, 28 Sep 2026 05:37:43 +0000</pubDate>
      <link>https://dev.to/unmeshed/ai-agent-infrastructure-explained-the-5-layers-you-need-in-production-27fb</link>
      <guid>https://dev.to/unmeshed/ai-agent-infrastructure-explained-the-5-layers-you-need-in-production-27fb</guid>
      <description>&lt;h2&gt;
  
  
  TLDR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Most agentic AI projects don't fail because the model is bad; they fail because the infrastructure underneath isn't there.&lt;/li&gt;
&lt;li&gt;"Agent infrastructure" isn't one thing. It's five layers: runtime/orchestration, identity/governance, data/memory, observability/evals, cost/resource management.&lt;/li&gt;
&lt;li&gt;No vendor, including Unmeshed, covers all five. Anyone claiming otherwise is drawing the map around their own product.&lt;/li&gt;
&lt;li&gt;Unmeshed handles runtime/orchestration and identity/governance. Data/memory and model routing are gaps, stated plainly.&lt;/li&gt;
&lt;li&gt;Fix the layer that's actually breaking for you first. Don't shop for a platform that claims to own the whole stack.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Over 40% of agentic AI projects will be canceled by the end of 2027. That's not a hot take; it's Gartner's own research.&lt;/p&gt;

&lt;p&gt;The reason isn't that the models got worse. Teams built agents that worked fine in a demo, then watched them fall apart the moment real traffic, real failures, and real audit requirements showed up.&lt;/p&gt;

&lt;p&gt;The model was never the problem. The infrastructure underneath it didn't exist yet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Search "AI agent infrastructure," and you'll land on five different answers&lt;/strong&gt; depending on whose guide you read. Four layers in one framework, five in another, seven in a third.&lt;/p&gt;

&lt;p&gt;Nobody selling a platform actually agrees on where their own product's edges are, and that tells you something. There's no established standard yet, just a handful of companies drawing the map to match what they already built.&lt;/p&gt;

&lt;p&gt;So the useful question isn't which platform covers all of agent infrastructure. Nothing does, honestly.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;useful question is which specific layer is actually breaking for you right now&lt;/strong&gt;, and what fixes that layer specifically. That's what this guide walks through.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. What AI Agent Infrastructure Actually Means
&lt;/h2&gt;

&lt;p&gt;AI agent infrastructure is the specialized layer of runtime, identity, data, observability, and cost-control systems that keeps autonomous agents running reliably once they leave the demo and hit production traffic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That's a mouthful, so here's the shorter version.&lt;/strong&gt; An agent framework, LangChain, CrewAI, Mastra, whatever you're using to define how the agent reasons and calls tools, decides what the agent does.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhdtz7h1kdybmwxvpoi29.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhdtz7h1kdybmwxvpoi29.png" alt="Comparison between AI agent frameworks and AI agent infrastructure showing that frameworks control reasoning while infrastructure provides runtime, governance, observability, and operational reliability." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;AI agent infrastructure decides whether it keeps doing that reliably at 2 am when a downstream API times out, whether anyone can tell what happened after the fact, and whether the bill for a bad loop shows up before it's too late to stop it.&lt;/p&gt;

&lt;p&gt;You can swap frameworks without touching infrastructure, and swap infrastructure without touching the framework. They're genuinely separate problems, which is part of why teams underbuild one while over-investing in the other.&lt;/p&gt;

&lt;p&gt;That distinction is worth holding onto for the rest of this guide, because most of the confusion around AI agent infrastructure comes from collapsing it back into the framework conversation.&lt;/p&gt;

&lt;p&gt;It's also why "the complete agent infra platform" is such a common pitch, and such an inconsistent one from vendor to vendor. Every version of that pitch draws the boundary in a different place, depending on what the company selling it already happened to build.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The Five Layers of AI Agent Infrastructure
&lt;/h2&gt;

&lt;p&gt;Pull up three different guides on this, and you'll get three different layer counts:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy1u3vwfejit8zgix0acm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy1u3vwfejit8zgix0acm.png" alt="AI agent infrastructure architecture diagram showing the five layers required for production AI agents: runtime orchestration, identity governance, data memory, observability evals, and cost management." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agentuity's guide&lt;/strong&gt; settles on four: runtime, orchestration, observability, and cost control&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MindStudio's framework&lt;/strong&gt; splits it into five: runtime orchestration, identity and authorization, data access and memory, payments and resource management, and observability and debugging&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Other guides&lt;/strong&gt; add a sixth or seventh cut specifically for memory or security&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that disagreement is really about the technology. It's about which product each company is selling, and where its edges happen to fall.&lt;/p&gt;

&lt;p&gt;For this guide, five layers cover the ground that actually matters when you're scoping an agent infrastructure stack from scratch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Runtime and orchestration&lt;/li&gt;
&lt;li&gt;Identity and governance&lt;/li&gt;
&lt;li&gt;Data and memory&lt;/li&gt;
&lt;li&gt;Observability and evals&lt;/li&gt;
&lt;li&gt;Cost and resource management&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Whichever count you land on, the point of mapping AI agent infrastructure this way is to stop treating it as one undifferentiated blob and start treating it as five separate buying, building, or fixing decisions.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What Breaks Without It&lt;/th&gt;
&lt;th&gt;Where To Look&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Runtime &amp;amp; orchestration&lt;/td&gt;
&lt;td&gt;Agents lose state mid-task and can't resume where they left off&lt;/td&gt;
&lt;td&gt;Durable execution, &lt;a href="https://unmeshed.io/blog/what-is-api-orchestration" rel="noopener noreferrer"&gt;API orchestration&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Identity &amp;amp; governance&lt;/td&gt;
&lt;td&gt;No audit trail, no record of what the agent actually touched&lt;/td&gt;
&lt;td&gt;Governed AI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data &amp;amp; memory&lt;/td&gt;
&lt;td&gt;Agent starts cold every session, or acts confidently on stale context&lt;/td&gt;
&lt;td&gt;Not a native layer for most orchestration platforms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Observability &amp;amp; evals&lt;/td&gt;
&lt;td&gt;A failed run is undebuggable, quality drifts unnoticed&lt;/td&gt;
&lt;td&gt;LLM observability, AI agent evals, &lt;a href="https://unmeshed.io/blog/llm-as-a-judge-explained" rel="noopener noreferrer"&gt;LLM as a judge&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost &amp;amp; resource management&lt;/td&gt;
&lt;td&gt;Runaway spend, no idea which agent is driving the bill&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://unmeshed.io/blog/llm-gateway-explained-production-ai" rel="noopener noreferrer"&gt;LLM gateway&lt;/a&gt;, &lt;a href="https://unmeshed.io/blog/llm-cost-optimization-9-techniques-to-reduce-token-usage-in-production" rel="noopener noreferrer"&gt;LLM cost optimization&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Runtime and Orchestration
&lt;/h3&gt;

&lt;p&gt;This is the layer most teams hit first, because it's the one that breaks the loudest. An agent runtime manages:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdu1fmyh043fe4g3ort5q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdu1fmyh043fe4g3ort5q.png" alt="AI agent runtime and durable execution diagram showing workflow recovery after a failed tool call, highlighting how agent infrastructure resumes execution instead of restarting from the beginning." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which step runs next&lt;/li&gt;
&lt;li&gt;How state carries across a multi-step task&lt;/li&gt;
&lt;li&gt;What happens when a tool call fails partway through&lt;/li&gt;
&lt;li&gt;How the agent resumes after pausing for a human approval or a slow webhook, instead of losing its place&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An agent that can't survive a downstream API timeout without restarting from zero isn't production software yet. It's an expensive demo with good uptime.&lt;/p&gt;

&lt;p&gt;Get this layer wrong, and every other part of your AI agent infrastructure inherits the instability, because nothing downstream can trust that a run actually finished the way it looks like it did.&lt;/p&gt;

&lt;p&gt;We've written about what durable execution actually requires in &lt;a href="https://unmeshed.io/blog/what-is-durable-execution" rel="noopener noreferrer"&gt;what is durable execution&lt;/a&gt;, and how this differs from plain API orchestration once a workflow includes AI steps and not just service calls.&lt;/p&gt;

&lt;h3&gt;
  
  
  Identity and Governance
&lt;/h3&gt;

&lt;p&gt;An agent is an identity now, not just a script. It holds credentials, calls tools on someone's behalf, and can rack up real consequences if nobody's watching what it does with that access.&lt;/p&gt;

&lt;p&gt;This layer covers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What the agent is allowed to touch&lt;/li&gt;
&lt;li&gt;What a human is allowed to authorize through it&lt;/li&gt;
&lt;li&gt;Whether there's an audit trail that would actually hold up if someone asked for one&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Skip it, and you get an agent with more standing access than any single person on the team, and no record of how it used it. We covered why this stopped being optional for enterprise teams in &lt;a href="https://unmeshed.io/blog/why-enterprises-are-moving-past-vibe-coding-to-governed-ai" rel="noopener noreferrer"&gt;governed AI&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Data and Memory
&lt;/h3&gt;

&lt;p&gt;Two different problems live under one label here.&lt;/p&gt;

&lt;p&gt;Data access is how the agent retrieves the right information at the right moment, through a vector store, a structured tool call, or context loaded straight into the prompt.&lt;/p&gt;

&lt;p&gt;Memory is whether the agent remembers anything between sessions instead of starting cold every single time.&lt;/p&gt;

&lt;p&gt;This is the one layer of agent infrastructure most orchestration-first platforms, Unmeshed included, don't own natively. Memory and retrieval usually come from a dedicated vector database or memory service, paired with whatever runs the rest of the workflow, not bundled into it.&lt;/p&gt;

&lt;p&gt;It's also the layer most frequently oversold in AI agent infrastructure marketing. A thin retrieval integration is easy to demo and hard to distinguish from real long-term memory until it's under real load. Worth naming plainly rather than stretching a feature to cover a gap that's genuinely there.&lt;/p&gt;

&lt;h3&gt;
  
  
  Observability and Evals
&lt;/h3&gt;

&lt;p&gt;When a normal API call fails, you check a log line. When an agent fails after thirty tool calls and a dozen model calls, a log line tells you almost nothing.&lt;/p&gt;

&lt;p&gt;This layer needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tracing across the entire run&lt;/li&gt;
&lt;li&gt;Token and cost tracking per session&lt;/li&gt;
&lt;li&gt;A way to tell whether the output was actually good, not just whether it ran without an error&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last part is evals, different from observability even though the two get lumped together constantly. We drew that line explicitly in &lt;a href="https://unmeshed.io/blog/llm-monitoring-vs-observability" rel="noopener noreferrer"&gt;LLM monitoring vs observability&lt;/a&gt; and went deep on the eval side in &lt;a href="https://unmeshed.io/blog/what-are-ai-agent-evals" rel="noopener noreferrer"&gt;what are AI agent evals&lt;/a&gt;, including how teams automate quality scoring with an LLM as a judge instead of a human reading every transcript.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cost and Resource Management
&lt;/h3&gt;

&lt;p&gt;A single LLM call costs fractions of a cent. An agent making a hundred calls across multiple providers, in a loop that doesn't know when to stop, costs real money, and the invoice usually arrives after the damage is done.&lt;/p&gt;

&lt;p&gt;This layer covers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Token budgets per run&lt;/li&gt;
&lt;li&gt;Rate limiting against downstream services&lt;/li&gt;
&lt;li&gt;Cost attribution, so you know which agent, user, or workflow is actually driving the number&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Multi-provider billing is its own headache here, which is exactly the problem an LLM gateway exists to solve. If runaway spend is the actual pain point, this is also where techniques like the ones in LLM cost optimization pay off fastest.&lt;/p&gt;

&lt;p&gt;Left unmanaged, this is the layer of AI agent infrastructure that turns a promising pilot into a budget conversation nobody wanted to have.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Where Unmeshed Fits, and Where It Doesn't
&lt;/h2&gt;

&lt;p&gt;No hedging on this one. Unmeshed covers two of the five layers directly:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff0z7jxetbmw6437g9zez.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff0z7jxetbmw6437g9zez.png" alt="Unmeshed AI agent infrastructure architecture showing coverage of runtime orchestration and governance layers while integrating with external memory systems and LLM routing platforms." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Runtime and orchestration:&lt;/strong&gt; Durable execution that survives a failed step and resumes instead of restarting, with workflow-level state kept automatically&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Identity and governance:&lt;/strong&gt; Human-in-the-loop approvals, a decision engine for branching logic, and audit logging on every step, so "what did the agent actually do" has a real answer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It does not cover two of the five:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Data and memory:&lt;/strong&gt; No built-in vector store or long-term memory layer, and pretending otherwise would be the kind of overclaiming this guide is arguing against&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model routing:&lt;/strong&gt; It doesn't function as an LLM gateway for multi-provider routing; that's a separate, complementary piece&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your gap is specifically in one of those two layers, you're pairing a dedicated tool with whatever runs your orchestration. That's true whether the orchestration layer is Unmeshed or anything else.&lt;/p&gt;

&lt;p&gt;Most AI agent tooling on the market is built to solve one of these layers well and imply, quietly, that it solves the rest. Being specific about which two is more useful than a feature list that stretches to cover the gaps.&lt;/p&gt;

&lt;blockquote&gt;
&lt;h3&gt;
  
  
  Two of These Five Layers, Handled
&lt;/h3&gt;

&lt;p&gt;Durable execution and governed, auditable agent steps are built in. See what that actually looks like before you build it yourself.&lt;br&gt;
&lt;a href="https://unmeshed.io/signup?utm_source=blog&amp;amp;utm_medium=organic&amp;amp;utm_campaign=what_is_ai_agent_infrastructure&amp;amp;utm_content=accent_cta" rel="noopener noreferrer"&gt;See It In Action&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. Build, Buy, or Assemble
&lt;/h2&gt;

&lt;p&gt;No single agent infra platform actually owns all five layers. The guides that claim to are drawing generous lines around their own product.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fruehygodvc636gxibh6z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fruehygodvc636gxibh6z.png" alt="AI agent infrastructure diagnostic framework helping teams identify whether agent failures are caused by runtime orchestration, governance controls, auditability, or AI cost management issues." width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
Most teams end up assembling from two or three specialized tools rather than buying one platform that does everything. That's not a failure of the market; it's just where the market actually is right now.&lt;/p&gt;

&lt;p&gt;A practical way to decide where to start: figure out which layer is causing you pain today, not which one sounds most important in the abstract.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What's Actually Happening&lt;/th&gt;
&lt;th&gt;The Layer That's Missing&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Agents keep losing state on failure&lt;/td&gt;
&lt;td&gt;Runtime and orchestration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nobody can answer what an agent did last Tuesday&lt;/td&gt;
&lt;td&gt;Identity and governance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The agent gives confidently wrong answers from stale context&lt;/td&gt;
&lt;td&gt;Data and memory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A production incident takes hours to diagnose&lt;/td&gt;
&lt;td&gt;Observability and evals&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Last month's bill had a number nobody can explain&lt;/td&gt;
&lt;td&gt;Cost and resource management&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Solve the layer that's actually on fire first. The rest of the agent infrastructure stack can wait until it's the one causing the pain.&lt;/p&gt;

&lt;p&gt;That's the practical reality of AI agent infrastructure today, whatever a sales deck implies about full coverage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;AI agent infrastructure isn't a single product category, whatever the pitch decks say.&lt;/p&gt;

&lt;p&gt;It's five layers that different vendors draw differently. The honest starting point is knowing which one is actually failing you, not shopping for a platform that claims to own all of them.&lt;/p&gt;

&lt;p&gt;Runtime and orchestration, identity and governance, data and memory, observability and evals, cost and resource management.&lt;/p&gt;

&lt;p&gt;Name the layer that's breaking, fix that one, and build out from there instead of trying to solve all five on day one.&lt;/p&gt;

&lt;p&gt;Still scoping your agent infrastructure stack? Talk to us about which layers you're missing, and whether Unmeshed covers the ones that matter most for your use case. &lt;a href="https://unmeshed.io/contact?utm_source=blog&amp;amp;utm_medium=organic&amp;amp;utm_campaign=what_is_ai_agent_infrastructure&amp;amp;utm_content=accent_cta" rel="noopener noreferrer"&gt;Start The Conversation&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>architecture</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Step Functions Worked Fine. Until Your Workflow Outgrew It.</title>
      <dc:creator>The Unmeshed Team</dc:creator>
      <pubDate>Fri, 25 Sep 2026 06:10:00 +0000</pubDate>
      <link>https://dev.to/unmeshed/step-functions-worked-fine-until-your-workflow-outgrew-it-2p8l</link>
      <guid>https://dev.to/unmeshed/step-functions-worked-fine-until-your-workflow-outgrew-it-2p8l</guid>
      <description>&lt;p&gt;Ask any AWS team how they ended up on Step Functions and you'll get the same answer: it was already there. Already in the console, already on the bill. Nobody had to make a case for it.&lt;/p&gt;

&lt;p&gt;The problems don't show up on day one. They show up slowly, a bill that's a little higher than expected a workflow that's a little slower than it should be for what you're building now. By the time someone brings it up in a meeting, half the room's already been thinking it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The billing model punishes exactly the growth you want
&lt;/h2&gt;

&lt;p&gt;Per-state billing is invisible at low volume. You don't notice it. Then workflow volume climbs into the millions and it stops being invisible, it becomes a number finance asks about. The frustrating part isn't that it's expensive. It's that the pricing model actively gets worse the more successful your usage of the tool becomes, which is backwards from how infrastructure spend is supposed to behave. A flat, fixed-price model doesn't have that problem: cost doesn't move just because volume did.&lt;/p&gt;

&lt;h2&gt;
  
  
  You're betting on one region staying up
&lt;/h2&gt;

&lt;p&gt;Step Functions is AWS-managed, all the way down. Fine, until a regional outage takes your workflows with it, or until ""we should probably not have all our eggs in one cloud"" moves from a hypothetical someone raised in a planning meeting to an actual mandate from above. Orchestration that runs across AWS, Azure, GCP, or on-prem means a provider having a bad day is an incident, not an outage for your customers too.&lt;/p&gt;

&lt;h2&gt;
  
  
  The JSON is the job, some days
&lt;/h2&gt;

&lt;p&gt;Step Functions' state machine definitions are genuinely powerful. They're also genuinely tedious to write and debug by hand, and every hour spent wrestling a state machine into shape is an hour not spent on the workflow's actual business logic. A visual builder and YAML definitions don't make the underlying problem simpler, they just get engineers out of the syntax and back into the logic faster.&lt;/p&gt;

&lt;h2&gt;
  
  
  It wasn't built to feel instant, because it wasn't built for that
&lt;/h2&gt;

&lt;p&gt;This is the one that actually surprises teams. Step Functions is an event-driven tool wearing a request-response costume when you need one, and the seams show:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Every state transition adds measurable latency. That's a non-issue for a nightly batch job and a real problem for a chat response or an AI agent call that needs to feel immediate.&lt;/li&gt;
&lt;li&gt;There's no native synchronous flow. Getting request-response behavior means wrapping the state machine in API Gateway or a Lambda, which means more infrastructure to babysit, not less.&lt;/li&gt;
&lt;li&gt;CloudWatch tells you what happened. It's genuinely bad at telling you what's happening right now: no live view of step progress or partial outputs while a workflow is actually stuck.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A platform built for real-time, API-driven orchestration from the start doesn't have these seams to paper over: no wait between steps, synchronous and asynchronous flows on equal footing, and tracing you can watch live instead of reconstruct afterward.&lt;/p&gt;

&lt;h2&gt;
  
  
  None of this means Step Functions is wrong for you
&lt;/h2&gt;

&lt;p&gt;It's a fine choice for AWS-native, event-driven, background-heavy workflows, and plenty of teams should keep using it exactly as-is. The point where it's worth reconsidering is specific: when volume growth starts hurting your bill, when a single-region dependency stops being acceptable, or when a workflow needs to respond in the time it takes someone to read a chat message rather than the time it takes a batch job to run. That's not a tooling preference. That's a different job than the one Step Functions was built to do."&lt;/p&gt;

</description>
      <category>aws</category>
      <category>stepfunctions</category>
      <category>architecture</category>
      <category>serverless</category>
    </item>
    <item>
      <title>Stop Passing Big Files Through Your Workflow. Here's What to Do Instead.</title>
      <dc:creator>The Unmeshed Team</dc:creator>
      <pubDate>Thu, 24 Sep 2026 06:10:32 +0000</pubDate>
      <link>https://dev.to/unmeshed/stop-passing-big-files-through-your-workflow-heres-what-to-do-instead-42hi</link>
      <guid>https://dev.to/unmeshed/stop-passing-big-files-through-your-workflow-heres-what-to-do-instead-42hi</guid>
      <description>&lt;p&gt;You've built the workflow, tested it with a sample file, and everything works. Then someone uploads a real spreadsheet, the kind with 40,000 rows instead of 10, and suddenly your "simple" step is dragging the whole process down.&lt;/p&gt;

&lt;p&gt;Here's what usually causes it: the workflow is carrying the file itself through every step, like it's just another piece of data. Passed as payload, dragged from step to step, worked with inline. It feels natural to build it that way. It's also the wrong move.&lt;/p&gt;

&lt;p&gt;Here's why, and what to do instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  Small payloads are fine. Big ones aren't.
&lt;/h2&gt;

&lt;p&gt;Passing a small JSON object between steps? No problem. Passing an actual file the same way? That's where things start breaking.&lt;/p&gt;

&lt;p&gt;You get bloated requests and responses. Slower step-to-step execution. Debugging turns into a nightmare once an output is too big to actually read. And you're burning memory on something that has nothing to do with the real work.&lt;/p&gt;

&lt;p&gt;None of this shows up while you're testing with a 10-row spreadsheet. It shows up the day a real file hits production.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: stop moving the file. Move a path instead.
&lt;/h2&gt;

&lt;p&gt;Rather than pushing file contents through the workflow, download the file once into shared storage and let each step that needs it read from disk. Steps pass metadata and results to each other, never the file itself.&lt;/p&gt;

&lt;p&gt;That's three steps, in practice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Download: stream it into storage, don't load it all into memory first&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;steps&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;file_url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://testing.s3.amazonaws.com/abcd/pqr/1234/myfile.xlsx&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;file_path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/app/files/myfiles/customFile.xlsx&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;makedirs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dirname&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;exist_ok&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
            &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;iter_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8192&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                        &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;statusMessage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;File downloaded and saved to &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;isSuccessful&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;statusMessage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Download failed: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;isSuccessful&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Streaming in chunks is the whole point here. It's what keeps a big file from spiking memory on the way in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Process: read it where it lives, return a summary, not the whole file&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pandas&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;steps&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;file_to_process&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/app/files/myfiles/customFile.xlsx&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_excel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_to_process&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fillna&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;N/A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;numeric_cols&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select_dtypes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;number&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]).&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;
        &lt;span class="n"&gt;numeric_summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;col&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;col&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;col&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;numeric_cols&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rowCount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;columnCount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;columns&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;numericSummary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;numeric_summary&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;preview&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;to_dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orient&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;records&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;statusMessage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Error processing file: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;isSuccessful&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;3. Delete: clean up once you're done, every single time&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;steps&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;file_path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/app/files/myfiles/customFile.xlsx&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;remove&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;File deleted successfully&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;File does not exist&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;statusMessage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;isSuccessful&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;statusMessage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;isSuccessful&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Skip this step and storage quietly fills up on any process that runs a few hundred times a day.&lt;/p&gt;

&lt;h2&gt;
  
  
  The step people actually get wrong
&lt;/h2&gt;

&lt;p&gt;It's step 2. The instinct is to return the parsed data as-is. Resist it.&lt;/p&gt;

&lt;p&gt;Figure out what the &lt;em&gt;next&lt;/em&gt; step genuinely needs, a row count, a list of columns, a few preview rows, and return only that. If something downstream truly needs the full dataset later, it's still sitting right there on disk.&lt;/p&gt;

&lt;p&gt;Get this one decision right and the rest of the workflow stays light. Everything after that step is working with a few KB of summary instead of megabytes of spreadsheet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Works past Excel too
&lt;/h2&gt;

&lt;p&gt;Nothing about this pattern is Excel-specific. Download → process → clean up works the same for CSVs, PDFs, generated reports, basically any bulky file your workflow has to touch. Only the processing logic in step 2 changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one habit that matters
&lt;/h2&gt;

&lt;p&gt;Big files stop being a headache the moment you stop treating them as data flowing through the workflow and start treating them as a resource the workflow borrows temporarily. Download it, use it, get rid of it. Nothing downstream has to carry weight it never needed.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>backend</category>
      <category>python</category>
      <category>automation</category>
    </item>
    <item>
      <title>Your Order Fulfillment Workflow Is One 24-Hour Wait Away From Chaos</title>
      <dc:creator>The Unmeshed Team</dc:creator>
      <pubDate>Wed, 23 Sep 2026 06:54:35 +0000</pubDate>
      <link>https://dev.to/unmeshed/your-order-fulfillment-workflow-is-one-24-hour-wait-away-from-chaos-3nc3</link>
      <guid>https://dev.to/unmeshed/your-order-fulfillment-workflow-is-one-24-hour-wait-away-from-chaos-3nc3</guid>
      <description>&lt;p&gt;Here's a fun fact about order fulfillment: it looks like one button click. "Buy now." Done, right?&lt;/p&gt;

&lt;p&gt;Wrong. Behind that button is a process that spans five systems, takes hours to finish, and has more opinions than your group chat. And almost every team builds it the same two ways both of which fall apart at the exact same moment: the moment it has to &lt;em&gt;wait&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;We watched this happen with a telecom order flow. Let's talk about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step one: chain some APIs together. What could go wrong?
&lt;/h2&gt;

&lt;p&gt;The classic move. Frontend calls Service A, Service A calls Service B, Service B calls Service C, and everyone's happy right up until you need to hold state for a while. Like, say, keeping someone's cart open for 24 hours because they got distracted by a raccoon on their porch camera and never finished checkout.&lt;/p&gt;

&lt;p&gt;There's no natural place to put that wait. So someone bolts on a cron job. Then a timer. Then a ""temporary"" flag in a database that becomes permanent the way all temporary things do. Congrats, you've built a haunted house of half-finished retry logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step two: ""let's just go event-driven!""
&lt;/h2&gt;

&lt;p&gt;This is the fix everyone reaches for next, and it does scale better. Publish events, let services react, store state wherever each service feels like storing it that day.&lt;/p&gt;

&lt;p&gt;The catch: your workflow logic is now scattered across every service that happens to be listening. So when someone (a customer, your boss, your own past self at 2am) asks ""where is this order stuck?"" buckle up. You're now spelunking through five different log files, lining up timestamps like you're solving a murder mystery nobody asked you to solve.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the workflow is actually asking for
&lt;/h2&gt;

&lt;p&gt;Strip away the telecom flavor and the order just wants four things, which is honestly not that much to ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hold state for hours&lt;/strong&gt; without forgetting the order exists (looking at you, 24-hour cart)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do independent stuff at the same time&lt;/strong&gt; — validating the plan, pulling customer data, running a credit check don't need to wait in line for each other&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pause for a human when it matters&lt;/strong&gt; — route the approval, let someone escalate it, resume without breaking anything&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't duplicate work on retry&lt;/strong&gt; — nobody wants to accidentally reserve two phones for one order because a request got retried&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Swap ""order"" for ""insurance claim"" or ""loan application"" and the list doesn't change. This isn't a telecom problem. It's a ""any process that takes longer than a coffee break"" problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part nobody brags about, but should
&lt;/h2&gt;

&lt;p&gt;Here's the underrated payoff: what happens when it breaks.&lt;/p&gt;

&lt;p&gt;If every step, every input, and every decision gets logged as the workflow runs, ""what happened to this order?"" becomes a five-second lookup by order ID. Which steps ran, what they got, where it died. If that record doesn't exist, you're back to reconstructing the crime scene from logs that were never meant to talk to each other.&lt;/p&gt;

&lt;p&gt;That's the actual pitch for orchestration. It's not that it connects your systems better anything can connect systems if you yell at it long enough. It's that your long-running, easily-distracted-by-a-24-hour-wait workflow finally has one address. Somewhere it can pause, get poked, resume, and be interrogated later without anyone crying.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>microservices</category>
      <category>distributedsystems</category>
    </item>
  </channel>
</rss>
