<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Paul-S</title>
    <description>The latest articles on DEV Community by Paul-S (@paul-s).</description>
    <link>https://dev.to/paul-s</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3112427%2F826409dc-379c-42dd-b263-05b74b80c4b9.png</url>
      <title>DEV Community: Paul-S</title>
      <link>https://dev.to/paul-s</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/paul-s"/>
    <language>en</language>
    <item>
      <title>The Flutter Upgrade Passed CI. Plugins Failed on Device</title>
      <dc:creator>Paul-S</dc:creator>
      <pubDate>Wed, 30 Sep 2026 11:25:51 +0000</pubDate>
      <link>https://dev.to/paul-s/the-flutter-upgrade-passed-ci-plugins-failed-on-device-16lm</link>
      <guid>https://dev.to/paul-s/the-flutter-upgrade-passed-ci-plugins-failed-on-device-16lm</guid>
      <description>&lt;p&gt;The upgrade pipeline looked perfect.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flutter analyze                         PASS
flutter test                            PASS
flutter build apk --release             PASS
flutter build ios --release --no-codesign PASS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The pull request was approved. The new build reached the testing team.&lt;/p&gt;

&lt;p&gt;Then the device reports arrived.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Android: Push notifications never register
iOS: Sign-in returns to a blank screen
Both: Unit tests remain green
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CI had not lied. It had successfully tested what we asked it to test.&lt;/p&gt;

&lt;p&gt;The problem was what we never asked it to test.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Green Build Does Not Prove Plugin Behaviour
&lt;/h2&gt;

&lt;p&gt;Most Flutter applications are not entirely Flutter.&lt;/p&gt;

&lt;p&gt;A plugin usually contains a Dart-facing API and a native implementation written in Kotlin, Java, Swift, or Objective-C. The two sides communicate through platform channels or native bindings.&lt;/p&gt;

&lt;p&gt;That creates a longer execution path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Dart code
   ↓
Platform channel
   ↓
Plugin registration
   ↓
Android or iOS lifecycle
   ↓
Native SDK
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A Dart unit test may validate the first layer while never loading the remaining four.&lt;/p&gt;

&lt;p&gt;Flutter’s documentation explains that plugin unit tests often replace the platform implementation with mocks. That is useful for testing application logic, but a mocked method channel cannot reveal that a native class failed to register or an operating-system callback stopped firing.&lt;/p&gt;

&lt;p&gt;The test passes because the native failure never enters the test process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Framework Upgrades Move Native Boundaries Too
&lt;/h2&gt;

&lt;p&gt;A Flutter upgrade is rarely just a Dart SDK change.&lt;/p&gt;

&lt;p&gt;On Android, it can affect the Android Gradle Plugin, Gradle itself, Java compatibility, Kotlin configuration, manifest handling, and plugin build scripts.&lt;/p&gt;

&lt;p&gt;Recent Flutter migration guidance illustrates this clearly. Built-in Kotlin support changes how Android projects and plugins are configured, while add-to-app projects may require manual migration because Flutter cannot rewrite the native host application automatically.&lt;/p&gt;

&lt;p&gt;On iOS, the move toward the &lt;code&gt;UIScene&lt;/code&gt; lifecycle changes where application events are delivered. Plugins relying on older &lt;code&gt;AppDelegate&lt;/code&gt; callbacks may compile correctly but fail when the application receives a URL, notification, or lifecycle event on a real device.&lt;/p&gt;

&lt;p&gt;Compilation answers one question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can these files become an application?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It does not answer another:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Will every native capability behave correctly when the operating system invokes it?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Debug the Boundary, Not Just the Dart Exception
&lt;/h2&gt;

&lt;p&gt;When a plugin fails after an upgrade, the Dart error is often the end of the trail rather than the beginning.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;MissingPluginException&lt;/code&gt;, frozen method call, or empty callback could originate from several places:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The plugin was not registered with the current engine.&lt;/li&gt;
&lt;li&gt;The Android activity or iOS scene changed before a callback returned.&lt;/li&gt;
&lt;li&gt;A required manifest entry, entitlement, or permission is missing.&lt;/li&gt;
&lt;li&gt;The plugin supports debug builds but fails after release optimization.&lt;/li&gt;
&lt;li&gt;Its native dependency is incompatible with the new build toolchain.&lt;/li&gt;
&lt;li&gt;The plugin’s Dart package was updated without a matching native migration.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Start by reproducing the failure on one known device. Then inspect native logs through Android Studio, &lt;code&gt;adb logcat&lt;/code&gt;, Xcode, or the macOS Console.&lt;/p&gt;

&lt;p&gt;Follow one call across the boundary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Did Dart send the method call?
Did the native handler receive it?
Did the operating system start the requested action?
Did the callback return?
Did Flutter receive the result?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is usually faster than repeatedly changing Dart code around a native failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put Plugin Workflows Into the Upgrade Gate
&lt;/h2&gt;

&lt;p&gt;A better upgrade pipeline includes tests that launch the application on the platforms it supports.&lt;/p&gt;

&lt;p&gt;Flutter’s &lt;code&gt;integration_test&lt;/code&gt; package can run tests on emulators, simulators, physical devices, and device farms. Use it for critical workflows that depend on plugins, such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Authentication redirects&lt;/li&gt;
&lt;li&gt;Camera and file access&lt;/li&gt;
&lt;li&gt;Push-notification registration&lt;/li&gt;
&lt;li&gt;Payments and in-app purchases&lt;/li&gt;
&lt;li&gt;Deep links and app links&lt;/li&gt;
&lt;li&gt;Location and background services&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not test only the happy path.&lt;/p&gt;

&lt;p&gt;Deny a permission and request it again. Put the app in the background during a callback. Launch it from a notification. Open a deep link after the process has been terminated. Run the signed release build, not only the debug build.&lt;/p&gt;

&lt;p&gt;The minimum upgrade evidence should look closer to this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Static analysis                 PASS
Dart unit tests                 PASS
Android release build           PASS
iOS release build               PASS
Android plugin workflows        PASS
iOS plugin workflows            PASS
Background and cold-start tests PASS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CI is still important. It simply needs coverage beyond compilation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Change the Definition of “Upgrade Complete”
&lt;/h2&gt;

&lt;p&gt;An upgrade is not complete when &lt;code&gt;flutter build&lt;/code&gt; exits with code zero.&lt;/p&gt;

&lt;p&gt;It is complete when the application’s important native paths have been exercised on Android and iOS, using the build modes and operating-system behaviours customers will actually encounter.&lt;/p&gt;

&lt;p&gt;That may reveal that a plugin must be upgraded, patched, replaced, or temporarily pinned. It may also expose a missing native migration in the application itself.&lt;/p&gt;

&lt;p&gt;This is the kind of evidence teams should expect from a &lt;a href="https://spaculus.com/services/hybrid-app-development-company/" rel="noopener noreferrer"&gt;hybrid app development company&lt;/a&gt;: not only a green pipeline, but a clear device matrix showing which plugin-backed workflows were tested.&lt;/p&gt;

&lt;p&gt;The Flutter upgrade passed CI because the pipeline proved that the project could build.&lt;/p&gt;

&lt;p&gt;The plugins failed because the application had not yet proved that it could run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which Flutter plugin has caused your most difficult device-only failure?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>flutter</category>
    </item>
    <item>
      <title>I Built a Laravel Agent That Must Ask Permission</title>
      <dc:creator>Paul-S</dc:creator>
      <pubDate>Mon, 28 Sep 2026 11:59:36 +0000</pubDate>
      <link>https://dev.to/paul-s/i-built-a-laravel-agent-that-must-ask-permission-366i</link>
      <guid>https://dev.to/paul-s/i-built-a-laravel-agent-that-must-ask-permission-366i</guid>
      <description>&lt;p&gt;The agent had already done the difficult work.&lt;/p&gt;

&lt;p&gt;It found the customer’s order, checked the payment status, read the refund policy, and calculated the amount.&lt;/p&gt;

&lt;p&gt;Then it proposed this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Action: Issue refund
Order: ORD-1842
Amount: $129.00
Reason: Customer received a damaged product
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The calculation looked reasonable. The customer qualified for a refund.&lt;/p&gt;

&lt;p&gt;I still did not want the model pressing the final button.&lt;/p&gt;

&lt;p&gt;A refund changes real business data and moves money. The agent could recommend the action, but a person needed to approve it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Dangerous Line Was One Method Call
&lt;/h2&gt;

&lt;p&gt;My first version exposed an &lt;code&gt;IssueRefund&lt;/code&gt; tool directly to the agent.&lt;/p&gt;

&lt;p&gt;The instructions included a sentence like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Always ask a manager before issuing refunds above $25.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That sounded safe until I considered what enforced it.&lt;/p&gt;

&lt;p&gt;The same model deciding to issue the refund was also responsible for remembering when it needed permission. A changed prompt, incomplete conversation, or different model response could bypass that instruction.&lt;/p&gt;

&lt;p&gt;The permission boundary belonged in PHP.&lt;/p&gt;

&lt;p&gt;Laravel released &lt;a href="https://laravel.com/blog/introducing-laravel-ai-sdk-v1" rel="noopener noreferrer"&gt;AI SDK v1.0&lt;/a&gt; on September 23, 2026. Among its production features is human approval for agent tools. An approvable tool pauses before its &lt;code&gt;handle()&lt;/code&gt; method executes.&lt;/p&gt;

&lt;p&gt;That was the behavior I needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making the Refund Tool Approvable
&lt;/h2&gt;

&lt;p&gt;Here is a simplified version of the tool:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="cp"&gt;&amp;lt;?php&lt;/span&gt;

&lt;span class="kn"&gt;namespace&lt;/span&gt; &lt;span class="nn"&gt;App\Ai\Tools&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Illuminate\Contracts\JsonSchema\JsonSchema&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Laravel\Ai\Approvals\Approval&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Laravel\Ai\Concerns\InteractsWithApprovals&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Laravel\Ai\Contracts\Approvable&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Laravel\Ai\Contracts\Tool&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Laravel\Ai\Tools\Request&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Stringable&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;IssueRefund&lt;/span&gt; &lt;span class="kd"&gt;implements&lt;/span&gt; &lt;span class="nc"&gt;Approvable&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;Tool&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;InteractsWithApprovals&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="kt"&gt;Stringable&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="n"&gt;string&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s1"&gt;'Issue an approved refund for an existing order.'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;protected&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;needsApproval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;Request&lt;/span&gt; &lt;span class="nv"&gt;$request&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kt"&gt;Approval&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="n"&gt;bool&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nv"&gt;$amount&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nv"&gt;$request&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'amount_cents'&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nv"&gt;$amount&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;2500&lt;/span&gt;
            &lt;span class="o"&gt;?&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
            &lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Approval&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;required&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="s1"&gt;'Refunds above $25 require manager approval.'&lt;/span&gt;
            &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;handle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;Request&lt;/span&gt; &lt;span class="nv"&gt;$request&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kt"&gt;Stringable&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="n"&gt;string&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;app&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;RefundService&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;class&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;issueOnce&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;orderId&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;$request&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'order_id'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
            &lt;span class="n"&gt;amountInCents&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;$request&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'amount_cents'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
            &lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;$request&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'reason'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;schema&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;JsonSchema&lt;/span&gt; &lt;span class="nv"&gt;$schema&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kt"&gt;array&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="s1"&gt;'order_id'&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nv"&gt;$schema&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;required&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
            &lt;span class="s1"&gt;'amount_cents'&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nv"&gt;$schema&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;integer&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nb"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;required&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
            &lt;span class="s1"&gt;'reason'&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nv"&gt;$schema&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;required&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="p"&gt;];&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Refunds of $25 or less can continue automatically. Anything above that amount produces an approval request.&lt;/p&gt;

&lt;p&gt;The threshold is ordinary application logic. The model cannot negotiate with it or reinterpret it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Pause Is a Real Workflow State
&lt;/h2&gt;

&lt;p&gt;When the agent selects &lt;code&gt;IssueRefund&lt;/code&gt;, Laravel returns the pending tool call instead of executing it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="nv"&gt;$response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;RefundAgent&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;forUser&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$user&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'Refund the damaged order ORD-1842.'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$response&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;hasPendingApprovals&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;foreach&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$response&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;pendingApprovals&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nv"&gt;$approval&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="c1"&gt;// Display the tool name, arguments and reason.&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The approval screen should show the exact order, amount, reason, and requested action. A vague “Allow this agent to continue?” message does not give the reviewer enough information.&lt;/p&gt;

&lt;p&gt;After the manager approves or rejects the request, the application resumes the conversation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Laravel\Ai\Approvals\Decision&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="nc"&gt;Laravel\Ai\Approvals\Decisions&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nv"&gt;$response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;RefundAgent&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$conversationId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;$manager&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="nf"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Decisions&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;from&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
        &lt;span class="nv"&gt;$toolCallId&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nc"&gt;Decision&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;approve&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;]));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Laravel also supports rejecting the request or editing its arguments before execution. The &lt;a href="https://laravel.com/framework/docs/ai-sdk#human-tool-approval" rel="noopener noreferrer"&gt;human tool approval documentation&lt;/a&gt; covers the full flow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Approval Is Not Authorization
&lt;/h2&gt;

&lt;p&gt;This was the most important lesson from the build.&lt;/p&gt;

&lt;p&gt;An approval button does not prove that the current user is allowed to approve the refund.&lt;/p&gt;

&lt;p&gt;Before resuming the agent, the application must verify that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The conversation belongs to the correct account&lt;/li&gt;
&lt;li&gt;The signed-in user has refund approval permission&lt;/li&gt;
&lt;li&gt;The order still qualifies for a refund&lt;/li&gt;
&lt;li&gt;The approved amount remains valid&lt;/li&gt;
&lt;li&gt;The action has not already been completed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The refund service also needs an idempotency rule. If a queue retries or a user submits the approval twice, the customer should still receive only one refund.&lt;/p&gt;

&lt;p&gt;The official documentation warns that paused turns are matched through the conversation and pending tool calls. The application is responsible for authorizing access before resuming them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Tests That Matter
&lt;/h2&gt;

&lt;p&gt;A happy-path test proves that an approved refund can succeed. That is only the beginning.&lt;/p&gt;

&lt;p&gt;I would also test these cases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An unauthorized employee attempts to approve the refund&lt;/li&gt;
&lt;li&gt;The order status changes while approval is pending&lt;/li&gt;
&lt;li&gt;A manager edits the amount before approval&lt;/li&gt;
&lt;li&gt;The same approval is submitted twice&lt;/li&gt;
&lt;li&gt;The approval references an unknown tool-call ID&lt;/li&gt;
&lt;li&gt;A rejected action returns a useful explanation&lt;/li&gt;
&lt;li&gt;An old approval expires before execution&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These tests check the business boundary around the model rather than the quality of its wording.&lt;/p&gt;

&lt;p&gt;If you are comparing teams under a search such as &lt;a href="https://spaculus.com/services/best-laravel-development-company/" rel="noopener noreferrer"&gt;best Laravel development company&lt;/a&gt;, ask how they would authorize, resume, audit, and retry this workflow. That answer reveals more than another framework quiz.&lt;/p&gt;

&lt;p&gt;The agent can collect evidence and recommend a refund.&lt;/p&gt;

&lt;p&gt;PHP decides whether the tool must pause. Laravel records the pending action. An authorized person makes the final decision. The service executes it once.&lt;/p&gt;

&lt;p&gt;That separation is what made the agent safe enough to use.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which action in your application should an AI agent never execute without approval?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>laravel</category>
      <category>agents</category>
      <category>ai</category>
    </item>
    <item>
      <title>Banks Know Their Customers. Now They Need to Know Their AI Agents.</title>
      <dc:creator>Paul-S</dc:creator>
      <pubDate>Thu, 17 Sep 2026 09:58:56 +0000</pubDate>
      <link>https://dev.to/paul-s/banks-know-their-customers-now-they-need-to-know-their-ai-agents-ka5</link>
      <guid>https://dev.to/paul-s/banks-know-their-customers-now-they-need-to-know-their-ai-agents-ka5</guid>
      <description>&lt;p&gt;“Who authorized this transaction?”&lt;/p&gt;

&lt;p&gt;“The customer.”&lt;/p&gt;

&lt;p&gt;“Which agent submitted it?”&lt;/p&gt;

&lt;p&gt;“We do not record that separately.”&lt;/p&gt;

&lt;p&gt;“Which tools was the agent allowed to use?”&lt;/p&gt;

&lt;p&gt;“It used the customer’s session.”&lt;/p&gt;

&lt;p&gt;“Can we revoke the agent without blocking the customer?”&lt;/p&gt;

&lt;p&gt;Silence.&lt;/p&gt;

&lt;p&gt;This is the architecture review many banks will eventually face.&lt;/p&gt;

&lt;p&gt;Banks already know how to identify customers, employees, applications, and service accounts. AI agents introduce a less familiar actor: software that can interpret a goal, select tools, collect data, and initiate actions on someone else’s behalf.&lt;/p&gt;

&lt;p&gt;Confirming the customer’s identity is no longer enough. The bank must also know which agent is acting, who owns it, what authority it received, and where that authority ends.&lt;/p&gt;

&lt;h2&gt;
  
  
  Know Your Agent Is a Chain of Trust
&lt;/h2&gt;

&lt;p&gt;At the Global Fintech Fest 2026, State Bank of India Chairman CS Setty proposed a “Know Your Agent” framework covering agent identity, authentication, customer consent, transaction limits, audit trails, and revocation.&lt;/p&gt;

&lt;p&gt;His warning was important. An autonomous error can travel through connected systems at machine speed. The same automation that reduces waiting time can also multiply the effect of a bad decision.&lt;/p&gt;

&lt;p&gt;A banking agent should therefore never appear in a transaction as an invisible extension of the customer.&lt;/p&gt;

&lt;p&gt;The trust chain should remain visible:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Customer → Consent → Agent → Policy Decision → Banking API&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;If one link is missing, the bank cannot confidently explain why an action happened.&lt;/p&gt;

&lt;p&gt;Authentication proves which agent made the request. Authorization determines whether that particular agent may perform that particular action for that customer at that moment.&lt;/p&gt;

&lt;p&gt;Banks need both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give the Agent a Passport, Not a Master Key
&lt;/h2&gt;

&lt;p&gt;An agent identity should carry a small, verifiable description of its authority. Think of it as a temporary passport rather than a reusable API key.&lt;/p&gt;

&lt;p&gt;Here is an illustrative manifest:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;agent&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;card-dispute-prod&lt;/span&gt;
  &lt;span class="na"&gt;owner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;disputes-team&lt;/span&gt;
  &lt;span class="na"&gt;acting_for&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;customer-8421&lt;/span&gt;

  &lt;span class="na"&gt;permitted&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;transactions.read&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;fee_reversal.draft&lt;/span&gt;

  &lt;span class="na"&gt;prohibited&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;money.transfer&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;beneficiary.create&lt;/span&gt;

  &lt;span class="na"&gt;approval_required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;fee_reversal.execute&lt;/span&gt;

  &lt;span class="na"&gt;expires_in&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5m&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This information should not exist only inside the model’s prompt. Prompts can be altered, misunderstood, or exposed to injection attacks.&lt;/p&gt;

&lt;p&gt;The passport must be issued, signed, and checked by infrastructure outside the model. Every banking API should evaluate it before accepting an action.&lt;/p&gt;

&lt;p&gt;This also separates three identities that teams often combine:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The customer requesting help&lt;/li&gt;
&lt;li&gt;The agent interpreting the request&lt;/li&gt;
&lt;li&gt;The service executing the action&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That separation makes investigation and revocation possible.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.nccoe.nist.gov/sites/default/files/2026-02/accelerating-the-adoption-of-software-and-ai-agent-identity-and-authorization-concept-paper.pdf" rel="noopener noreferrer"&gt;NIST National Cybersecurity Center of Excellence&lt;/a&gt; is exploring similar questions around agent identification, authentication, least privilege, delegated access, human approval, and verifiable audit records. Its 2026 concept paper also considers established technologies such as OAuth, OpenID Connect, SPIFFE, and SCIM.&lt;/p&gt;

&lt;p&gt;Banks may not need a completely new identity system. They may need to adapt proven identity controls to a new class of non-human user.&lt;/p&gt;

&lt;h2&gt;
  
  
  Authority Should Shrink as Consequence Grows
&lt;/h2&gt;

&lt;p&gt;Not every agent action carries equal risk.&lt;/p&gt;

&lt;p&gt;A banking assistant might operate in four lanes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Explain:&lt;/strong&gt; Describe a fee or summarize account activity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prepare:&lt;/strong&gt; Collect details and draft a dispute.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recommend:&lt;/strong&gt; Suggest a decision to an employee.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Execute:&lt;/strong&gt; Change data, move money, or approve an outcome.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Moving from explanation to execution should not simply unlock more tools. It should narrow the acceptable conditions.&lt;/p&gt;

&lt;p&gt;A balance explanation may require authenticated account access. Adding a beneficiary should require fresh customer confirmation. Reversing a fee may need an employee’s approval. Moving money should require stricter limits, stronger authentication, and an independent fraud check.&lt;/p&gt;

&lt;p&gt;My view is that banks should automate evidence gathering before automating irreversible decisions. It produces meaningful efficiency while keeping responsibility with a person who is accountable for the outcome.&lt;/p&gt;

&lt;h2&gt;
  
  
  Financial-Crime Work Shows a Safer Starting Point
&lt;/h2&gt;

&lt;p&gt;FIS offers a useful example. Its &lt;a href="https://www.fisglobal.com/about-us/media-room/press-release/2026/fis-brings-agentic-ai-to-banking-with-anthropic-starting-with-financial-crimes" rel="noopener noreferrer"&gt;Financial Crimes AI Agent&lt;/a&gt; is designed to assemble evidence from banking systems, evaluate activity, and surface higher-risk cases.&lt;/p&gt;

&lt;p&gt;The investigator still controls the decision.&lt;/p&gt;

&lt;p&gt;That division of work matters. The agent handles the repetitive search across disconnected systems. A qualified human evaluates the evidence and accepts responsibility for the conclusion.&lt;/p&gt;

&lt;p&gt;This is a more defensible starting point than allowing a general-purpose agent to perform every available banking action simply because the customer is authenticated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Revocation Should Be Boring
&lt;/h2&gt;

&lt;p&gt;If an agent behaves unexpectedly, stopping it should not require disabling the customer’s account or rotating a shared production credential.&lt;/p&gt;

&lt;p&gt;Agent authority should be short-lived and revocable independently.&lt;/p&gt;

&lt;p&gt;A bank should be able to answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which agent is active?&lt;/li&gt;
&lt;li&gt;Who or what delegated its authority?&lt;/li&gt;
&lt;li&gt;Which resources can it access?&lt;/li&gt;
&lt;li&gt;What actions has it attempted?&lt;/li&gt;
&lt;li&gt;When does its permission expire?&lt;/li&gt;
&lt;li&gt;Can the bank terminate it immediately?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are ordinary identity-management questions. What changes is the speed and flexibility of the software being controlled.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Next Banking User May Not Be Human
&lt;/h2&gt;

&lt;p&gt;This shift also changes how conversational banking products should be built.&lt;/p&gt;

&lt;p&gt;For an &lt;a href="https://spaculus.com/ai-chatbot-development-company/" rel="noopener noreferrer"&gt;AI chatbot development company&lt;/a&gt;, improving response quality is no longer the complete assignment. Once a chatbot can call tools and modify business systems, its identity, permissions, approval boundaries, and audit trail become part of the product.&lt;/p&gt;

&lt;p&gt;At Spaculus Software, that means treating an AI agent as a governed system actor, not merely a conversational interface.&lt;/p&gt;

&lt;p&gt;Banks have spent decades learning who their customers are. The next challenge is knowing which software is standing beside those customers, what it has been asked to do, and whether it should be allowed to continue.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If an AI agent contacted your banking API today, could your system identify the agent separately from the customer?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>fintech</category>
      <category>security</category>
      <category>architecture</category>
    </item>
    <item>
      <title>An AI Engineer’s Hardest Problem Is Knowing When the Model Should Stop</title>
      <dc:creator>Paul-S</dc:creator>
      <pubDate>Tue, 15 Sep 2026 12:05:21 +0000</pubDate>
      <link>https://dev.to/paul-s/an-ai-engineers-hardest-problem-is-knowing-when-the-model-should-stop-3fjo</link>
      <guid>https://dev.to/paul-s/an-ai-engineers-hardest-problem-is-knowing-when-the-model-should-stop-3fjo</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;RUN 4821
Goal: Process customer refund
Order found: Yes
Refund policy found: Yes
Policy versions found: 2
Amounts calculated: $89 and $129
Tool retries: 4
Next action selected: Issue $129 refund
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent found conflicting policies, calculated two different amounts, and still selected an action.&lt;/p&gt;

&lt;p&gt;That last line might look like successful task completion.&lt;/p&gt;

&lt;p&gt;It is actually the failure.&lt;/p&gt;

&lt;p&gt;The agent should have stopped when it discovered that the evidence could not support one safe decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Missing State in Most Agent Workflows
&lt;/h2&gt;

&lt;p&gt;Many agent workflows define only two outcomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SUCCESS
FAILURE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But an agent can execute every tool successfully and still make the wrong decision. The API responded. The database was available. No exception occurred.&lt;/p&gt;

&lt;p&gt;The uncertainty existed in the business context, not the infrastructure.&lt;/p&gt;

&lt;p&gt;Production agents need more meaningful terminal states:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ASK_FOR_INPUT
WAIT_FOR_APPROVAL
ESCALATE_TO_HUMAN
DENY_ACTION
STOP_LIMIT_REACHED
COMPLETE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A stop is not necessarily a failure. Sometimes it is the most correct output the system can produce.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Another Prompt Will Not Fix This
&lt;/h2&gt;

&lt;p&gt;A common response is to add instructions such as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Only proceed when you are confident.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That delegates the boundary decision to the same probabilistic model creating the uncertainty.&lt;/p&gt;

&lt;p&gt;The model may report high confidence even when its evidence is incomplete. It may interpret repeated tool calls as a reason to keep investigating. A different model version could interpret the instruction differently.&lt;/p&gt;

&lt;p&gt;Prompts can guide reasoning. They should not be the only mechanism protecting consequential actions.&lt;/p&gt;

&lt;p&gt;The application needs a deterministic gate between the model’s recommendation and the real operation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turn Stop Conditions Into Code
&lt;/h2&gt;

&lt;p&gt;Here is a framework-independent example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;next_state&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;authorized&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DENY_ACTION&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;conflicting_evidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ESCALATE_TO_HUMAN&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;missing_required_data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ASK_FOR_INPUT&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;irreversible&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;approved&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;WAIT_FOR_APPROVAL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_attempts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;STOP_LIMIT_REACHED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMPLETE&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice what this function does not check: the model’s self-reported confidence.&lt;/p&gt;

&lt;p&gt;It checks observable conditions such as authorization, missing data, conflicting evidence, reversibility, approval, and execution limits.&lt;/p&gt;

&lt;p&gt;The model can propose the next action. The surrounding software decides whether that action is allowed.&lt;/p&gt;

&lt;p&gt;OWASP calls the combination of excessive functionality, permissions, and autonomy &lt;a href="https://genai.owasp.org/llmrisk/llm062025-excessive-agency/" rel="noopener noreferrer"&gt;Excessive Agency&lt;/a&gt;. Its recommended controls include least-privilege access, limited tools, independent authorization checks, and human approval for high-impact actions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Break the Exit Paths Before Users Do
&lt;/h2&gt;

&lt;p&gt;Happy-path tests ask whether an agent can complete a task.&lt;/p&gt;

&lt;p&gt;Boundary tests ask whether it refuses to complete the wrong task.&lt;/p&gt;

&lt;p&gt;A small &lt;code&gt;pytest&lt;/code&gt; suite can make those decisions visible:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pytest&lt;/span&gt;


&lt;span class="nd"&gt;@pytest.mark.parametrize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;change, expected&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;authorized&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DENY_ACTION&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;conflicting_evidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ESCALATE_TO_HUMAN&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;missing_required_data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ASK_FOR_INPUT&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;irreversible&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;approved&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;WAIT_FOR_APPROVAL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_attempts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;STOP_LIMIT_REACHED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_agent_exit_paths&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_run&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;change&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;expected&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;run&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;base_run&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;change&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="nf"&gt;next_state&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;expected&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These tests are simple, but they force the team to define expected behavior before an unusual request reaches production.&lt;/p&gt;

&lt;p&gt;The next layer is adversarial evaluation. Give the agent contradictory documents, expired authorization, unavailable tools, repeated timeouts, ambiguous user instructions, and requests that cross account boundaries.&lt;/p&gt;

&lt;p&gt;Do not score only the final answer. Record whether the agent chose to continue, ask, pause, deny, or escalate.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf" rel="noopener noreferrer"&gt;NIST AI Risk Management Framework&lt;/a&gt; recommends documenting knowledge limits, defining human oversight, testing under deployment-like conditions, and ensuring AI systems can fail safely when operating beyond their limits.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Hiring Question I Would Ask
&lt;/h2&gt;

&lt;p&gt;I would not evaluate an AI engineer only by asking which models or agent frameworks they have used.&lt;/p&gt;

&lt;p&gt;I would give them the refund trace from the beginning of this article and ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Where should this workflow stop, and which component should enforce that decision?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A strong answer should discuss permissions, evidence quality, action reversibility, approval requirements, retry budgets, logging, and human handoff.&lt;/p&gt;

&lt;p&gt;For teams planning to &lt;a href="https://spaculus.com/hire-ai-engineers/" rel="noopener noreferrer"&gt;hire AI engineers&lt;/a&gt;, this boundary-first evaluation is more revealing than another prompt-engineering exercise. It is also how Spaculus Software approaches production AI workflows: the model reasons, but the system retains control.&lt;/p&gt;

&lt;p&gt;The hardest part of agent engineering is not helping a model finish more tasks.&lt;/p&gt;

&lt;p&gt;It is making sure the system recognizes the tasks it must not finish alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What stop condition would you add before allowing an AI agent to change real business data?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>architecture</category>
      <category>security</category>
    </item>
    <item>
      <title>Your AI Engineer Is Debugging Workflows, Not Prompts</title>
      <dc:creator>Paul-S</dc:creator>
      <pubDate>Fri, 11 Sep 2026 11:10:44 +0000</pubDate>
      <link>https://dev.to/paul-s/your-ai-engineer-is-debugging-workflows-not-prompts-2e8l</link>
      <guid>https://dev.to/paul-s/your-ai-engineer-is-debugging-workflows-not-prompts-2e8l</guid>
      <description>&lt;p&gt;The support agent looked fine during testing.&lt;/p&gt;

&lt;p&gt;It understood refund requests, found the correct order, explained the policy, and asked for confirmation before taking action.&lt;/p&gt;

&lt;p&gt;Then it reached production and refunded the same order twice.&lt;/p&gt;

&lt;p&gt;The first reaction was predictable: rewrite the prompt.&lt;/p&gt;

&lt;p&gt;Add “never issue the same refund twice.” Put it in capital letters. Repeat it near the tool instructions. Lower the temperature.&lt;/p&gt;

&lt;p&gt;None of that fixed the real problem.&lt;/p&gt;

&lt;p&gt;The payment API had completed the first refund, but its response timed out. The workflow interpreted the missing response as failure and retried the tool call. The model followed its instructions both times.&lt;/p&gt;

&lt;p&gt;This was not a prompt bug. It was a distributed-systems bug wearing an AI costume.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Prompt Is Only One Part of the System
&lt;/h2&gt;

&lt;p&gt;A demo often looks like this:&lt;/p&gt;

&lt;p&gt;user request -&amp;gt; model -&amp;gt; answer&lt;/p&gt;

&lt;p&gt;A production agent looks more like this:&lt;/p&gt;

&lt;p&gt;request -&amp;gt; context -&amp;gt; model -&amp;gt; tool -&amp;gt; external API&lt;br&gt;
        -&amp;gt; state update -&amp;gt; model -&amp;gt; approval -&amp;gt; final response&lt;/p&gt;

&lt;p&gt;Every arrow is another place where the workflow can fail.&lt;/p&gt;

&lt;p&gt;The wrong customer record may enter the context. A tool may receive valid JSON with the wrong business meaning. Authentication may expire halfway through a run. A background worker may retry after the external action has already succeeded.&lt;/p&gt;

&lt;p&gt;Prompt changes cannot repair those failures.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tool Calls Create Real Side Effects
&lt;/h2&gt;

&lt;p&gt;Generating a poor paragraph is inconvenient. Sending an email, cancelling an order, or updating a CRM record changes something outside the model.&lt;/p&gt;

&lt;p&gt;That means consequential tools need the same protections as any other production service: input validation, authorization, timeouts, retry policies, idempotency, and audit logs.&lt;/p&gt;

&lt;p&gt;A simplified refund tool might protect repeated calls with an action identifier:&lt;/p&gt;

&lt;p&gt;async def refund_order(order_id: str, action_id: str):&lt;br&gt;
    previous = await actions.find(action_id)&lt;br&gt;
    if previous:&lt;br&gt;
        return previous.result&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;return await actions.run_once(
    action_id,
    lambda: payments.refund(order_id),
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The important instruction is not hidden in the prompt. It is enforced by the application: one logical action should produce one external effect.&lt;/p&gt;

&lt;p&gt;The production implementation still needs atomic storage and protection against concurrent requests, but the boundary is clear. The model can propose the action. The system decides whether that action is safe to execute.&lt;/p&gt;

&lt;h2&gt;
  
  
  Traces Are More Useful Than Chat Logs
&lt;/h2&gt;

&lt;p&gt;A chat transcript shows what the user and model said. It rarely explains why an agent selected a tool, what arguments it generated, how long the API took, or which state was restored after a retry.&lt;/p&gt;

&lt;p&gt;An end-to-end trace should connect the model turn, tool call, approval, API response, state change, and final answer under one run identifier.&lt;/p&gt;

&lt;p&gt;The OpenAI Agents SDK tracing documentation reflects this approach. Its traces can capture generations, function calls, handoffs, guardrails, and custom events as parts of one workflow.&lt;/p&gt;

&lt;p&gt;This makes a useful debugging question possible:&lt;/p&gt;

&lt;p&gt;Where did the actual behavior first diverge from the expected behavior?&lt;/p&gt;

&lt;p&gt;That is much more precise than asking why the model was “confused.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Some Actions Should Pause
&lt;/h2&gt;

&lt;p&gt;Not every decision should be automated simply because the agent can call the necessary tool.&lt;/p&gt;

&lt;p&gt;Reading an order status may be low risk. Issuing a large refund, deleting a record, publishing content, or changing account permissions may require human approval.&lt;/p&gt;

&lt;p&gt;The human-in-the-loop guidance for the OpenAI Agents SDK demonstrates a pause, approve or reject, and resume pattern for sensitive tool calls. The specific framework is optional. The control is not.&lt;/p&gt;

&lt;p&gt;Experienced &lt;a href="https://spaculus.com/hire-ai-engineers/" rel="noopener noreferrer"&gt;AI engineers&lt;/a&gt; define these boundaries before production incidents define them instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evals Should Test the Workflow
&lt;/h2&gt;

&lt;p&gt;Testing ten prompts against ten expected answers is useful, but it does not test the complete system.&lt;/p&gt;

&lt;p&gt;Workflow evaluations should also ask whether the correct tool was selected, whether its arguments were valid, whether unauthorized data was excluded, whether duplicate actions were prevented, and whether low-confidence situations were escalated.&lt;/p&gt;

&lt;p&gt;Every production failure should become a regression case. Over time, the evaluation set becomes a record of how the system has failed in the real world.&lt;/p&gt;

&lt;p&gt;At Spaculus Software, this is why AI engineering work extends beyond model integration. Reliable systems require workflow design, tool boundaries, evaluation, observability, and safe recovery paths around the model.&lt;/p&gt;

&lt;p&gt;The prompt still matters. It guides intent, tone, reasoning, and tool selection.&lt;/p&gt;

&lt;p&gt;But when an AI application breaks in production, the most valuable engineer in the room is usually not asking, “How should we reword this?”&lt;/p&gt;

&lt;p&gt;They are asking, “What exactly happened between the request and the result?”&lt;/p&gt;

&lt;p&gt;What workflow failure has been hardest for you to diagnose?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>promptengineering</category>
      <category>programming</category>
    </item>
    <item>
      <title>I Connected an LLM to Production. Here’s What Broke</title>
      <dc:creator>Paul-S</dc:creator>
      <pubDate>Tue, 08 Sep 2026 10:54:30 +0000</pubDate>
      <link>https://dev.to/paul-s/i-connected-an-llm-to-production-heres-what-broke-4949</link>
      <guid>https://dev.to/paul-s/i-connected-an-llm-to-production-heres-what-broke-4949</guid>
      <description>&lt;p&gt;The demo looked ready.&lt;/p&gt;

&lt;p&gt;The application could search company documents, answer support questions, classify requests, and create tickets through an API. It handled every test prompt we gave it.&lt;/p&gt;

&lt;p&gt;Then real users arrived.&lt;/p&gt;

&lt;p&gt;They asked incomplete questions. They pasted entire email threads. They used terms that did not exist in our documentation. Some requests matched several policies, while others required information the system could not access.&lt;/p&gt;

&lt;p&gt;The LLM was still producing fluent answers. The workflow around it was falling apart.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flmwotnlb9v25r8rjnu2f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flmwotnlb9v25r8rjnu2f.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Output Was Valid Until It Wasn’t
&lt;/h2&gt;

&lt;p&gt;During testing, the model returned predictable JSON:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  "category": "billing",&lt;br&gt;
  "priority": "high",&lt;br&gt;
  "requires_human": true&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;Production inputs were less predictable. Sometimes the model changed a field name, returned an unsupported category, or added an explanation outside the JSON object.&lt;/p&gt;

&lt;p&gt;A response that looks reasonable to a person can still break an application.&lt;/p&gt;

&lt;p&gt;The first improvement was to treat model output as untrusted input. Every response had to pass schema validation before the application could use it.&lt;/p&gt;

&lt;p&gt;from typing import Literal&lt;br&gt;
from pydantic import BaseModel&lt;/p&gt;

&lt;p&gt;class TicketAction(BaseModel):&lt;br&gt;
    category: Literal["billing", "technical", "account"]&lt;br&gt;
    priority: Literal["low", "medium", "high"]&lt;br&gt;
    requires_human: bool&lt;/p&gt;

&lt;p&gt;def validate_action(model_output: dict):&lt;br&gt;
    return TicketAction.model_validate(model_output)&lt;/p&gt;

&lt;p&gt;Structured output reduced parsing failures, but it did not prove that the selected category was correct. Format validation and business validation became separate steps.&lt;/p&gt;

&lt;h2&gt;
  
  
  RAG Retrieved Something Relevant but Not Correct
&lt;/h2&gt;

&lt;p&gt;The knowledge base contained current policies, archived documents, internal notes, and several pages describing similar processes.&lt;/p&gt;

&lt;p&gt;The retrieval system often returned a document related to the question. That did not mean it returned the document needed to answer it.&lt;/p&gt;

&lt;p&gt;One customer asked about cancelling a subscription. The system retrieved an older cancellation policy because it shared more words with the question than the current policy did.&lt;/p&gt;

&lt;p&gt;The solution was not simply increasing the number of retrieved chunks.&lt;/p&gt;

&lt;p&gt;Documents needed version metadata, ownership, effective dates, and access rules. Archived content had to be removed from normal retrieval. The application also needed to recognise conflicting evidence and stop instead of asking the LLM to choose a convenient answer.&lt;/p&gt;

&lt;p&gt;RAG improved the model’s access to information. It did not remove the need to manage that information.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tool Calling Turned Small Errors Into Real Actions
&lt;/h2&gt;

&lt;p&gt;A wrong answer is harmful. A wrong action can be worse.&lt;/p&gt;

&lt;p&gt;Once the LLM could create tickets, update CRM records, and trigger notifications, every uncertain decision had operational consequences.&lt;/p&gt;

&lt;p&gt;Retries created duplicate records. Incorrect parameters sent requests to the wrong queue. A broad service account gave the agent access to functions it never needed.&lt;/p&gt;

&lt;p&gt;We reduced this risk by giving each tool one narrow purpose. Tool arguments were validated outside the model, write operations used idempotency keys, and high-impact actions required confirmation.&lt;/p&gt;

&lt;p&gt;This follows the principle behind OWASP’s guidance on excessive agency: limit available tools, permissions, functionality, and autonomy.&lt;/p&gt;

&lt;p&gt;The model could recommend an action. The application decided whether that action was allowed.&lt;/p&gt;

&lt;h2&gt;
  
  
  One User Request Became Eight Model Calls
&lt;/h2&gt;

&lt;p&gt;The original demo made one request to an LLM.&lt;/p&gt;

&lt;p&gt;The production version classified the user’s intention, rewrote the search query, retrieved documents, reranked the results, generated an answer, checked the answer, selected a tool, and summarized the result.&lt;/p&gt;

&lt;p&gt;Each step appeared reasonable on its own. Together, they created noticeable latency and unpredictable costs.&lt;/p&gt;

&lt;p&gt;Agents made the problem harder because the number of steps could change for every request. A simple question might finish immediately, while an ambiguous request could enter a loop of repeated searches and tool calls.&lt;/p&gt;

&lt;p&gt;We added limits for execution time, model calls, retries, retrieved context, and total tokens. Smaller models handled basic classification, while stronger models were reserved for decisions that needed deeper reasoning.&lt;/p&gt;

&lt;p&gt;The goal was not to minimize every model call. It was to ensure that each call justified its cost and delay.&lt;/p&gt;

&lt;h2&gt;
  
  
  Our Logs Said Everything Was Successful
&lt;/h2&gt;

&lt;p&gt;The API returned 200. The workflow was still wrong.&lt;/p&gt;

&lt;p&gt;Traditional logs showed that the request completed, but they did not explain which documents were retrieved, why a tool was selected, or where the answer changed.&lt;/p&gt;

&lt;p&gt;Production debugging required a trace of the complete workflow. We recorded the prompt version, model version, retrieved document IDs, tool arguments, tool results, latency, token usage, validation failures, and final outcome.&lt;/p&gt;

&lt;p&gt;Sensitive customer data was removed or masked before storage. Observability should help investigate failures without creating a new privacy problem.&lt;/p&gt;

&lt;p&gt;The NIST Generative AI Profile also emphasizes ongoing monitoring, documented responsibilities, incident handling, and human oversight for generative AI systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Evaluation Set Was Too Polite
&lt;/h2&gt;

&lt;p&gt;Our early tests asked clear questions with known answers.&lt;/p&gt;

&lt;p&gt;Real users did not behave like an evaluation dataset.&lt;/p&gt;

&lt;p&gt;They submitted vague instructions, conflicting requests, spelling mistakes, old account details, pasted web content, and questions requiring permissions they did not have.&lt;/p&gt;

&lt;p&gt;The evaluation set had to include those conditions. We added cases involving stale documents, conflicting sources, API timeouts, repeated actions, missing information, prompt injection, and requests that should be refused or escalated.&lt;/p&gt;

&lt;p&gt;Every production failure became a new regression test.&lt;/p&gt;

&lt;p&gt;That changed evaluation from a task completed before launch into a process that continued throughout the product’s life.&lt;/p&gt;

&lt;h2&gt;
  
  
  The LLM Was Only One Part of the Product
&lt;/h2&gt;

&lt;p&gt;The biggest lesson was simple: connecting an LLM to an application is not the same as building a production AI system.&lt;/p&gt;

&lt;p&gt;Reliable AI requires structured outputs, managed knowledge, secure tools, permission checks, human approvals, observability, cost controls, and realistic evaluations.&lt;/p&gt;

&lt;p&gt;This is also what teams should examine when they plan to &lt;a href="https://spaculus.com/hire-ai-engineers/" rel="noopener noreferrer"&gt;hire AI engineers&lt;/a&gt;. Prompting skills matter, but production AI development also requires backend engineering, API design, security, data management, testing, and operational thinking.&lt;/p&gt;

&lt;p&gt;The LLM did not suddenly become less capable after deployment. Production simply exposed every assumption that the demo never tested.&lt;/p&gt;

&lt;p&gt;What broke first when you moved an LLM application into production?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>The Laravel App Worked Fine Until the Customer Base Started Growing</title>
      <dc:creator>Paul-S</dc:creator>
      <pubDate>Tue, 01 Sep 2026 12:43:55 +0000</pubDate>
      <link>https://dev.to/paul-s/the-laravel-app-worked-fine-until-the-customer-base-started-growing-3iko</link>
      <guid>https://dev.to/paul-s/the-laravel-app-worked-fine-until-the-customer-base-started-growing-3iko</guid>
      <description>&lt;p&gt;At launch, the Laravel application felt fast. Pages loaded quickly, orders moved through the system, and the support team heard few complaints. Then the customer base grew. Reports took longer to open, checkout occasionally stalled, and background tasks began to pile up. The application had not suddenly become badly written. Its original assumptions simply no longer matched the workload.&lt;/p&gt;

&lt;p&gt;This is a common stage in a successful product. More customers do not just create more page views. They create more data, simultaneous sessions, uploads, notifications, searches, reports, API calls, and support activity. A design that works comfortably for hundreds of users may reveal bottlenecks when thousands arrive at the same time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Growth changes the shape of the workload
&lt;/h2&gt;

&lt;p&gt;Teams often respond to slower performance by adding a larger server. That can provide temporary relief, but it does not explain where time is being spent. One customer action may trigger several database queries, an email, a webhook, an audit entry, and a file operation. As traffic rises, each small cost is multiplied.&lt;/p&gt;

&lt;p&gt;The first task is therefore measurement. Request duration, error rates, memory use, database load, slow queries, queue depth, and third-party response times help reveal the actual constraint. Without that visibility, scaling becomes an expensive guessing exercise.&lt;/p&gt;

&lt;h2&gt;
  
  
  The database often feels the pressure first
&lt;/h2&gt;

&lt;p&gt;Database problems can stay hidden while tables are small. Missing indexes, repeated queries, and loading more records than a page needs may barely be noticeable during early development. With years of orders or a rapidly growing product catalogue, the same patterns become costly.&lt;/p&gt;

&lt;p&gt;Laravel teams should examine slow-query logs and common user journeys. N+1 queries can often be reduced through appropriate eager loading. Large lists should use pagination, and searches should avoid scanning unnecessary columns or rows. Indexes should support real filtering and sorting patterns, not be added blindly. Optimizing the query behind a busy endpoint is usually more valuable than increasing hardware without investigation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Synchronous work makes customers wait
&lt;/h2&gt;

&lt;p&gt;A web request should finish the work the customer needs immediately and move suitable follow-up tasks elsewhere. Sending emails, generating reports, processing imports, resizing images, or notifying external systems can often run through queues. This keeps the user-facing response short even when the total amount of work grows.&lt;/p&gt;

&lt;p&gt;Queues are not a place to hide unreliable code. Jobs need sensible timeouts, controlled retries, failure monitoring, and protection against running the same action twice. Workers must also scale with demand. A fast website paired with a six-hour queue backlog is still a poor customer experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caching helps when its rules are clear
&lt;/h2&gt;

&lt;p&gt;Frequently read information can be cached to reduce repeated database and API work. Configuration, catalogue summaries, and expensive calculations are common candidates. However, caching introduces a new question: when does the stored value become outdated?&lt;/p&gt;

&lt;p&gt;A cache strategy should define what is stored, how long it remains valid, and which event refreshes or removes it. Caching everything can trade a performance issue for stale prices, incorrect permissions, or confusing customer data. The safest approach starts with measured hot paths and clear invalidation rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  External services become part of reliability
&lt;/h2&gt;

&lt;p&gt;Payment providers, shipping platforms, CRMs, and AI services may work perfectly in a low-volume test. At scale, rate limits, network timeouts, and intermittent failures become normal operating conditions. Integrations should fail gracefully, use appropriate retries, record enough context for diagnosis, and avoid blocking an entire page when a non-essential service is unavailable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scaling requires operational discipline
&lt;/h2&gt;

&lt;p&gt;Adding application instances works best when the application is stateless. Sessions, cache data, and user uploads should not depend on the local disk of one server. Deployments should be repeatable, and database changes should be planned so old and new application versions can operate safely during a release.&lt;/p&gt;

&lt;p&gt;Businesses investing in &lt;a href="https://spaculus.com/services/best-laravel-development-company/" rel="noopener noreferrer"&gt;Custom Laravel Development&lt;/a&gt; should plan for expected traffic, data growth, background processing, integrations, failure handling, and observability before adding infrastructure. These decisions make later growth less disruptive and help teams scale the parts of the system that genuinely need it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with evidence, not a rewrite
&lt;/h2&gt;

&lt;p&gt;A growing application does not automatically need microservices or a complete rebuild. A well-structured Laravel monolith can support substantial demand when its database access, queues, caching, and deployment model are designed carefully. Splitting a system too early can add network failures, duplicated data, and operational complexity without improving the customer experience.&lt;/p&gt;

&lt;p&gt;Begin with profiling and realistic load tests. Fix the most expensive query, move slow non-essential work to queues, add monitoring, and test again. Repeat until the system meets a defined performance target. Architecture should change only when evidence shows that a clear boundary needs independent scaling, ownership, or release cycles.&lt;/p&gt;

&lt;p&gt;Customer growth is not proof that Laravel failed. It is proof that the product reached a workload its early design never had to handle. The strongest teams treat that moment as a signal: measure what changed, remove the real bottlenecks, and build the operational habits needed for the next stage of growth.&lt;/p&gt;

</description>
      <category>laravel</category>
      <category>coding</category>
    </item>
    <item>
      <title>AI Chatbot vs Traditional Mobile App: What Do Customers Actually Need?</title>
      <dc:creator>Paul-S</dc:creator>
      <pubDate>Tue, 01 Sep 2026 11:51:41 +0000</pubDate>
      <link>https://dev.to/paul-s/ai-chatbot-vs-traditional-mobile-app-what-do-customers-actually-need-296m</link>
      <guid>https://dev.to/paul-s/ai-chatbot-vs-traditional-mobile-app-what-do-customers-actually-need-296m</guid>
      <description>&lt;p&gt;When businesses discuss digital customer experience, the conversation often becomes a contest: should they build an AI chatbot or a traditional mobile app? That framing sounds simple, but it overlooks the most important person in the decision—the customer.&lt;/p&gt;

&lt;p&gt;Customers rarely care which technology sits behind an experience. They want to complete a task quickly, understand what is happening, and feel confident that their information is safe. Sometimes a conversation is the easiest route. In other situations, a familiar screen with clear buttons is faster and more dependable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a traditional mobile app does well
&lt;/h2&gt;

&lt;p&gt;A mobile app is strong when customers perform the same structured tasks regularly. Banking customers may want to check balances, review transactions, transfer money, or manage cards. Fitness users may want dashboards, progress charts, saved routines, and device tracking. These actions benefit from a stable visual layout that users can learn over time.&lt;/p&gt;

&lt;p&gt;Apps are also useful when an experience depends on phone features such as the camera, GPS, biometric login, push notifications, or offline access. A well-designed app gives users direct control. They can see available options, compare information, move backward, and confirm an action before submitting it.&lt;/p&gt;

&lt;p&gt;That predictability matters for high-value or sensitive tasks. Customers may prefer a visible form and a clear confirmation screen when making a payment, changing account settings, or uploading personal documents.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where an AI chatbot can reduce friction
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbfoahsv7terklanhbulx.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbfoahsv7terklanhbulx.jpg" alt=" " width="608" height="328"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Chatbots are more helpful when customers do not know where to begin. Instead of searching through menus, users can explain what they need in their own words. The chatbot can ask follow-up questions, clarify the request, and guide them toward a relevant answer or action.&lt;/p&gt;

&lt;p&gt;This approach works well for product discovery, appointment scheduling, onboarding, troubleshooting, and questions that involve several possible routes. A customer could describe the type of insurance cover they need, ask which subscription suits their team, or explain an unusual delivery problem without first learning the company's navigation structure.&lt;/p&gt;

&lt;p&gt;Conversation can also make digital services easier for people who struggle with complex menus or unfamiliar terminology. Voice input, translation, and step-by-step explanations can improve access when they are designed carefully.&lt;/p&gt;

&lt;h2&gt;
  
  
  A chatbot is not automatically the simpler choice
&lt;/h2&gt;

&lt;p&gt;Natural language feels easy, but it can introduce uncertainty. A customer may not know what the chatbot can do, how to phrase a request, or whether an answer is accurate. Long conversations can also become slower than tapping a few familiar buttons.&lt;/p&gt;

&lt;p&gt;Chatbots are weak when users need to scan, compare, or control several items at once. Choosing airline seats, reviewing financial charts, editing a detailed profile, or comparing product specifications usually works better through a visual interface. Customers should not have to conduct a lengthy conversation for a task that a simple screen can complete in seconds.&lt;/p&gt;

&lt;p&gt;Trust is another concern. If a chatbot can access accounts or perform actions, users need clear confirmation, visible limits, and an easy way to reach a person. A confident but incorrect answer can damage the experience faster than a confusing menu.&lt;/p&gt;

&lt;h2&gt;
  
  
  The practical answer is often a hybrid experience
&lt;/h2&gt;

&lt;p&gt;Businesses do not always need to choose one interface. A mobile app can provide the stable structure, while an AI assistant helps users navigate it. The chatbot might explain a feature, find a transaction, summarize account activity, or prepare an action. The app can then display the details and ask the user to review and confirm.&lt;/p&gt;

&lt;p&gt;This combination uses conversation for discovery and guidance while keeping visual controls for comparison, editing, and approval. It also allows customers to switch methods. Someone may start with a question, move to a form, and return to the conversation if they need help.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with customer tasks, not technology
&lt;/h2&gt;

&lt;p&gt;Businesses considering &lt;a href="https://spaculus.com/ai-chatbot-development-company/" rel="noopener noreferrer"&gt;AI chatbot app development services&lt;/a&gt; should begin by studying the tasks customers are trying to complete. The right design depends on how often the task occurs, how much information it involves, how serious an error would be, and whether customers need to compare options visually.&lt;/p&gt;

&lt;p&gt;Teams should map the complete journey, including what happens when the system does not understand, when data is missing, or when human judgment is required. They should also test the experience with real customers rather than assuming that a conversational interface is naturally easier for everyone.&lt;/p&gt;

&lt;p&gt;Useful measures include completion rate, time required, abandonment, repeated attempts, customer satisfaction, and the quality of human handoffs. These reveal whether the interface is solving a genuine problem or simply adding a fashionable AI layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give customers the shortest reliable path
&lt;/h2&gt;

&lt;p&gt;AI chatbots and traditional mobile apps serve different needs. Apps provide structure, visibility, and control. Chatbots provide flexibility, guidance, and a natural starting point for uncertain requests. Neither is universally better.&lt;/p&gt;

&lt;p&gt;What customers actually need is the shortest reliable path to their goal. The best digital products choose the interfac&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mobile</category>
    </item>
    <item>
      <title>Five Signs Your Python Application Needs an Experienced Engineer</title>
      <dc:creator>Paul-S</dc:creator>
      <pubDate>Wed, 26 Aug 2026 11:56:55 +0000</pubDate>
      <link>https://dev.to/paul-s/five-signs-your-python-application-needs-an-experienced-engineer-6k0</link>
      <guid>https://dev.to/paul-s/five-signs-your-python-application-needs-an-experienced-engineer-6k0</guid>
      <description>&lt;p&gt;A Python application can appear healthy while problems are quietly growing underneath it. &lt;/p&gt;

&lt;p&gt;Features are being released. Customers can log in. The API responds. Nothing seems urgent. &lt;/p&gt;

&lt;p&gt;Then traffic increases, a dependency is updated, or a new developer joins the project. Suddenly, simple changes take days, errors become difficult to reproduce, and nobody wants to touch certain parts of the code. &lt;/p&gt;

&lt;p&gt;These problems do not necessarily mean Python was the wrong choice. They usually mean the application has grown beyond the engineering practices used to build its first version. &lt;/p&gt;

&lt;p&gt;Here are five signs that your Python application may need a more experienced engineer. &lt;/p&gt;

&lt;h2&gt;
  
  
  1. Every Small Change Breaks Something Else
&lt;/h2&gt;

&lt;p&gt;Adding a field to a form should not break the reporting system. Updating a payment method should not affect user registration. &lt;/p&gt;

&lt;p&gt;When unrelated features repeatedly fail after small changes, the code probably contains tight dependencies. One function may be handling validation, database operations, business rules, and external API calls at the same time. &lt;/p&gt;

&lt;p&gt;This often happens in early-stage products. The team moves quickly because proving the idea matters more than creating perfect architecture. &lt;/p&gt;

&lt;p&gt;That approach can work for a prototype. It becomes dangerous once customers depend on the application. &lt;/p&gt;

&lt;p&gt;An experienced Python engineer can separate responsibilities, introduce clearer boundaries, and reduce the chance that one change creates failures elsewhere. The objective is not to rewrite everything. It is to make future changes safer. &lt;/p&gt;

&lt;h2&gt;
  
  
  2. Nobody Trusts the Test Suite
&lt;/h2&gt;

&lt;p&gt;A test suite should give developers confidence before a release. Instead, some teams have tests that fail randomly, take too long, or cover only the easiest parts of the application. &lt;/p&gt;

&lt;p&gt;The team eventually stops paying attention to failures. Developers rerun the same test until it passes or disable it to complete a deployment. &lt;/p&gt;

&lt;p&gt;At that point, testing becomes decoration rather than protection. &lt;/p&gt;

&lt;p&gt;A senior engineer will first identify which workflows create the greatest business risk. Authentication, payments, permissions, data processing, and third-party integrations usually deserve attention before minor interface details. &lt;/p&gt;

&lt;p&gt;Good testing is not about reaching an impressive coverage percentage. It is about detecting failures that could affect customers or business operations. &lt;/p&gt;

&lt;h2&gt;
  
  
  3. Production Errors Are Difficult to Investigate
&lt;/h2&gt;

&lt;p&gt;Consider this error handling: &lt;/p&gt;

&lt;p&gt;try: &lt;br&gt;
    process_payment(order) &lt;br&gt;
except Exception: &lt;br&gt;
    pass &lt;/p&gt;

&lt;p&gt;The application does not crash, but the payment may fail without leaving useful evidence. The customer sees an incomplete order, while the support team has no information about what happened. &lt;/p&gt;

&lt;p&gt;Changing pass to a log statement is not enough if the system still lacks request IDs, structured logs, alerts, performance metrics, and relevant business context. &lt;/p&gt;

&lt;p&gt;Experienced engineers think about how software will be investigated after deployment. They design applications to answer practical questions: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which customer was affected? &lt;/li&gt;
&lt;li&gt;Which request failed? &lt;/li&gt;
&lt;li&gt;Did an external service time out? &lt;/li&gt;
&lt;li&gt;Can the operation be retried safely? &lt;/li&gt;
&lt;li&gt;Did the same error affect other users? &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If finding the cause of an error requires hours of guesswork, the application needs better observability, not just more debugging. &lt;/p&gt;

&lt;h2&gt;
  
  
  4. Performance Problems Are Solved by Adding Servers
&lt;/h2&gt;

&lt;p&gt;Adding infrastructure can temporarily hide inefficient code, but it does not always solve the underlying problem. &lt;/p&gt;

&lt;p&gt;A slow Python application may be making repeated database queries, loading too much data into memory, calling external services one after another, or performing heavy work inside a web request. &lt;/p&gt;

&lt;p&gt;An experienced engineer measures the system before changing it. The real bottleneck might be a database index, an inefficient query, a blocking network call, or a task that belongs in a background queue. &lt;/p&gt;

&lt;p&gt;This distinction matters because each problem requires a different solution. Adding servers to compensate for a poor database query can increase cloud costs without providing reliable performance. &lt;/p&gt;

&lt;p&gt;Optimization should start with evidence. &lt;/p&gt;

&lt;h2&gt;
  
  
  5. One Developer Holds the Entire System Together
&lt;/h2&gt;

&lt;p&gt;Sometimes only one person knows how deployments work, why a particular workaround exists, or which background job must be restarted manually. &lt;/p&gt;

&lt;p&gt;That person becomes the project’s unofficial documentation. &lt;/p&gt;

&lt;p&gt;This situation creates a serious business risk. If the developer is unavailable, releases slow down and production incidents become harder to resolve. New team members may avoid important parts of the application because they do not understand the consequences of changing them. &lt;/p&gt;

&lt;p&gt;An experienced engineer can reduce this dependency through clearer documentation, automated deployments, code reviews, architecture notes, and repeatable operational processes. &lt;/p&gt;

&lt;p&gt;The goal is not to make every developer know everything. It is to ensure that critical knowledge belongs to the team rather than one individual. &lt;/p&gt;

&lt;h2&gt;
  
  
  Experience Is More Than Writing Advanced Python
&lt;/h2&gt;

&lt;p&gt;Experienced Python engineers do not simply write more complicated code. In many cases, they make the application simpler. &lt;/p&gt;

&lt;p&gt;They know when a small refactor is enough and when an architectural change is necessary. They consider security, testing, deployment, monitoring, and maintainability alongside feature delivery. &lt;/p&gt;

&lt;p&gt;If you plan to &lt;a href="https://spaculus.com/services/hire-python-developers/" rel="noopener noreferrer"&gt;hire Python developers&lt;/a&gt; for an existing application, evaluate more than framework knowledge. Ask candidates how they have diagnosed production failures, improved legacy code, reduced deployment risk, and handled systems that grew beyond their original design. &lt;/p&gt;

&lt;p&gt;At Spaculus Software, this is often the first step when joining an existing Python project: understand the current system before recommending changes. A careful technical review can reveal whether the application needs focused improvements, gradual modernization, or a larger architectural update. &lt;/p&gt;

&lt;h2&gt;
  
  
  Final Thought
&lt;/h2&gt;

&lt;p&gt;A working application is not always a healthy application. &lt;/p&gt;

&lt;p&gt;Frequent regressions, unreliable tests, unclear production errors, rising infrastructure costs, and knowledge concentrated in one person are signs that technical risk is accumulating. &lt;/p&gt;

&lt;p&gt;The right engineer will not begin by rebuilding everything. They will identify the most important risks, protect what already works, and help the application become easier to change as the business grows. &lt;/p&gt;

&lt;p&gt;Which of these warning signs have you encountered in a Python project?&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>python</category>
      <category>softwareengineering</category>
      <category>backenddevelopment</category>
    </item>
    <item>
      <title>The Chatbot Worked in Testing, Then Real Users Arrived</title>
      <dc:creator>Paul-S</dc:creator>
      <pubDate>Tue, 25 Aug 2026 11:53:23 +0000</pubDate>
      <link>https://dev.to/paul-s/the-chatbot-worked-in-testing-then-real-users-arrived-2aa9</link>
      <guid>https://dev.to/paul-s/the-chatbot-worked-in-testing-then-real-users-arrived-2aa9</guid>
      <description>&lt;p&gt;The demo looked ready.&lt;/p&gt;

&lt;p&gt;The chatbot answered every question from the testing document. It explained product features, found account information, and politely handed difficult conversations to a human.&lt;/p&gt;

&lt;p&gt;The team tested it again before launch.&lt;/p&gt;

&lt;p&gt;“Where can I download my invoice?”&lt;/p&gt;

&lt;p&gt;“Can I change my subscription?”&lt;/p&gt;

&lt;p&gt;“What is your refund policy?”&lt;/p&gt;

&lt;p&gt;Every answer was correct.&lt;/p&gt;

&lt;p&gt;Then real customers arrived.&lt;/p&gt;

&lt;p&gt;One user wrote, “charged twice pls fix.” Another sent three messages instead of one complete question. Someone pasted an entire email thread into the chat. A customer referred to “the plan I had before,” although the chatbot had no access to that history.&lt;/p&gt;

&lt;p&gt;By the end of the first day, the team had discovered something its test script never showed:&lt;/p&gt;

&lt;p&gt;The chatbot understood the test cases. It did not yet understand the messiness of real conversations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing Questions Were Too Clean
&lt;/h2&gt;

&lt;p&gt;Development teams often test chatbots with complete, well-written questions because those cases are easy to review.&lt;/p&gt;

&lt;p&gt;Real users do not behave that way.&lt;/p&gt;

&lt;p&gt;They misspell words, change topics halfway through a message, use internal product names, and assume the chatbot remembers information from earlier sessions. They may provide too little context or far more context than the system can process effectively.&lt;/p&gt;

&lt;p&gt;A chatbot tested only with ideal questions is like a payment form tested only with valid cards. It proves that the happy path works, not that the product is ready.&lt;/p&gt;

&lt;p&gt;A stronger evaluation set should include incomplete messages, spelling mistakes, conflicting requests, long conversations, unsupported languages, angry customers, and questions with no verified answer.&lt;/p&gt;

&lt;p&gt;OpenAI’s official evaluation guidance describes evals as an essential part of checking whether model outputs meet defined content and style expectations. The important word is defined. “The answer looks good” is not a measurable production requirement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrieval Worked Until the Wording Changed
&lt;/h2&gt;

&lt;p&gt;The chatbot used retrieval-augmented generation to answer from company documents. During testing, the user’s wording closely matched the documentation.&lt;/p&gt;

&lt;p&gt;The knowledge base said “subscription cancellation.” Testers asked, “How do I cancel my subscription?”&lt;/p&gt;

&lt;p&gt;Customers asked:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“How do I stop getting billed next month?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The intent was the same, but retrieval did not always return the correct document.&lt;/p&gt;

&lt;p&gt;This is why teams should inspect more than the final answer. They need to know which documents were retrieved, how relevant those documents were, and whether the model had enough evidence to respond.&lt;/p&gt;

&lt;p&gt;A useful production trace might record:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  "conversation_id": "conv_1842",&lt;br&gt;
  "intent": "cancel_subscription",&lt;br&gt;
  "retrieved_documents": 3,&lt;br&gt;
  "top_relevance_score": 0.62,&lt;br&gt;
  "response_time_ms": 2480,&lt;br&gt;
  "used_fallback": false,&lt;br&gt;
  "human_handoff": true&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;The exact fields will vary, but the principle remains: if the chatbot produces a bad answer and the team cannot reconstruct what happened, debugging becomes guesswork.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bot Answered When It Should Have Stopped
&lt;/h2&gt;

&lt;p&gt;One customer asked about a policy that was not present in the approved knowledge base.&lt;/p&gt;

&lt;p&gt;The chatbot still responded.&lt;/p&gt;

&lt;p&gt;The answer sounded reasonable, confident, and completely invented.&lt;/p&gt;

&lt;p&gt;This is one of the most dangerous production failures because fluent language can hide missing evidence. Anthropic’s official guidance on reducing hallucinations recommends allowing the model to express uncertainty and grounding responses in direct source material.&lt;/p&gt;

&lt;p&gt;A production chatbot should have a clear refusal or escalation rule:&lt;/p&gt;

&lt;p&gt;if retrieval_score &amp;lt; MIN_CONFIDENCE:&lt;br&gt;
    return {&lt;br&gt;
        "answer": "I don’t have enough verified information to answer that.",&lt;br&gt;
        "action": "handoff_to_human"&lt;br&gt;
    }&lt;/p&gt;

&lt;p&gt;The threshold should be tested using real conversations. Setting it too low increases unsupported answers. Setting it too high sends too many customers to human support.&lt;/p&gt;

&lt;h2&gt;
  
  
  Some Users Tested the Boundaries on Purpose
&lt;/h2&gt;

&lt;p&gt;Not every unexpected input is accidental.&lt;/p&gt;

&lt;p&gt;Users may ask the chatbot to ignore previous instructions, reveal its system prompt, expose private information, or perform actions outside their permissions. If the chatbot reads uploaded files, webpages, emails, or support tickets, malicious instructions may also enter indirectly through that content.&lt;/p&gt;

&lt;p&gt;The OWASP GenAI Security Project lists prompt injection as a major risk for LLM applications. It also makes an important point: retrieval and fine-tuning do not completely remove the problem.&lt;/p&gt;

&lt;p&gt;Teams therefore need controls outside the prompt. Tools should enforce user permissions independently. Sensitive actions should require confirmation. Retrieved content should be treated as untrusted input, and the chatbot should never receive broader system access than the task requires.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production Quality Is More Than Answer Accuracy
&lt;/h2&gt;

&lt;p&gt;A chatbot can answer correctly and still create a poor experience.&lt;/p&gt;

&lt;p&gt;A response that arrives after twelve seconds may cause the customer to leave. A correct answer written in five dense paragraphs may be useless on mobile. A bot that forgets the previous message forces the customer to start again. A handoff that loses the conversation history makes human support repeat the same questions.&lt;/p&gt;

&lt;p&gt;Production monitoring should therefore cover:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Answer correctness and source support&lt;/li&gt;
&lt;li&gt;Retrieval quality&lt;/li&gt;
&lt;li&gt;Response time and failures&lt;/li&gt;
&lt;li&gt;Cost per conversation&lt;/li&gt;
&lt;li&gt;Fallback and handoff rates&lt;/li&gt;
&lt;li&gt;Repeated questions after an answer&lt;/li&gt;
&lt;li&gt;Customer feedback and unresolved conversations&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These metrics should be reviewed by intent. A chatbot may perform well for opening hours and order tracking while failing badly on billing or account access. One overall success rate can hide those differences.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real Conversations Should Become New Tests
&lt;/h2&gt;

&lt;p&gt;The most valuable evaluation set is not created once before launch. It grows from production.&lt;/p&gt;

&lt;p&gt;Failed searches, poor answers, unusual wording, escalated conversations, and negative feedback should become new test cases. Before changing the prompt, model, retrieval settings, or knowledge base, teams can run those cases again and check whether the update fixes one problem without creating another.&lt;/p&gt;

&lt;p&gt;That feedback loop is what separates a chatbot demo from a maintained product.&lt;/p&gt;

&lt;p&gt;It is also the work businesses should examine when choosing an &lt;a href="https://spaculus.com/ai-chatbot-development-company/" rel="noopener noreferrer"&gt;AI chatbot development company&lt;/a&gt;. Spaculus Software supports chatbot architecture, RAG pipelines, integrations, evaluation, security controls, deployment, and ongoing monitoring—not only the chat interface customers see.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Launch Begins After Launch
&lt;/h2&gt;

&lt;p&gt;The chatbot did not suddenly become less intelligent when customers arrived.&lt;/p&gt;

&lt;p&gt;The environment changed.&lt;/p&gt;

&lt;p&gt;Testing gave it clean questions, known answers, and predictable conversations. Production introduced ambiguity, missing context, unusual language, security risks, latency, integration failures, and genuine consequences for being wrong.&lt;/p&gt;

&lt;p&gt;The team’s mistake was not launching too early.&lt;/p&gt;

&lt;p&gt;It was treating launch as the end of testing.&lt;/p&gt;

&lt;p&gt;For an AI chatbot, real users do more than use the product. They reveal the test cases the development team never knew it needed.&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>AI Developers Can Build the Demo in a Week. Production Is Where the Real Work Starts</title>
      <dc:creator>Paul-S</dc:creator>
      <pubDate>Fri, 21 Aug 2026 10:39:55 +0000</pubDate>
      <link>https://dev.to/paul-s/ai-developers-can-build-the-demo-in-a-week-production-is-where-the-real-work-starts-14mi</link>
      <guid>https://dev.to/paul-s/ai-developers-can-build-the-demo-in-a-week-production-is-where-the-real-work-starts-14mi</guid>
      <description>&lt;p&gt;The demo looks perfect. &lt;/p&gt;

&lt;p&gt;A user asks a question. The AI finds the right information, writes a clear answer, and returns it within seconds. &lt;/p&gt;

&lt;p&gt;Everyone in the meeting is impressed. &lt;/p&gt;

&lt;p&gt;Then the application goes live. &lt;/p&gt;

&lt;p&gt;Real users ask unclear questions. Some upload broken files. Others try prompts nobody expected. Responses become slower, API costs rise, and the AI occasionally gives a confident answer that is completely wrong. &lt;/p&gt;

&lt;p&gt;The demo proved that the idea could work. &lt;/p&gt;

&lt;p&gt;Production reveals whether it can keep working. &lt;/p&gt;

&lt;h2&gt;
  
  
  Why Is an AI Demo Easy to Build?
&lt;/h2&gt;

&lt;p&gt;Modern AI APIs make it possible to create a working prototype quickly. A developer can connect a language model, add a simple interface, provide a few instructions, and have something impressive within days. &lt;/p&gt;

&lt;p&gt;That is valuable. A quick demo helps a company test an idea before spending heavily on it. &lt;/p&gt;

&lt;p&gt;But a demo normally works with: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Selected questions &lt;/li&gt;
&lt;li&gt;Clean data &lt;/li&gt;
&lt;li&gt;Limited users &lt;/li&gt;
&lt;li&gt;Controlled conditions &lt;/li&gt;
&lt;li&gt;Little concern about cost or scale &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Production removes all those protections. &lt;/p&gt;

&lt;h2&gt;
  
  
  What Changes When Real Users Arrive?
&lt;/h2&gt;

&lt;p&gt;Real users do not follow the demo script. &lt;/p&gt;

&lt;p&gt;They misspell words, leave out important details, switch topics, upload unexpected formats, and sometimes ask the AI to do things it should never do. &lt;/p&gt;

&lt;p&gt;This is where AI development becomes less about prompts and more about engineering. &lt;/p&gt;

&lt;p&gt;A production AI system needs to know: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;When it has enough information to answer &lt;/li&gt;
&lt;li&gt;When it should search company data &lt;/li&gt;
&lt;li&gt;When it should ask another question &lt;/li&gt;
&lt;li&gt;When it should refuse a request &lt;/li&gt;
&lt;li&gt;When a human should take control &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A better prompt may improve the response, but it cannot solve every production problem. &lt;/p&gt;

&lt;h2&gt;
  
  
  Production AI Must Be Measured, Not Trusted
&lt;/h2&gt;

&lt;p&gt;Traditional software usually gives the same result when it receives the same input. AI systems can behave differently, even when a request looks similar. &lt;/p&gt;

&lt;p&gt;That means testing a few successful examples is not enough. &lt;/p&gt;

&lt;p&gt;AI developers need evaluation sets containing normal questions, difficult cases, incomplete requests, unsafe prompts, and examples collected from real usage. They must measure whether answers are correct, grounded in approved data, useful, fast, and affordable. &lt;/p&gt;

&lt;p&gt;Logging is equally important. If an answer goes wrong, the team should be able to see what the user asked, what information was retrieved, which model responded, and where the process failed. &lt;/p&gt;

&lt;p&gt;Without that visibility, improving the system becomes guesswork. &lt;/p&gt;

&lt;h2&gt;
  
  
  Reliability Is More Than Preventing Hallucinations
&lt;/h2&gt;

&lt;p&gt;Wrong answers receive the most attention, but production AI can fail in quieter ways. &lt;/p&gt;

&lt;p&gt;A response may be correct but arrive too late. An AI agent may repeat a tool call and increase costs. A retrieval system may find an outdated document. A model update may change behaviour that worked yesterday. &lt;/p&gt;

&lt;p&gt;Production engineering therefore includes: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Data quality and retrieval &lt;/li&gt;
&lt;li&gt;Security and access control &lt;/li&gt;
&lt;li&gt;Response evaluation &lt;/li&gt;
&lt;li&gt;Monitoring and alerts &lt;/li&gt;
&lt;li&gt;Cost and latency control &lt;/li&gt;
&lt;li&gt;Fallbacks and human review &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is often the point when companies decide to hire AI engineers instead of treating AI as one more API integration. &lt;/p&gt;

&lt;h2&gt;
  
  
  When Should You Hire AI Developers?
&lt;/h2&gt;

&lt;p&gt;You probably do not need a large AI team to test an early idea. One focused prototype can answer an important question: Does this solve a real user problem? &lt;/p&gt;

&lt;p&gt;Once the answer is yes, the requirements change. &lt;/p&gt;

&lt;p&gt;If you plan to &lt;a href="https://spaculus.com/hire-ai-engineers/" rel="noopener noreferrer"&gt;Hire AI engineers&lt;/a&gt;, look beyond model knowledge and prompt writing. Ask how they test output quality, protect private data, control costs, handle model failures, and monitor the complete user request. &lt;/p&gt;

&lt;p&gt;While exploring production AI at Spaculus Software, one lesson keeps coming up: getting the first answer is easy, but making the complete system reliable takes real engineering. The goal is not simply to make AI answer once. It is to make the complete system useful, measurable, and dependable when real people start using it. &lt;/p&gt;

&lt;p&gt;A demo earns attention. Production earns trust. &lt;/p&gt;

&lt;p&gt;*&lt;em&gt;For developers who have shipped an AI feature: what failed first when real users arrived? *&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>developers</category>
      <category>programmers</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Can I Hire MEAN Stack Developers to Upgrade an Existing Angular or Node.js Application?</title>
      <dc:creator>Paul-S</dc:creator>
      <pubDate>Thu, 20 Aug 2026 13:17:30 +0000</pubDate>
      <link>https://dev.to/paul-s/can-i-hire-mean-stack-developers-to-upgrade-an-existing-angular-or-nodejs-application-4hdb</link>
      <guid>https://dev.to/paul-s/can-i-hire-mean-stack-developers-to-upgrade-an-existing-angular-or-nodejs-application-4hdb</guid>
      <description>&lt;p&gt;Yes. You can hire MEAN Stack developers to upgrade an existing Angular frontend, Node.js backend, or complete JavaScript application.&lt;/p&gt;

&lt;p&gt;A MEAN developer works across MongoDB, Express.js, Angular, and Node.js. This full-stack knowledge is useful when an upgrade affects both the user interface and backend APIs.&lt;/p&gt;

&lt;p&gt;However, the MEAN Stack label alone does not guarantee a successful upgrade. The developer should have practical experience with your current Angular or Node.js version, dependencies, testing environment, database, deployment process, and application architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Should You Upgrade an Existing Application?
&lt;/h2&gt;

&lt;p&gt;Framework upgrades are not only about accessing new features. They help applications continue receiving security fixes, maintain compatibility with third-party packages, and avoid growing technical debt.&lt;/p&gt;

&lt;p&gt;As of August 2026, Angular 22 is under active support. Angular 21 and Angular 20 are receiving long-term support, although Angular 20’s LTS period ends on November 28, 2026. Angular 19 and earlier versions are no longer supported.&lt;/p&gt;

&lt;p&gt;Angular provides approximately 12 months of active support followed by 12 months of long-term support. Once a version becomes unsupported, it no longer receives regular framework fixes or security patches. Angular release policy&lt;/p&gt;

&lt;p&gt;Node.js has a different release model. Node.js 26 is the current release, while Node.js 24 and Node.js 22 are supported LTS versions. Node.js 20 and earlier versions have reached end of life.&lt;/p&gt;

&lt;p&gt;The Node.js project recommends using only Active LTS or Maintenance LTS releases in production. Node.js release schedule&lt;/p&gt;

&lt;p&gt;If your application uses an unsupported version, postponing the upgrade can increase security, compatibility, and maintenance risks.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Is a MEAN Stack Developer the Right Choice?
&lt;/h2&gt;

&lt;p&gt;A MEAN developer is a strong fit when an Angular upgrade also affects Node.js APIs, Express middleware, authentication, MongoDB queries, or shared TypeScript models.&lt;/p&gt;

&lt;p&gt;For example, upgrading Angular may require a newer TypeScript version. That change can affect shared packages used by the Node.js backend. A full-stack developer can review these dependencies together instead of treating the frontend and backend as separate systems.&lt;/p&gt;

&lt;p&gt;MEAN developers can also help when the project needs more than a version update. They can replace deprecated packages, improve API performance, review MongoDB queries, modernize Angular components, and strengthen automated testing.&lt;/p&gt;

&lt;p&gt;If the application contains only a small Angular interface without Node.js or MongoDB, a dedicated Angular developer may be more efficient. A complex Node.js microservices platform may similarly require a backend specialist with distributed-systems experience.&lt;/p&gt;

&lt;p&gt;The right choice depends on the actual architecture, not the name of the stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Should Be Checked Before the Upgrade?
&lt;/h2&gt;

&lt;p&gt;An upgrade should begin with an application audit rather than immediately changing package versions.&lt;/p&gt;

&lt;p&gt;For an Angular application, developers can start with:&lt;/p&gt;

&lt;p&gt;ng version&lt;br&gt;
ng update&lt;br&gt;
npm outdated&lt;br&gt;
npm audit&lt;/p&gt;

&lt;p&gt;For a Node.js application:&lt;/p&gt;

&lt;p&gt;node --version&lt;br&gt;
npm outdated&lt;br&gt;
npm audit&lt;br&gt;
npm test&lt;/p&gt;

&lt;p&gt;These commands identify the current versions, outdated dependencies, known package vulnerabilities, and the condition of the existing test suite.&lt;/p&gt;

&lt;p&gt;The audit should also cover custom build configurations, deprecated APIs, UI libraries, authentication packages, database drivers, environment variables, and CI/CD workflows.&lt;/p&gt;

&lt;p&gt;This assessment helps determine whether the project needs a routine update, a broader modernization effort, or selected parts of the application to be rebuilt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can Angular Be Upgraded Across Several Versions at Once?
&lt;/h2&gt;

&lt;p&gt;Angular recommends upgrading one major version at a time.&lt;/p&gt;

&lt;p&gt;An application moving from Angular 19 to Angular 22 should normally follow this path:&lt;/p&gt;

&lt;p&gt;*&lt;em&gt;Angular 19 → Angular 20 → Angular 21 → Angular 22&lt;br&gt;
*&lt;/em&gt;&lt;br&gt;
Angular’s migration tooling supports updates between adjacent major versions. The official guidance says that the version being upgraded should be within one major version of the target. Angular Update Guide&lt;/p&gt;

&lt;p&gt;This staged approach makes problems easier to isolate. Jumping across several versions at once can combine dependency conflicts, TypeScript changes, removed APIs, and test failures into one difficult debugging process.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Happens During an Upgrade Project?
&lt;/h2&gt;

&lt;p&gt;The team should first document the current versions, architecture, package dependencies, business-critical workflows, and known defects.&lt;/p&gt;

&lt;p&gt;Next, developers identify compatible versions of Angular, Node.js, TypeScript, MongoDB drivers, and third-party packages. Unsupported libraries may need to be replaced before the main framework upgrade can continue.&lt;/p&gt;

&lt;p&gt;The application is then upgraded in controlled stages. After each stage, developers run unit tests, integration tests, build checks, and important user journeys.&lt;/p&gt;

&lt;p&gt;The updated application should be released to a staging environment before production. Performance, errors, API behaviour, authentication, database operations, and customer-facing workflows should be monitored after deployment.&lt;/p&gt;

&lt;p&gt;An upgrade is complete only when the application works reliably in production. Successfully installing new package versions is not enough.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should You Upgrade or Rewrite the Application?
&lt;/h2&gt;

&lt;p&gt;In most cases, an incremental upgrade is safer than a complete rewrite.&lt;/p&gt;

&lt;p&gt;A rewrite creates new risks around feature parity, data migration, testing, delivery time, and business continuity. Existing business rules that took years to develop may be overlooked during reconstruction.&lt;/p&gt;

&lt;p&gt;A rewrite may be appropriate when the current architecture blocks essential changes, core dependencies have no supported upgrade path, serious security problems are deeply embedded, or maintaining the existing code costs more than replacing it.&lt;/p&gt;

&lt;p&gt;The decision should be based on a technical audit rather than the age of the application alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Should You Ask Before Hiring MEAN Stack Developers?
&lt;/h2&gt;

&lt;p&gt;Ask developers how they would assess the application before providing a final estimate.&lt;/p&gt;

&lt;p&gt;They should be able to explain their approach to incremental Angular upgrades, Node.js LTS migration, dependency conflicts, automated testing, rollback planning, staging, and production monitoring.&lt;/p&gt;

&lt;p&gt;Request examples of previous upgrade projects and ask what unexpected problems occurred. Experience resolving real dependency, testing, and deployment issues is often more valuable than familiarity with a long list of technologies.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Can Spaculus Help?
&lt;/h2&gt;

&lt;p&gt;Spaculus Software provides MEAN Stack developers for Angular interfaces, Node.js and Express APIs, MongoDB applications, performance optimization, and legacy-system modernization.&lt;/p&gt;

&lt;p&gt;The team can audit an existing application, prepare a staged upgrade plan, replace unsupported dependencies, improve testing, optimize APIs, and support production deployment. Spaculus also provides custom software, SaaS, Cloud and DevOps, UI/UX, QA, mobile, and AI development services when the upgrade is part of a broader product roadmap.&lt;/p&gt;

&lt;p&gt;Businesses can &lt;a href="https://spaculus.com/services/hire-mean-stack-developers/" rel="noopener noreferrer"&gt;hire MEAN Stack developers&lt;/a&gt; from Spaculus without automatically rebuilding the complete application.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;p&gt;MEAN Stack developers can upgrade Angular-only, Node.js-only, or complete MEAN applications.&lt;br&gt;
Angular 19 and earlier versions are unsupported as of August 2026.&lt;br&gt;
Node.js 20 and earlier versions have reached end of life.&lt;br&gt;
Angular upgrades should normally move through one major version at a time.&lt;br&gt;
A technical audit should determine whether the application needs an update, modernization, or rewrite.&lt;/p&gt;

&lt;p&gt;Upgrading an application is rarely just an npm install command. A safe project begins by understanding what the current system does, which workflows users cannot afford to lose, and how every change will be tested before production.&lt;/p&gt;

</description>
      <category>node</category>
      <category>angular</category>
      <category>meanstack</category>
    </item>
  </channel>
</rss>
