<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Muhammad Umair Ashraf</title>
    <description>The latest articles on DEV Community by Muhammad Umair Ashraf (@omayrashraf).</description>
    <link>https://dev.to/omayrashraf</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4134004%2F5d108082-c685-44a3-9afd-f94e943f3e26.jpg</url>
      <title>DEV Community: Muhammad Umair Ashraf</title>
      <link>https://dev.to/omayrashraf</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/omayrashraf"/>
    <language>en</language>
    <item>
      <title>Designing a Simulation-First Safety Gate for Strands Robots</title>
      <dc:creator>Muhammad Umair Ashraf</dc:creator>
      <pubDate>Thu, 08 Oct 2026 08:00:19 +0000</pubDate>
      <link>https://dev.to/omayrashraf/designing-a-simulation-first-safety-gate-for-strands-robots-4o9p</link>
      <guid>https://dev.to/omayrashraf/designing-a-simulation-first-safety-gate-for-strands-robots-4o9p</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# What If a Robot Had to Earn Permission Before Reaching Real Hardware? 🤖&lt;/span&gt;

Physical AI changes the meaning of failure.

For an AI chatbot, a bad response might waste a few seconds.

For a robot, a bad action could cause a collision, damage hardware, drop an object, or create an unsafe situation.

That made me think about a different approach to Strands Robots:
&lt;span class="gt"&gt;
&amp;gt; What if a robot policy had to pass a safety gate before it was allowed to run on real hardware?&lt;/span&gt;

Instead of:

&lt;span class="p"&gt;```&lt;/span&gt;&lt;span class="nl"&gt;text
&lt;/span&gt;Strands Agent → Robot
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I would like to see a workflow like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Strands Agent
      ↓
Simulation
      ↓
Episode Recording
      ↓
Safety Evaluation
      ↓
┌─────────┬─────────┬─────────┐
│  PASS   │  FAIL   │ UNKNOWN │
└────┬────┴────┬────┴────┬────┘
     ↓         ↓         ↓
 Hardware    Reject    Human
  Test                  Review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What should the safety gate evaluate?
&lt;/h2&gt;

&lt;p&gt;For each recorded robot episode, I would like to evaluate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Workspace boundaries&lt;/li&gt;
&lt;li&gt;Collision/contact events&lt;/li&gt;
&lt;li&gt;Joint velocity&lt;/li&gt;
&lt;li&gt;Execution time&lt;/li&gt;
&lt;li&gt;Action stability&lt;/li&gt;
&lt;li&gt;Task completion&lt;/li&gt;
&lt;li&gt;Unexpected behavior&lt;/li&gt;
&lt;li&gt;Camera/video evidence&lt;/li&gt;
&lt;li&gt;Robot telemetry&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Some checks should be deterministic.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;joint_velocity&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;MAX_VELOCITY&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;FAIL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;collision_detected&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;FAIL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;outside_workspace&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;FAIL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Other questions could use an AI evaluator:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Did the robot actually accomplish the requested task?&lt;/p&gt;

&lt;p&gt;Was the behavior consistent with the task?&lt;/p&gt;

&lt;p&gt;Is there enough evidence to approve this episode?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The part I find most interesting: UNKNOWN
&lt;/h2&gt;

&lt;p&gt;I don't think physical AI should only have:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PASS
FAIL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It should also have:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;UNKNOWN
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the system doesn't have enough evidence to determine whether an episode is safe, it should not automatically approve it.&lt;/p&gt;

&lt;p&gt;Instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;UNKNOWN → Human Review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This creates a much safer progression from simulation to physical hardware.&lt;/p&gt;

&lt;h2&gt;
  
  
  Connecting Agent Evaluation and Robot Evaluation
&lt;/h2&gt;

&lt;p&gt;There are actually two things we need to evaluate.&lt;/p&gt;

&lt;h3&gt;
  
  
  Agent
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Did it select the correct tool?&lt;/li&gt;
&lt;li&gt;Did it generate the correct instruction?&lt;/li&gt;
&lt;li&gt;Did it complete the requested goal?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Robot
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Did it stay within the workspace?&lt;/li&gt;
&lt;li&gt;Did it avoid collisions?&lt;/li&gt;
&lt;li&gt;Did it respect movement limits?&lt;/li&gt;
&lt;li&gt;Did it behave safely?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The combined workflow could look like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;              Strands Agent
                    ↓
             Amazon Bedrock
                    ↓
              Robot Tool
                    ↓
              Simulation
                    ↓
            Record Episode
                    ↓
          Safety Evaluation
                    ↓
          ┌─────────┴─────────┐
         FAIL               PASS
          ↓                   ↓
       Reject           Agent Evaluation
                              ↓
                       Quality Threshold
                              ↓
                        Hardware Test
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Amazon Bedrock AgentCore Evaluations already provides online, on-demand, and batch evaluation, as well as custom evaluators and code-based evaluators. That makes the agent-evaluation side increasingly practical. (&lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/evaluations-types.html?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;AWS Documentation&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;The opportunity I see is going one step further for physical AI:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;simulation → episode → safety evaluation → approval → hardware&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters
&lt;/h2&gt;

&lt;p&gt;A successful simulation should not automatically mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Ship it to the robot."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead, it should mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The behavior has passed the defined safety and quality gates."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And failures should become useful evaluation data:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Run
 ↓
Evaluate
 ↓
Find Failure
 ↓
Record Failure
 ↓
Improve Policy
 ↓
Run Again
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That creates a continuous improvement loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  My AWS Wishlist idea
&lt;/h2&gt;

&lt;p&gt;I would love to see a native &lt;strong&gt;Simulation-to-Hardware Safety Gate for Strands Robots&lt;/strong&gt;, integrated with Amazon Bedrock AgentCore Evaluations.&lt;/p&gt;

&lt;p&gt;It could provide:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Configurable safety policies&lt;/li&gt;
&lt;li&gt;PASS / FAIL / UNKNOWN decisions&lt;/li&gt;
&lt;li&gt;Episode-level evidence&lt;/li&gt;
&lt;li&gt;Human approval workflows&lt;/li&gt;
&lt;li&gt;Robot telemetry + video evaluation&lt;/li&gt;
&lt;li&gt;Custom safety evaluators&lt;/li&gt;
&lt;li&gt;CloudWatch integration&lt;/li&gt;
&lt;li&gt;Evaluation history&lt;/li&gt;
&lt;li&gt;Regression testing across recorded episodes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal isn't to prevent AI agents from being autonomous.&lt;/p&gt;

&lt;p&gt;It's to make autonomy &lt;strong&gt;progressive and measurable&lt;/strong&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Don't let the robot be the first place where we discover that an AI policy is unsafe.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Start in simulation.&lt;/p&gt;

&lt;p&gt;Evaluate the behavior.&lt;/p&gt;

&lt;p&gt;Learn from failures.&lt;/p&gt;

&lt;p&gt;Then earn the right to reach the real world.&lt;/p&gt;

&lt;h1&gt;
  
  
  AWS #Strands #PhysicalAI #Robotics #AmazonBedrock #AgenticAI
&lt;/h1&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
This version is much better suited to **DEV.to as a community post**: it presents an idea, explains the technical gap, gives a concrete architecture, and ends with a discussion-worthy AWS improvement rather than reading like a formal documentation article.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>aws</category>
      <category>ai</category>
      <category>robotics</category>
      <category>agenticai</category>
    </item>
  </channel>
</rss>
