<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: ZedIoT</title>
    <description>The latest articles on DEV Community by ZedIoT (@zediot).</description>
    <link>https://dev.to/zediot</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3399208%2Fd2ba7c47-d4b2-4057-969d-cee59480b9eb.png</url>
      <title>DEV Community: ZedIoT</title>
      <link>https://dev.to/zediot</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/zediot"/>
    <language>en</language>
    <item>
      <title>RKNN ONNX Opset Compatibility: Constraints, Failure Patterns, and Baselines for Edge NPU Deployment</title>
      <dc:creator>ZedIoT</dc:creator>
      <pubDate>Thu, 01 Oct 2026 12:30:00 +0000</pubDate>
      <link>https://dev.to/zediot/rknn-onnx-opset-compatibility-constraints-failure-patterns-and-baselines-for-edge-npu-deployment-3kho</link>
      <guid>https://dev.to/zediot/rknn-onnx-opset-compatibility-constraints-failure-patterns-and-baselines-for-edge-npu-deployment-3kho</guid>
      <description>&lt;h1&gt;
  
  
  RKNN ONNX Opset Compatibility: Constraints, Failure Patterns, and Baselines for Edge NPU Deployment
&lt;/h1&gt;

&lt;p&gt;Most edge AI projects start with a simple question: can the model be exported to ONNX and run on the board? Success at that stage is usually defined as "the first demo works." Then real delivery begins, and the nature of the problem changes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The model needs structural tweaks for a new scenario.&lt;/li&gt;
&lt;li&gt;The algorithm team upgrades the base framework or model version.&lt;/li&gt;
&lt;li&gt;The same product line must reuse the model across multiple SoCs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now the ONNX opset — previously treated as a neutral intermediate format — becomes a hard engineering constraint. Whether a model can keep evolving is often decided not by accuracy or compute, but by whether the conversion pipeline stays stable. And "stable" here doesn't mean "can it convert today" — it means "will it stay controllable over the next 6–12 months?"&lt;/p&gt;

&lt;p&gt;In RKNN scenarios, opset selection is effectively locking in your future engineering freedom in advance.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Why ONNX Generality Breaks Down on NPUs
&lt;/h2&gt;

&lt;p&gt;ONNX was designed to solve cross-framework model exchange — not to guarantee executability on specific hardware. That works fine in CPU/GPU ecosystems because:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Runtimes can rely on kernel fallback paths.&lt;/li&gt;
&lt;li&gt;Graph optimizations and operator fusion can be adjusted at runtime.&lt;/li&gt;
&lt;li&gt;There is significant buffer space between operator semantics and execution.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In NPU scenarios, most of those assumptions no longer hold. NPUs behave much closer to ASICs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Supported operator sets are limited and fixed.&lt;/li&gt;
&lt;li&gt;Tensor shapes, layouts, and operator combinations have strict constraints.&lt;/li&gt;
&lt;li&gt;There is no "run first and fix later" runtime compromise.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A fully valid ONNX model — even one verified on CPU — can still be outright rejected during NPU conversion. On Rockchip platforms, RKNN's job is not to "interpret ONNX graphs as best as possible," but to compile ONNX graphs into static, NPU-executable representations. This isn't a toolchain maturity issue; it's a structural mismatch between generic IRs and hardware execution models.&lt;/p&gt;

&lt;h3&gt;
  
  
  What opset really means in RKNN projects
&lt;/h3&gt;

&lt;p&gt;For Rockchip NPUs, the conversion stage must decide up front:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Whether every operator has a hardware mapping&lt;/li&gt;
&lt;li&gt;Whether operator attributes satisfy NPU constraints&lt;/li&gt;
&lt;li&gt;Whether the entire graph can be fully offloaded to the NPU&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Opset is no longer just a syntax version — it becomes an upstream constraint on how the graph is expressed. Across different opsets, the same operator may differ in attribute definitions, default behavior, or shape inference rules, and RKNN amplifies those differences at compile time. Opset selection isn't a parameter you can casually roll back. It's a platform-level decision: once fixed, the freedom of future model structures is implicitly constrained.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Why RKNN Behaves Like a Compiler
&lt;/h2&gt;

&lt;p&gt;On paper the pipeline looks linear:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PyTorch → ONNX (with opset) → RKNN Toolkit → NPU Binary&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In practice, success is determined not by the linear flow but by what information is preserved or lost at each stage. The most fragile — and irreversible — step is &lt;strong&gt;ONNX → RKNN&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Once inside RKNN conversion, the model is no longer treated as a dynamically interpretable graph. It must become a fully compilable static structure. Any node that cannot be mapped to the NPU fails the entire conversion — there is no partial fallback.&lt;/p&gt;

&lt;p&gt;Unlike many GPU inference engines, RKNN resolves all operator mappings at compile time, performs no runtime operator substitution, and treats conversion failure as an invalid design assumption. That's why engineers new to RKNN often find it "overly strict." GPU-era intuition — "if an operator isn't supported, it'll just be slower" — does not apply.&lt;/p&gt;

&lt;p&gt;This strictness is not a flaw; it's the price paid for determinism and efficiency. Once conversion succeeds, execution paths, latency, and resource usage become highly predictable.&lt;/p&gt;

&lt;p&gt;In the ONNX ecosystem, newer opsets usually mean more flexible operator definitions, richer attribute combinations, and better dynamic-shape semantics. In RKNN scenarios, those "improvements" often introduce uncertainty — new opsets may expose attributes RKNN doesn't support or change default behaviors, leading to immediate unsupported-attribute errors, models that convert but behave incorrectly at runtime, and dramatically different stability across opsets for the same model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Newer opsets are not necessarily better. Verified opsets are safer.&lt;/strong&gt; Stability comes from well-defined constraints, not maximal expressiveness.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. ONNX → RKNN Failure Patterns in Practice
&lt;/h2&gt;

&lt;p&gt;Most failures aren't caused by exotic operators — they're caused by how model structures are expressed. They stem from mismatches between toolchain assumptions and model design assumptions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Conversion-time failure vs runtime anomaly
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Conversion-Time Failure&lt;/th&gt;
&lt;th&gt;Runtime Anomaly&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;When it occurs&lt;/td&gt;
&lt;td&gt;ONNX → RKNN conversion&lt;/td&gt;
&lt;td&gt;NPU inference runtime&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical symptom&lt;/td&gt;
&lt;td&gt;Unsupported op / attribute&lt;/td&gt;
&lt;td&gt;Incorrect outputs, accuracy collapse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Debug difficulty&lt;/td&gt;
&lt;td&gt;Relatively clear&lt;/td&gt;
&lt;td&gt;Extremely high&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avoidable?&lt;/td&gt;
&lt;td&gt;Yes, via structural constraints&lt;/td&gt;
&lt;td&gt;Very hard, often requires redesign&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Engineering risk&lt;/td&gt;
&lt;td&gt;Exposed early&lt;/td&gt;
&lt;td&gt;Late-stage "time bomb"&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In practice, the most dangerous situation is not "can't convert" — it's "converts successfully but produces unreliable results."&lt;/p&gt;

&lt;h3&gt;
  
  
  High-risk structural patterns
&lt;/h3&gt;

&lt;p&gt;These are perfectly legal in ONNX but problematic for NPUs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Dynamic shape propagation&lt;/strong&gt; — shapes cannot be resolved at compile time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stacked reshape / permute chains&lt;/strong&gt; — data layouts cannot be mapped to fixed hardware paths&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Post-processing logic embedded in detection heads&lt;/strong&gt; — operator fusion limits exceeded&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Implicit broadcast behaviors&lt;/strong&gt; — layout assumptions clash with NPU paths&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The issue isn't that ONNX is "wrong." It's that NPUs require fully deterministic graphs.&lt;/p&gt;

&lt;h3&gt;
  
  
  The hidden combination risk
&lt;/h3&gt;

&lt;p&gt;A frequently underestimated reality: &lt;strong&gt;an opset can be valid, a model structure can be valid, and the combination still fails.&lt;/strong&gt; Opset changes may alter default operator behavior or attribute expression, directly affecting RKNN's compile-time decisions.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Combination&lt;/th&gt;
&lt;th&gt;Surface Status&lt;/th&gt;
&lt;th&gt;Actual Risk&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;New opset + dynamic shape&lt;/td&gt;
&lt;td&gt;ONNX-valid&lt;/td&gt;
&lt;td&gt;Compile-time indeterminacy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New opset + complex detection head&lt;/td&gt;
&lt;td&gt;Exportable&lt;/td&gt;
&lt;td&gt;NPU mapping failure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Old opset + simplified structure&lt;/td&gt;
&lt;td&gt;Conservative&lt;/td&gt;
&lt;td&gt;Highest stability&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is why many teams find that rolling back opset restores control rather than "downgrading capability."&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Why YOLOv8 Tensions With RKNN
&lt;/h2&gt;

&lt;p&gt;YOLOv8 is not "unsuitable" for RKNN — but its design goals inherently conflict with NPU execution models.&lt;/p&gt;

&lt;p&gt;YOLOv8 exhibits several engineering traits: highly modular head structures, heavy use of reshape / concat / split, friendly support for dynamic input sizes, and increasingly integrated post-processing. These are strengths on GPU/CPU, but they significantly increase compile-time complexity on NPUs. The breaking points are not sporadic bugs — they're direct manifestations of design mismatch.&lt;/p&gt;

&lt;h3&gt;
  
  
  Risk by task type
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task Type&lt;/th&gt;
&lt;th&gt;Risk Level&lt;/th&gt;
&lt;th&gt;Engineering Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Detection&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Head complexity must be controlled&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Segmentation&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Mask branches are structurally complex&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pose&lt;/td&gt;
&lt;td&gt;Very High&lt;/td&gt;
&lt;td&gt;Keypoint dimensions are highly dynamic&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This doesn't mean YOLOv8 is "bad" — it means NPU compilation was not its primary design target.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Engineering Tradeoffs: Model Freedom vs NPU Determinism
&lt;/h2&gt;

&lt;p&gt;Once the failure mechanisms are clear, the real question becomes: do you keep forcing models through RKNN, or redesign the system with NPU constraints as first-class citizens?&lt;/p&gt;

&lt;p&gt;Discussions about "RKNN adaptation" often mask a deeper question: what are you optimizing — model freedom or delivery certainty?&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If your product requires frequent structural iteration, you need evolution space.&lt;/li&gt;
&lt;li&gt;If your product demands predictable latency, power, and cost, you need determinism.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;RKNN's value lies not in flexibility, but in predictability.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Focus&lt;/th&gt;
&lt;th&gt;GPU/CPU-Friendly ONNX&lt;/th&gt;
&lt;th&gt;RKNN/NPU-Friendly&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model iteration&lt;/td&gt;
&lt;td&gt;High freedom&lt;/td&gt;
&lt;td&gt;Constrained upfront&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Performance predictability&lt;/td&gt;
&lt;td&gt;Runtime-dependent&lt;/td&gt;
&lt;td&gt;Highly stable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Debugging&lt;/td&gt;
&lt;td&gt;Rich tools&lt;/td&gt;
&lt;td&gt;Constraint-driven&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mass production stability&lt;/td&gt;
&lt;td&gt;Version-sensitive&lt;/td&gt;
&lt;td&gt;Strong once converted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Team coordination&lt;/td&gt;
&lt;td&gt;Algorithm-led&lt;/td&gt;
&lt;td&gt;Joint algorithm–engineering&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A counterintuitive but common conclusion: &lt;strong&gt;in RKNN projects, it is often cheaper to design for hardware early than to patch errors later.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Opset locking and product lifecycle
&lt;/h3&gt;

&lt;p&gt;In RKNN projects, opset functions like an interface contract. Once validated, upgrades must be treated like system dependency upgrades. The typical lifecycle:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;PoC&lt;/strong&gt; — make it run; pick a workable opset&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MVP&lt;/strong&gt; — lock structure and prioritize stability&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production&lt;/strong&gt; — freeze opset, tools, export scripts&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Iteration&lt;/strong&gt; — move variability to the system layer&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Which systems fit RKNN
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System Type&lt;/th&gt;
&lt;th&gt;Fit&lt;/th&gt;
&lt;th&gt;Reason&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Single-task, stable detection/classification&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Determinism pays off&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frequent A/B testing, algorithm-driven&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Toolchain limits iteration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dynamic input sizes / batch&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Compile-time fixation hard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Power- and cost-constrained edge products&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;NPU advantages realized&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Heavy in-graph post-processing&lt;/td&gt;
&lt;td&gt;Medium–Low&lt;/td&gt;
&lt;td&gt;Requires refactoring&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A practical rule of thumb: if iteration comes from rules and thresholds, RKNN is friendly. If it comes from model structure, RKNN becomes a production line requiring dedicated maintenance.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. When to Stop Forcing RKNN Adaptation
&lt;/h2&gt;

&lt;p&gt;When two or three of the following signals appear, the return on continued adaptation is usually starting to decline:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Every small model change introduces new incompatible nodes that can't be resolved through local replacements.&lt;/li&gt;
&lt;li&gt;You're writing more and more export-specific scripts and only a few people on the team can maintain them.&lt;/li&gt;
&lt;li&gt;Conversion technically succeeds, but inference anomalies can't be consistently reproduced or explained (the most dangerous case).&lt;/li&gt;
&lt;li&gt;The roadmap requires frequent backbone/head changes or new task branches (e.g., detection → segmentation or pose).&lt;/li&gt;
&lt;li&gt;Version upgrades turn into a "game of chance" with no repeatable validation baseline.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When that happens, the pragmatic approach is usually a binary choice: converge the model structure toward an NPU-friendly form, or shrink the role of the NPU so it handles only what it's good at.&lt;/p&gt;

&lt;h3&gt;
  
  
  Model design principles for RKNN
&lt;/h3&gt;

&lt;p&gt;The value of these principles isn't that they "sound right" — it's that they reduce organizational friction, giving algorithm and engineering teams a shared language around the same constraints.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prefer statically determinable shape paths.&lt;/strong&gt; Avoid bringing dynamic behavior into the NPU compilation stage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Minimize stacked permute / reshape operations&lt;/strong&gt;, especially near the head.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Place post-processing outside the model&lt;/strong&gt; whenever possible (CPU or lightweight operators), and treat NPU output as raw prediction tensors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Establish traceable baselines&lt;/strong&gt; for opset, export scripts, and toolchain versions to avoid "same model name, different graph."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat "can be compiled by the NPU" as an acceptance criterion&lt;/strong&gt;, not "the error was patched."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These points sound conservative, but they often determine whether, at mass-production time, you're reusing a stable pipeline or firefighting every week.&lt;/p&gt;

&lt;h3&gt;
  
  
  A practical early-stage validation method
&lt;/h3&gt;

&lt;p&gt;Early in a project, the most effective strategy is not to push accuracy to the limit immediately, but to first establish a stable, repeatable validation loop:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fix the export entry point&lt;/strong&gt; — same PyTorch commit + same export script + same opset&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix reference inputs&lt;/strong&gt; — a small set of repeatable sample tensors so data noise doesn't affect judgments&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix conversion outputs&lt;/strong&gt; — record RKNN conversion logs, graph optimization summaries, quantization configs, and artifact hashes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix on-device validation&lt;/strong&gt; — at minimum, capture output tensor statistics (min / max / mean / distribution); don't rely only on visual inspection&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix regression gates&lt;/strong&gt; — every model change must first pass "compilable + output consistency" before discussing accuracy improvements&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Once this baseline is in place, opset selection stops being a matter of experience or guesswork — it becomes locked in by evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Common Errors → Structural Causes → Strategies
&lt;/h2&gt;

&lt;p&gt;Error messages vary across RKNN Toolkit versions, SoCs, and ONNX exporters. Grouping by typical keywords speeds up root-cause identification.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Error Keyword&lt;/th&gt;
&lt;th&gt;Likely Structural Cause&lt;/th&gt;
&lt;th&gt;Engineering Strategy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Unsupported operator&lt;/td&gt;
&lt;td&gt;NPU does not support op or attribute combination&lt;/td&gt;
&lt;td&gt;Replace structure, offload subgraph, redesign head&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Attribute not supported&lt;/td&gt;
&lt;td&gt;Opset introduced unsupported attributes&lt;/td&gt;
&lt;td&gt;Roll back opset, adjust export params&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cannot infer shape&lt;/td&gt;
&lt;td&gt;Dynamic shapes in critical path&lt;/td&gt;
&lt;td&gt;Fix input size, remove -1, simplify head&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Concat axis mismatch&lt;/td&gt;
&lt;td&gt;Feature map misalignment&lt;/td&gt;
&lt;td&gt;Align branches, reduce cross-scale concat&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reshape failed&lt;/td&gt;
&lt;td&gt;Dynamic target shapes&lt;/td&gt;
&lt;td&gt;Use static shapes or move reshape outside&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transpose not supported&lt;/td&gt;
&lt;td&gt;Excessive layout changes&lt;/td&gt;
&lt;td&gt;Unify layout early, move permutes outside&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gather / Scatter&lt;/td&gt;
&lt;td&gt;Index-based ops in graph&lt;/td&gt;
&lt;td&gt;Externalize logic to CPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NonMaxSuppression&lt;/td&gt;
&lt;td&gt;NMS embedded in model&lt;/td&gt;
&lt;td&gt;Always externalize NMS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TopK / Sort&lt;/td&gt;
&lt;td&gt;Sorting in post-processing&lt;/td&gt;
&lt;td&gt;Replace with thresholds or external logic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reduce* issues&lt;/td&gt;
&lt;td&gt;Unsupported axis combinations&lt;/td&gt;
&lt;td&gt;Restructure reduce or replace with pooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pad not supported&lt;/td&gt;
&lt;td&gt;Complex padding modes&lt;/td&gt;
&lt;td&gt;Use constant pad or redesign&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resize not supported&lt;/td&gt;
&lt;td&gt;Unsupported interpolation&lt;/td&gt;
&lt;td&gt;Use nearest or external resize&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quantization failed&lt;/td&gt;
&lt;td&gt;Calibration mismatch&lt;/td&gt;
&lt;td&gt;Align data, FP first, mixed precision&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Large accuracy drop&lt;/td&gt;
&lt;td&gt;Quantization or numeric mismatch&lt;/td&gt;
&lt;td&gt;Layer-wise comparison, redesign sensitive heads&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Final Thought
&lt;/h2&gt;

&lt;p&gt;RKNN / ONNX opset compatibility is not just a toolchain issue — it's an engineering contract problem. Once this constraint is understood and controlled, NPU deployment becomes predictable instead of fragile.&lt;/p&gt;

&lt;p&gt;The more expressive freedom you demand from the model, the harder it becomes for static NPU backends to guarantee executability. Once you accept constraints and push variability into the system layer, the deterministic advantages of NPUs can finally be realized.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q1: Why does RKNN ONNX opset compatibility cause conversion failures?&lt;/strong&gt;&lt;br&gt;
Because RKNN compiles ONNX models into a static NPU execution graph. Many ONNX opsets introduce dynamic semantics or attributes that cannot be resolved at compile time, causing failures even when the model is ONNX-valid.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q2: Why can an ONNX model run on CPU but fail on an RKNN NPU?&lt;/strong&gt;&lt;br&gt;
CPU runtimes allow dynamic execution and operator fallback at runtime, while RKNN requires all operators, shapes, and attributes to be fully determined during compilation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q3: Which ONNX opset should be used with RKNN Toolkit2?&lt;/strong&gt;&lt;br&gt;
A verified opset already proven compatible with the target Rockchip NPU and RKNN Toolkit version. Newer opsets often increase conversion risk rather than improving stability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q4: Why does YOLOv8 frequently fail when converted to RKNN?&lt;/strong&gt;&lt;br&gt;
YOLOv8 relies heavily on dynamic reshape, concat operations, and embedded post-processing logic, which conflict with the static graph and compile-time constraints required by RKNN.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q5: When should teams stop forcing RKNN compatibility?&lt;/strong&gt;&lt;br&gt;
When repeated model changes introduce non-local failures, inference becomes unstable or unexplainable, or opset upgrades lack a reproducible validation baseline.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Where has RKNN's compile-time strictness bitten you hardest — an opset upgrade that broke a working export, or a model that converted cleanly but produced wrong outputs at runtime? And if you've been through it, what's your minimum baseline before you trust a converted model?&lt;/em&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>iot</category>
      <category>computervision</category>
      <category>python</category>
    </item>
    <item>
      <title>ESP32-S3 Edge AI in Practice: Deep Optimization of TensorFlow Lite Micro Inference Performance</title>
      <dc:creator>ZedIoT</dc:creator>
      <pubDate>Tue, 29 Sep 2026 12:30:00 +0000</pubDate>
      <link>https://dev.to/zediot/esp32-s3-edge-ai-in-practice-deep-optimization-of-tensorflow-lite-micro-inference-performance-4b0g</link>
      <guid>https://dev.to/zediot/esp32-s3-edge-ai-in-practice-deep-optimization-of-tensorflow-lite-micro-inference-performance-4b0g</guid>
      <description>&lt;h1&gt;
  
  
  ESP32-S3 Edge AI in Practice: Deep Optimization of TensorFlow Lite Micro Inference Performance
&lt;/h1&gt;

&lt;p&gt;Getting a TensorFlow Lite Micro model to &lt;em&gt;run&lt;/em&gt; on an ESP32-S3 is easy. Getting it to run &lt;em&gt;fast enough for real-time work&lt;/em&gt; is a different problem entirely — one that comes down to three things: full INT8 quantization, ESP-NN hardware acceleration, and careful SRAM/PSRAM placement. This article walks through each of those levers, with real benchmark numbers and the pitfalls that silently eat your performance.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Why Run TensorFlow Lite Micro on ESP32-S3?
&lt;/h2&gt;

&lt;p&gt;In TinyML systems, compute limits are always the main constraint. Compared to earlier chips like ESP32 or ESP32-S2, ESP32-S3 introduces major improvements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Dual-core Xtensa® 32-bit LX7&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dedicated vector instructions&lt;/strong&gt; for AI workloads&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  1.1 Hardware-Level AI Acceleration
&lt;/h3&gt;

&lt;p&gt;ESP32-S3 supports SIMD operations, allowing multiple 8-bit or 16-bit MAC operations in a single clock cycle. For convolution and fully connected layers, this delivers &lt;strong&gt;5–10× inference speedup&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.2 Balanced Memory Architecture
&lt;/h3&gt;

&lt;p&gt;TFLM is designed for devices with less than 1 MB RAM. ESP32-S3 provides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;512 KB on-chip SRAM&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Up to 1 GB external Flash&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Optional PSRAM expansion&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This flexible design allows larger models, such as lightweight MobileNet or custom CNNs, without sacrificing accuracy.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.3 Seamless Ecosystem Integration
&lt;/h3&gt;

&lt;p&gt;Espressif's &lt;strong&gt;esp-nn&lt;/strong&gt; library is deeply integrated into TFLM. When using standard TFLM APIs, optimized ESP32-S3 kernels are automatically selected — no hand-written assembly required.&lt;/p&gt;

&lt;p&gt;In real deployments, ESP32-S3 marks the shift from "barely usable" MCU AI to production-grade edge inference, and is one of the most cost-effective edge computing platforms for AIoT applications.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. TensorFlow Lite Micro Architecture Overview
&lt;/h2&gt;

&lt;p&gt;TensorFlow Lite Micro is a stripped-down version of TFLite that runs directly on bare metal or RTOS, without Linux dependencies. Understanding its architecture is key to optimization.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.1 Core Components
&lt;/h3&gt;

&lt;p&gt;TFLM consists of four main parts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Interpreter&lt;/strong&gt; — controls graph execution, memory allocation, and operator dispatch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Op Resolver&lt;/strong&gt; — defines which operators are included. Only required ops should be enabled.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tensor Arena&lt;/strong&gt; — a static memory region used for intermediate tensors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kernels&lt;/strong&gt; — mathematical implementations. ESP32-S3 replaces reference kernels with optimized ones.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2.2 Inference Lifecycle
&lt;/h3&gt;

&lt;p&gt;The full TFLM inference workflow on ESP32-S3 follows this sequence: load model → allocate tensors → fill input → invoke → read output.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key note:&lt;/strong&gt; &lt;code&gt;AllocateTensors&lt;/code&gt; calculates tensor lifetimes and reuses memory. Tensor Arena size must be carefully tuned to the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. From Keras Model to ESP32-S3 Firmware
&lt;/h2&gt;

&lt;p&gt;Deploying a model requires compression and conversion.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.1 Model Training and Conversion
&lt;/h3&gt;

&lt;p&gt;Models are trained in TensorFlow/Keras and exported as &lt;code&gt;.h5&lt;/code&gt; or SavedModel, then converted to &lt;code&gt;.tflite&lt;/code&gt; using the TFLite Converter. &lt;strong&gt;Quantization is mandatory.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3.2 Why INT8 Quantization Matters
&lt;/h3&gt;

&lt;p&gt;ESP32-S3 hardware acceleration is optimized for INT8. Converting an FP32 model to INT8 offers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;75% smaller model size&lt;/strong&gt; — parameters shrink from 4 bytes to 1 byte.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;4–10× faster inference&lt;/strong&gt; — avoids expensive floating-point operations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lower power consumption&lt;/strong&gt; — integer arithmetic is far more energy-efficient than floating-point.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3.3 ESP-IDF Integration
&lt;/h3&gt;

&lt;p&gt;In ESP-IDF, TFLM is included as a component. The &lt;code&gt;.tflite&lt;/code&gt; model is converted into a C array using &lt;code&gt;xxd&lt;/code&gt; and linked into firmware:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;unsigned&lt;/span&gt; &lt;span class="kt"&gt;char&lt;/span&gt; &lt;span class="n"&gt;g_model&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="mh"&gt;0x1c&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mh"&gt;0x00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mh"&gt;0x00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="n"&gt;tflite&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;MicroMutableOpResolver&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;resolver&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="n"&gt;resolver&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddConv2D&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="n"&gt;resolver&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddFullyConnected&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="n"&gt;tflite&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;MicroInterpreter&lt;/span&gt; &lt;span class="nf"&gt;interpreter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;resolver&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tensor_arena&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;kTensorArenaSize&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  4. ESP-NN: Unlocking ESP32-S3 Performance
&lt;/h2&gt;

&lt;p&gt;If you use the standard open-source TFLM library without optimization, inference runs on the Xtensa core using generic instructions — it does not fully utilize the ESP32-S3's hardware.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ESP-NN&lt;/strong&gt; is Espressif's low-level library optimized specifically for AI inference. It provides hand-written assembly optimizations for high-frequency operators such as convolution, pooling, and activation functions (e.g., ReLU).&lt;/p&gt;

&lt;p&gt;During compilation, TFLM detects the target hardware platform. If it identifies ESP32-S3, it automatically replaces the default Reference Kernels with optimized ESP-NN Kernels.&lt;/p&gt;

&lt;p&gt;In a standard 2D convolution benchmark, enabling ESP-NN acceleration made the ESP32-S3 approximately &lt;strong&gt;7.2× faster&lt;/strong&gt; compared to running without optimization. This directly impacts the feasibility of real-time voice processing and high-frame-rate gesture recognition.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Memory Optimization: Coordinating SRAM and PSRAM
&lt;/h2&gt;

&lt;p&gt;When deploying TFLM on ESP32-S3, memory (RAM) is often more limited than compute power. The ESP32-S3 provides about 512 KB of on-chip SRAM — very fast, but quickly insufficient for vision models.&lt;/p&gt;

&lt;p&gt;Balancing internal SRAM and external PSRAM is critical for TinyML performance.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.1 Static Allocation Strategy for Tensor Arena
&lt;/h3&gt;

&lt;p&gt;TFLM uses a continuous memory block called the &lt;strong&gt;Tensor Arena&lt;/strong&gt; to store all intermediate tensors during inference.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prioritize on-chip SRAM&lt;/strong&gt; — for small models (audio recognition, sensor classification), allocate the entire Tensor Arena in internal SRAM for the lowest latency.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PSRAM expansion strategy&lt;/strong&gt; — for models like Person Detection with large feature maps, allocate the Tensor Arena in external PSRAM. PSRAM is slightly slower (accessed via SPI/Octal), but the ESP32-S3 cache mechanism reduces the performance impact.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5.2 Separating Model Weights (Flash) from Runtime Memory (RAM)
&lt;/h3&gt;

&lt;p&gt;To save RAM, model weights should remain in Flash and be mapped using &lt;strong&gt;XIP (Execute In Place)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Use the &lt;code&gt;TFLITE_SCHEMA_RESERVED_BUFFER&lt;/code&gt; macro to ensure model parameters are not copied into RAM at startup, reserving the full 512 KB of SRAM for dynamic tensors.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key tip:&lt;/strong&gt; In ESP-IDF, enable &lt;code&gt;CONFIG_SPIRAM_USE_MALLOC&lt;/code&gt; and use &lt;code&gt;heap_caps_malloc(size, MALLOC_CAP_SPIRAM)&lt;/code&gt; to precisely control where tensor buffers are allocated.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Performance Tuning: Maximizing ESP32-S3 Vector Compute
&lt;/h2&gt;

&lt;p&gt;At the edge, every millisecond matters. Maximum inference speed comes down to quantization strategy and operator optimization.&lt;/p&gt;

&lt;h3&gt;
  
  
  6.1 Full Integer Quantization
&lt;/h3&gt;

&lt;p&gt;The ESP32-S3 vector instruction set is optimized specifically for INT8 arithmetic. If a model includes floating-point (FP32) operators, TFLM falls back to slower software-based execution.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Post-Training Quantization (PTQ)&lt;/strong&gt; — when exporting, provide a representative dataset to map the weight dynamic range to -128..127.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quantization-Aware Training (QAT)&lt;/strong&gt; — for accuracy-sensitive models, simulate quantization during training.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Benchmark results show that fully quantized models on ESP32-S3 can run over &lt;strong&gt;6× faster&lt;/strong&gt; than floating-point models.&lt;/p&gt;

&lt;h3&gt;
  
  
  6.2 Profiling Tools
&lt;/h3&gt;

&lt;p&gt;Use &lt;code&gt;esp_timer_get_time()&lt;/code&gt; to measure the execution time of &lt;code&gt;interpreter.Invoke()&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Typical inference results on ESP32-S3:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model Type&lt;/th&gt;
&lt;th&gt;Parameters&lt;/th&gt;
&lt;th&gt;Input Size&lt;/th&gt;
&lt;th&gt;Quantization&lt;/th&gt;
&lt;th&gt;Inference (SRAM)&lt;/th&gt;
&lt;th&gt;Inference (PSRAM)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Keyword Spotting (KWS)&lt;/td&gt;
&lt;td&gt;20K&lt;/td&gt;
&lt;td&gt;1s Audio (MFCC)&lt;/td&gt;
&lt;td&gt;INT8&lt;/td&gt;
&lt;td&gt;~12 ms&lt;/td&gt;
&lt;td&gt;~15 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gesture Recognition (IMU)&lt;/td&gt;
&lt;td&gt;5K&lt;/td&gt;
&lt;td&gt;128Hz Accel&lt;/td&gt;
&lt;td&gt;INT8&lt;/td&gt;
&lt;td&gt;~2 ms&lt;/td&gt;
&lt;td&gt;~2.5 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Person Detection (MobileNet)&lt;/td&gt;
&lt;td&gt;250K&lt;/td&gt;
&lt;td&gt;96×96 Grayscale&lt;/td&gt;
&lt;td&gt;INT8&lt;/td&gt;
&lt;td&gt;N/A (Overflow)&lt;/td&gt;
&lt;td&gt;~145 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Digit Classification (MNIST)&lt;/td&gt;
&lt;td&gt;60K&lt;/td&gt;
&lt;td&gt;28×28 Image&lt;/td&gt;
&lt;td&gt;INT8&lt;/td&gt;
&lt;td&gt;~8 ms&lt;/td&gt;
&lt;td&gt;~10 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Data based on 240 MHz CPU frequency with hardware vector acceleration enabled.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Typical Use Cases: TinyML in Real AIoT Deployments
&lt;/h2&gt;

&lt;p&gt;The ESP32-S3 + TFLM combination supports a wide range of edge AI applications, from voice to vision.&lt;/p&gt;

&lt;h3&gt;
  
  
  7.1 Voice Interaction: Offline Keyword Spotting (KWS)
&lt;/h3&gt;

&lt;p&gt;One of the most mature TFLM use cases. Raw audio is captured from a mic, an FFT generates MFCC features, and these feed into a CNN for classification. ESP32-S3's vector instructions accelerate FFT, enabling real-time wake-word detection at very low power.&lt;/p&gt;

&lt;h3&gt;
  
  
  7.2 Edge Vision: Smart Doorbells and Face Detection
&lt;/h3&gt;

&lt;p&gt;With the ESP32-S3 camera interface, TFLM can run lightweight vision models:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Low-power sensing&lt;/strong&gt; — a PIR sensor wakes the chip, which captures an image and uses TFLM to determine if a human is present.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Advantages&lt;/strong&gt; — local pre-filtering reduces Wi-Fi power consumption by ~90% compared to cloud upload, and improves privacy.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  7.3 Industrial Predictive Maintenance: Vibration Analysis
&lt;/h3&gt;

&lt;p&gt;A three-axis accelerometer collects motor vibration data; a TFLM model analyzes frequency-domain features locally and detects early signs of wear, imbalance, or overheating. Devices only send alerts on anomaly, not continuous raw data.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Practical Advice: Three Steps to Optimize TFLM Projects
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Trim unused operators&lt;/strong&gt; — the default &lt;code&gt;AllOpsResolver&lt;/code&gt; includes all supported operators and can consume 100–200 KB of Flash. Use &lt;code&gt;MicroMutableOpResolver&lt;/code&gt; and add only the required operators (e.g., &lt;code&gt;AddConv2D&lt;/code&gt;, &lt;code&gt;AddReshape&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Balance clock speed and power&lt;/strong&gt; — ESP32-S3 supports up to 240 MHz. In battery scenarios, adjust frequency dynamically: higher clock shortens inference time, letting the chip enter Deep Sleep sooner.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Leverage dual-core architecture&lt;/strong&gt; — run Wi-Fi stack and sensor acquisition on Core 0, and TFLM inference independently on Core 1, so network interruptions don't affect inference stability.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Key takeaway:&lt;/strong&gt; High performance on ESP32-S3 requires understanding the memory hierarchy. Careful SRAM management and full integer quantization are essential to pushing MCU-level AI to its limits.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Implementation Example: Integrating TFLM in ESP-IDF
&lt;/h2&gt;

&lt;p&gt;Running inference requires proper &lt;code&gt;MicroInterpreter&lt;/code&gt; configuration and linking the esp-nn library.&lt;/p&gt;

&lt;h3&gt;
  
  
  9.1 Model Loading and Interpreter Initialization
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;.tflite&lt;/code&gt; model is converted into a hex C array stored in Flash. Use a pointer referencing the Flash address directly, rather than copying the model into RAM:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;uint8_t&lt;/span&gt; &lt;span class="n"&gt;tensor_arena&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;kTensorArenaSize&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="n"&gt;__attribute__&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;aligned&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="p"&gt;)));&lt;/span&gt;

&lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="n"&gt;tflite&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;MicroMutableOpResolver&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;resolver&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="n"&gt;resolver&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddConv2D&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="n"&gt;resolver&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AddFullyConnected&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="n"&gt;tflite&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;MicroInterpreter&lt;/span&gt; &lt;span class="nf"&gt;interpreter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;resolver&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tensor_arena&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;kTensorArenaSize&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;error_reporter&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="n"&gt;interpreter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AllocateTensors&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  9.2 Critical Step: Input Preprocessing
&lt;/h3&gt;

&lt;p&gt;Raw data collected by the ESP32-S3 (ADC samples or camera pixels) is typically &lt;code&gt;uint8&lt;/code&gt; or &lt;code&gt;int16&lt;/code&gt;. Before feeding it into a quantized model, ensure the &lt;strong&gt;scale and zero-point match the values used during training&lt;/strong&gt; — getting this wrong can drop accuracy dramatically.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. TFLM vs Other Edge AI Frameworks
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Framework&lt;/th&gt;
&lt;th&gt;Strengths&lt;/th&gt;
&lt;th&gt;Weaknesses&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;TFLM (Native)&lt;/td&gt;
&lt;td&gt;Strong ecosystem, rich operators, native ESP-NN integration&lt;/td&gt;
&lt;td&gt;Steeper learning curve, manual memory management&lt;/td&gt;
&lt;td&gt;General TinyML tasks, research&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Edge Impulse&lt;/td&gt;
&lt;td&gt;User-friendly UI, automated data pipeline, integrated TFLM&lt;/td&gt;
&lt;td&gt;Limited advanced customization, partially closed-source&lt;/td&gt;
&lt;td&gt;Rapid prototyping, non-AI specialists&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ESP-DL&lt;/td&gt;
&lt;td&gt;Official Espressif framework, deeply optimized for S3&lt;/td&gt;
&lt;td&gt;Smaller operator library, more complex conversion&lt;/td&gt;
&lt;td&gt;Vision/speech needing max performance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MicroTVM&lt;/td&gt;
&lt;td&gt;Compile-time optimization, extremely compact code&lt;/td&gt;
&lt;td&gt;Limited operator coverage, complex config&lt;/td&gt;
&lt;td&gt;Ultra resource-constrained MCUs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Recommendation:&lt;/strong&gt; value development efficiency and community support → TFLM. Need to extract every bit of S3 performance with a simple model → ESP-DL.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Deployment Pitfalls: Three Common Mistakes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ignoring memory alignment&lt;/strong&gt; — ESP32-S3 SIMD requires tensor addresses to be 16-byte aligned. A misaligned &lt;code&gt;tensor_arena&lt;/code&gt; can trigger &lt;code&gt;StoreProhibited&lt;/code&gt; exceptions or serious performance loss.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Operator shadowing&lt;/strong&gt; — when integrating esp-nn, check &lt;code&gt;CMakeLists.txt&lt;/code&gt; to ensure the optimized library is linked, not the default reference implementation. If a small convolution takes &amp;gt;50 ms, hardware acceleration is likely not active.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ignoring quantization parameters&lt;/strong&gt; — don't feed raw 0–255 pixel values into an INT8 model. Map input using &lt;code&gt;input-&amp;gt;params.scale&lt;/code&gt; and &lt;code&gt;input-&amp;gt;params.zero_point&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  12. Where ESP32-S3 TinyML Is Heading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Multi-modal fusion&lt;/strong&gt; — dual-core processing: audio wake-word on one core, visual gesture on the other.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On-device learning&lt;/strong&gt; — partial weight update techniques may allow local fine-tuning based on user behavior.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Advanced model compression&lt;/strong&gt; — Neural Architecture Search (NAS) will produce more efficient backbones tailored for ESP32-S3.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  13. System Execution Diagram: ESP32-S3 Memory and TFLM
&lt;/h2&gt;

&lt;p&gt;The relationship between Flash, PSRAM, SRAM, Tensor Arena, and ESP-NN defines the actual memory architecture of an ESP32-S3 running TFLM:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Flash&lt;/strong&gt; → holds model weights (&lt;code&gt;.tflite&lt;/code&gt; as C array), accessed via XIP.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SRAM (512 KB)&lt;/strong&gt; → fast on-chip memory, ideal for Tensor Arena of small models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PSRAM (optional)&lt;/strong&gt; → external expansion for large feature maps, slower but cached.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tensor Arena&lt;/strong&gt; → static block for intermediate tensors, placed in SRAM or PSRAM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ESP-NN&lt;/strong&gt; → optimized kernels replacing reference implementations during compilation.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;This article walked through the complete workflow for deploying TensorFlow Lite Micro on ESP32-S3: how the Xtensa LX7 vector instruction set accelerates deep learning via SIMD, INT8 quantization, and memory hierarchy optimization (SRAM/PSRAM). Benchmark results and code examples show the significant gains from enabling esp-nn.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key takeaways:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best practices&lt;/strong&gt; — ensure memory alignment and trim unused operators.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Core architecture&lt;/strong&gt; — TFLM interpreter with ESP-NN hardware acceleration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Performance impact&lt;/strong&gt; — INT8 quantization delivers up to 6× speedup.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory optimization&lt;/strong&gt; — careful Tensor Arena allocation is critical.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  FAQ: ESP32-S3 and TensorFlow Lite Micro
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q1: Which hardware-accelerated operators are supported?&lt;/strong&gt;&lt;br&gt;
The esp-nn library accelerates depthwise convolution, standard convolution, fully connected layers, pooling, and some activations (ReLU, Leaky ReLU), optimized with the S3's 128-bit vector instructions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q2: How can I tell if my Tensor Arena size is right?&lt;/strong&gt;&lt;br&gt;
After &lt;code&gt;AllocateTensors()&lt;/code&gt;, use &lt;code&gt;interpreter.arena_used_bytes()&lt;/code&gt; to check actual usage. Leave a 10–20% margin for runtime stack overhead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q3: Why does my model work on PC but produce wrong results on the S3?&lt;/strong&gt;&lt;br&gt;
In 90% of cases, it's quantization mismatch — check that the Representative Dataset reflects real sensor data distribution, and that input data is scaled with the correct scale/zero-point.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q4: Does PSRAM significantly reduce inference speed?&lt;/strong&gt;&lt;br&gt;
Yes, typically 10–30% additional latency. Enabling Octal SPI and cache prefetching minimizes the impact. For large models, PSRAM is often the only viable option.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q5: Can ESP32-S3 run floating-point models?&lt;/strong&gt;&lt;br&gt;
Yes, but strongly discouraged. The S3 has a single-precision FPU but no vectorized FP acceleration, so FP32 models run significantly slower than INT8.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;What's the biggest performance bottleneck you've hit with TinyML on ESP32-S3 — memory, quantization, or getting ESP-NN to actually kick in? And have you found PSRAM worth the latency trade-off for vision models?&lt;/em&gt;&lt;/p&gt;

</description>
      <category>esp32</category>
      <category>tinyml</category>
      <category>machinelearning</category>
      <category>iot</category>
    </item>
    <item>
      <title>How to Deploy YOLOv8 on RK3566: Build Efficient Edge AI Inference from Scratch</title>
      <dc:creator>ZedIoT</dc:creator>
      <pubDate>Thu, 24 Sep 2026 12:30:00 +0000</pubDate>
      <link>https://dev.to/zediot/how-to-deploy-yolov8-on-rk3566-build-efficient-edge-ai-inference-from-scratch-4aio</link>
      <guid>https://dev.to/zediot/how-to-deploy-yolov8-on-rk3566-build-efficient-edge-ai-inference-from-scratch-4aio</guid>
      <description>&lt;p&gt;Running a real-time object detector on a $30 ARM board sounds like a stretch — until you realize the RK3566's NPU was built specifically for int8 models. The catch is that YOLOv8 ships in FP32, so the entire job is a conversion-and-quantization pipeline rather than a plug-and-play demo. This article walks that pipeline end to end: hardware, ONNX export, RKNN conversion, post-processing, and production deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Understanding the Hardware: RK3566 at a Glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Specification&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CPU&lt;/td&gt;
&lt;td&gt;Quad-core ARM Cortex-A55 (1.8GHz)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NPU&lt;/td&gt;
&lt;td&gt;0.8–1.0 TOPS (Rockchip 3rd-gen NPU)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU&lt;/td&gt;
&lt;td&gt;Mali-G52 (optional for OpenCL acceleration)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory&lt;/td&gt;
&lt;td&gt;Up to 4GB LPDDR4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OS&lt;/td&gt;
&lt;td&gt;Linux / Android (Debian, Ubuntu, or Buildroot variants)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SDK&lt;/td&gt;
&lt;td&gt;Rockchip RKNN Toolkit 2.x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The NPU inside RK3566 is specifically designed for int8 quantized models, meaning you'll need to convert YOLOv8 (originally in FP32) to RKNN format with proper calibration and optimization.&lt;/p&gt;

&lt;h2&gt;
  
  
  YOLOv8 Overview
&lt;/h2&gt;

&lt;p&gt;YOLOv8, developed by Ultralytics, is the latest iteration of the popular "You Only Look Once" object detection series.&lt;/p&gt;

&lt;p&gt;Compared with YOLOv5 and YOLOv7, YOLOv8 introduces:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Improved architecture with &lt;strong&gt;CSPDarknet + C2f blocks&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dynamic shape support&lt;/strong&gt; for flexible resolutions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ONNX export compatibility&lt;/strong&gt; for cross-platform deployment&lt;/li&gt;
&lt;li&gt;Smaller model sizes (&lt;strong&gt;YOLOv8-n, YOLOv8-s&lt;/strong&gt;) ideal for edge devices&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On RK3566, the &lt;strong&gt;YOLOv8-n (Nano)&lt;/strong&gt; model is recommended for achieving real-time inference while maintaining decent accuracy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deployment Workflow Overview
&lt;/h2&gt;

&lt;p&gt;Before diving into code, the high-level process looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Model preparation&lt;/strong&gt; — train or download YOLOv8 weights.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ONNX export&lt;/strong&gt; — use Ultralytics CLI or API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conversion&lt;/strong&gt; — convert ONNX → RKNN using Rockchip RKNN Toolkit 2.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deployment&lt;/strong&gt; — run inference via RKNN runtime on the device.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Visualization&lt;/strong&gt; — render detection boxes on camera input.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Environment Preparation
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Hardware Requirements
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;RK3566 board (e.g., Radxa Zero 3W, Pine64 Quartz64, or custom industrial SBC)&lt;/li&gt;
&lt;li&gt;5V/3A power supply&lt;/li&gt;
&lt;li&gt;USB serial cable or SSH access&lt;/li&gt;
&lt;li&gt;Camera (USB / MIPI)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Software Environment
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Tool / Version&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Host PC&lt;/td&gt;
&lt;td&gt;Ubuntu 20.04 / 22.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Python&lt;/td&gt;
&lt;td&gt;3.8+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;YOLOv8&lt;/td&gt;
&lt;td&gt;Ultralytics &amp;gt;= 8.0.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ONNX&lt;/td&gt;
&lt;td&gt;1.12+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RKNN Toolkit&lt;/td&gt;
&lt;td&gt;v2.3.0+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RKNN Runtime&lt;/td&gt;
&lt;td&gt;for ARM64 / Debian&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;💡 &lt;strong&gt;Tip:&lt;/strong&gt; You need both the RKNN Toolkit (for model conversion on PC) and RKNN Runtime (for deployment on the device).&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Install RKNN Toolkit on Host PC
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Create virtual environment&lt;/span&gt;
python3 &lt;span class="nt"&gt;-m&lt;/span&gt; venv rknn_env
&lt;span class="nb"&gt;source &lt;/span&gt;rknn_env/bin/activate

&lt;span class="c"&gt;# Install dependencies&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;torch onnx onnxsim ultralytics
pip &lt;span class="nb"&gt;install &lt;/span&gt;rknn-toolkit2&lt;span class="o"&gt;==&lt;/span&gt;2.3.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After installation, verify by running:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; rknn.api.rknn &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should see a valid version output confirming the toolkit is correctly installed.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Export YOLOv8 Model to ONNX
&lt;/h3&gt;

&lt;p&gt;If you've trained your custom YOLOv8 model (or downloaded pre-trained weights), export it with the following command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;yolo &lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;yolov8n.pt &lt;span class="nv"&gt;format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;onnx &lt;span class="nv"&gt;opset&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;12
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This will produce a &lt;code&gt;yolov8n.onnx&lt;/code&gt; file, which can now be optimized for RK3566.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deploy YOLOv8 on RK3566 (Step-by-Step)
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Convert ONNX to RKNN Format
&lt;/h3&gt;

&lt;p&gt;Once you have the YOLOv8 ONNX model (&lt;code&gt;yolov8n.onnx&lt;/code&gt;), the next step is converting it into Rockchip's RKNN format, which is optimized for the NPU.&lt;/p&gt;

&lt;p&gt;Below is a sample Python script using RKNN Toolkit 2:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;rknn.api&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;RKNN&lt;/span&gt;

&lt;span class="n"&gt;rknn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNN&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# 1. Load ONNX model
&lt;/span&gt;&lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_onnx&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;yolov8n.onnx&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 2. Configure preprocessing parameters
&lt;/span&gt;&lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;mean_values&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt;
    &lt;span class="n"&gt;std_values&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt;
    &lt;span class="n"&gt;target_platform&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;rk3566&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;quantized_dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;asymmetric_affine-u8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 3. Build RKNN model
&lt;/span&gt;&lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;build&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;do_quantization&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dataset&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;./dataset.txt&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 4. Export model
&lt;/span&gt;&lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;export_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;yolov8n_rk3566.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Explanation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;target_platform='rk3566'&lt;/code&gt; ensures compatibility with the RK3566 NPU.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;dataset.txt&lt;/code&gt; should contain a list of sample image paths for quantization calibration.&lt;/li&gt;
&lt;li&gt;The quantization step converts FP32 → INT8, significantly improving inference speed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;⚠️ &lt;strong&gt;Tip:&lt;/strong&gt; Choose representative images for quantization to minimize accuracy loss.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Prepare the Dataset for Quantization
&lt;/h3&gt;

&lt;p&gt;To generate &lt;code&gt;dataset.txt&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;find ./images/ &lt;span class="nt"&gt;-type&lt;/span&gt; f &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*.jpg"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; dataset.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use around 100–300 images covering your main object categories and lighting variations. The more representative your dataset, the better the quantization accuracy.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Verify RKNN Model on Host PC
&lt;/h3&gt;

&lt;p&gt;Before deploying to the board, test the converted RKNN model locally to ensure correctness:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;rknn.api&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;RKNN&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;

&lt;span class="n"&gt;rknn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNN&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;yolov8n_rk3566.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;init_runtime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;img&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;imread&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;test.jpg&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;outputs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inputs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;outputs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the model runs without errors and produces detection tensors, you're ready to deploy it onto RK3566.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Deploying to RK3566 Board
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Step 1. Transfer Files&lt;/strong&gt; — copy the following to your RK3566 device via SCP or USB:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;yolov8n_rk3566.rknn&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;test.jpg&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;inference_rk3566.py&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Step 2. Install Runtime&lt;/strong&gt; — on RK3566 (Debian/Ubuntu system):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;apt update
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install &lt;/span&gt;python3-opencv
pip3 &lt;span class="nb"&gt;install &lt;/span&gt;rknn-runtime&lt;span class="o"&gt;==&lt;/span&gt;2.3.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 3. Run Inference:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 inference_rk3566.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example minimal code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;rknnlite.api&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;RKNNLite&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;

&lt;span class="n"&gt;rknn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNNLite&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;yolov8n_rk3566.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;init_runtime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;img&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;imread&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;test.jpg&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;outputs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inputs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="c1"&gt;# Visualize results
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Inference output shape:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;outputs&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;💡 &lt;strong&gt;Pro Tip:&lt;/strong&gt; &lt;code&gt;RKNNLite&lt;/code&gt; is optimized for on-device inference and uses less memory than RKNN Toolkit.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Real-Time Camera Inference
&lt;/h3&gt;

&lt;p&gt;For applications like smart surveillance or factory inspection, connect a USB/MIPI camera to the RK3566 device.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;rknnlite.api&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;RKNNLite&lt;/span&gt;

&lt;span class="n"&gt;rknn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNNLite&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;yolov8n_rk3566.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;init_runtime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;cap&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;VideoCapture&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;ret&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;frame&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cap&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;ret&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="n"&gt;outputs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inputs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="c1"&gt;# TODO: Add YOLOv8 postprocessing (NMS + bbox drawing)
&lt;/span&gt;    &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;imshow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;YOLOv8 RK3566&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;waitKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;27&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;  &lt;span class="c1"&gt;# ESC
&lt;/span&gt;        &lt;span class="k"&gt;break&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The postprocessing stage involves decoding model output tensors and applying Non-Max Suppression (NMS) to draw bounding boxes.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Measuring Performance
&lt;/h3&gt;

&lt;p&gt;You can use Python's &lt;code&gt;time&lt;/code&gt; module to benchmark inference time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="n"&gt;t0&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;outputs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inputs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Inference time:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;t0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ms&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Typical results for YOLOv8n (INT8) on RK3566:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Resolution&lt;/th&gt;
&lt;th&gt;FPS&lt;/th&gt;
&lt;th&gt;CPU Usage&lt;/th&gt;
&lt;th&gt;Power Consumption&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;YOLOv8n (INT8)&lt;/td&gt;
&lt;td&gt;320×320&lt;/td&gt;
&lt;td&gt;~18–22 FPS&lt;/td&gt;
&lt;td&gt;&amp;lt;50%&lt;/td&gt;
&lt;td&gt;~3.5W&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;YOLOv8s (INT8)&lt;/td&gt;
&lt;td&gt;640×640&lt;/td&gt;
&lt;td&gt;~8–10 FPS&lt;/td&gt;
&lt;td&gt;&amp;lt;70%&lt;/td&gt;
&lt;td&gt;~4.2W&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;👉 You can achieve real-time detection for small to medium models on RK3566, ideal for IoT cameras, kiosks, and embedded vision systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  Optimization Tips
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Use Fixed Input Resolution&lt;/strong&gt; (e.g. 320×320) — avoid dynamic resizing on-device to save CPU cycles.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quantization-Aware Training (QAT)&lt;/strong&gt; — if possible, retrain YOLOv8 with quantization awareness to preserve accuracy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batch Normalization Folding&lt;/strong&gt; — enable folding during conversion for better NPU compatibility.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use RKNN Precompiled Postprocessing&lt;/strong&gt; — Rockchip SDK provides C++ utilities for YOLO postprocessing (NMS) with NEON optimization.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Building Full Inference Pipeline on RK3566
&lt;/h2&gt;

&lt;p&gt;Now that YOLOv8 is successfully running on RK3566, the system needs post-processing, real-time visualization, and deployment automation to be production-ready.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Understanding YOLOv8 Output Structure
&lt;/h3&gt;

&lt;p&gt;After running inference, the RKNN model outputs one or more tensors, typically shaped like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(1, 84, 8400)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Where:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;8400&lt;/strong&gt; = total number of anchor points (depending on input size)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;84&lt;/strong&gt; = (4 bbox coordinates + 80 class probabilities)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The next step is to decode these raw tensors into bounding boxes, class labels, and confidence scores.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Implementing YOLOv8 Post-Processing
&lt;/h3&gt;

&lt;p&gt;Here's a simplified Python example for YOLOv8 output decoding and Non-Max Suppression (NMS) on RK3566:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;xywh2xyxy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;zeros_like&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;nms&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;iou_thresh&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.45&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;idxs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;argsort&lt;/span&gt;&lt;span class="p"&gt;()[::&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;keep&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;idxs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;idxs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="n"&gt;keep&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;idxs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;break&lt;/span&gt;
        &lt;span class="n"&gt;iou&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;bbox_iou&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;idxs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:]])&lt;/span&gt;
        &lt;span class="n"&gt;idxs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;idxs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:][&lt;/span&gt;&lt;span class="n"&gt;iou&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;iou_thresh&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;keep&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;bbox_iou&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;box1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;inter_x1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;maximum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;box1&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;inter_y1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;maximum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;box1&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;inter_x2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;minimum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;box1&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;inter_y2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;minimum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;box1&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;inter_area&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;maximum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inter_x2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;inter_x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;maximum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inter_y2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;inter_y1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;area1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;box1&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;box1&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;box1&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;box1&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;area2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;inter_area&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;area1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;area2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;inter_area&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This lightweight NMS function works efficiently on RK3566 for up to 8–10 objects per frame.&lt;/p&gt;

&lt;p&gt;⚙️ If you need higher performance, Rockchip provides a C++ YOLO post-processing SDK (&lt;code&gt;rknn_yolov5_postprocess.cc&lt;/code&gt;) that can be easily adapted to YOLOv8.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Visualizing Real-Time Detection
&lt;/h3&gt;

&lt;p&gt;Integrate NMS with OpenCV to visualize detections from your live camera feed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;rknnlite.api&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;RKNNLite&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;draw_boxes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cls_ids&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;class_names&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;box&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cls&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cls_ids&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;box&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;label&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;class_names&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;cls&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;rectangle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;putText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;label&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;FONT_HERSHEY_SIMPLEX&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.6&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;img&lt;/span&gt;

&lt;span class="n"&gt;rknn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNNLite&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;yolov8n_rk3566.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;init_runtime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;cap&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;VideoCapture&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;class_names&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;coco.names&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;ret&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;frame&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cap&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;ret&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="n"&gt;outputs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inputs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cls_ids&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;postprocess_yolov8&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;outputs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;frame&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;draw_boxes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cls_ids&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;class_names&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;imshow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;YOLOv8 Edge AI - RK3566&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;waitKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;27&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives you a real-time camera detection demo running directly on RK3566's NPU — typically reaching 15–20 FPS with YOLOv8n.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Deploying as a Service (Production Mode)
&lt;/h3&gt;

&lt;p&gt;Once the system runs smoothly, you can automate it using systemd so that it launches automatically after power-on — suitable for kiosks, industrial cameras, or unattended IoT devices.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1. Create a Service File:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;nano /etc/systemd/system/yolov8_inference.service
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[Unit]&lt;/span&gt;
&lt;span class="py"&gt;Description&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;YOLOv8 Edge Inference on RK3566&lt;/span&gt;
&lt;span class="py"&gt;After&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;network.target&lt;/span&gt;

&lt;span class="nn"&gt;[Service]&lt;/span&gt;
&lt;span class="py"&gt;ExecStart&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;/usr/bin/python3 /home/pi/yolo/inference_rk3566.py&lt;/span&gt;
&lt;span class="py"&gt;Restart&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;always&lt;/span&gt;
&lt;span class="py"&gt;User&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;pi&lt;/span&gt;
&lt;span class="py"&gt;WorkingDirectory&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;/home/pi/yolo/&lt;/span&gt;

&lt;span class="nn"&gt;[Install]&lt;/span&gt;
&lt;span class="py"&gt;WantedBy&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;multi-user.target&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 2. Enable &amp;amp; Start Service:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl &lt;span class="nb"&gt;enable &lt;/span&gt;yolov8_inference.service
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl start yolov8_inference.service
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The service now runs automatically on boot — ensuring headless operation for real-world edge deployments.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Optional: Dockerized Deployment
&lt;/h3&gt;

&lt;p&gt;For developers building scalable solutions (e.g., batch camera inference), containerization can simplify setup and updates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dockerfile Example:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; arm64v8/python:3.9&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;apt update &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; apt &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; python3-opencv
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; requirements.txt /tmp/&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; /tmp/requirements.txt
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; yolov8n_rk3566.rknn /app/&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; inference_rk3566.py /app/&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;
&lt;span class="k"&gt;CMD&lt;/span&gt;&lt;span class="s"&gt; ["python3", "inference_rk3566.py"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Build &amp;amp; Run:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker build &lt;span class="nt"&gt;-t&lt;/span&gt; yolov8-rk3566 &lt;span class="nb"&gt;.&lt;/span&gt;
docker run &lt;span class="nt"&gt;--privileged&lt;/span&gt; &lt;span class="nt"&gt;--device&lt;/span&gt; /dev/video0:/dev/video0 yolov8-rk3566
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This encapsulates all dependencies and ensures consistent runtime behavior across devices.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Integration Possibilities
&lt;/h3&gt;

&lt;p&gt;Once YOLOv8 inference is stable on RK3566, you can expand functionality:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Integration&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MQTT / WebSocket&lt;/td&gt;
&lt;td&gt;Stream detection results to cloud dashboard or local server.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RTSP Video Stream&lt;/td&gt;
&lt;td&gt;Use GStreamer or ffmpeg to output processed video.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Edge–Cloud Hybrid&lt;/td&gt;
&lt;td&gt;Combine RK3566 inference with cloud analytics via REST API.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Custom Models&lt;/td&gt;
&lt;td&gt;Replace YOLOv8n with your own trained models (e.g., defect detection).&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This flexibility makes RK3566 suitable for smart retail, factory inspection, traffic monitoring, and AIoT gateways.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Verification Checklist
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;✅ Model converted successfully (&lt;code&gt;.rknn&lt;/code&gt; file valid)&lt;/li&gt;
&lt;li&gt;✅ Quantization accuracy acceptable (mAP loss &amp;lt;3%)&lt;/li&gt;
&lt;li&gt;✅ Real-time performance achieved (&amp;gt;15 FPS)&lt;/li&gt;
&lt;li&gt;✅ Auto-start service working correctly&lt;/li&gt;
&lt;li&gt;✅ Integration tests (MQTT, camera, Docker) passed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When all these boxes are checked, your RK3566-powered device is ready for production-grade YOLOv8 edge inference.&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;This tutorial walked through a complete end-to-end YOLOv8 Edge AI pipeline on RK3566 — from ONNX export to production deployment.&lt;/p&gt;

&lt;p&gt;Key takeaways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;RKNN Toolkit 2&lt;/strong&gt; simplifies ONNX → RKNN conversion and quantization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RKNN Runtime&lt;/strong&gt; enables fast, low-power inference on embedded hardware.&lt;/li&gt;
&lt;li&gt;With proper post-processing and service automation, &lt;strong&gt;RK3566 becomes a reliable edge vision platform&lt;/strong&gt; for commercial applications.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Have you hit a wall with RKNN quantization accuracy, or with real-time NMS post-processing on a low-power board? What input resolution did you settle on to balance FPS and detection quality on RK3566?&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>machinelearning</category>
      <category>ai</category>
      <category>deeplearning</category>
    </item>
    <item>
      <title>From Wake Word Detection to Edge Intelligence: The Technical Potential of ESP32-S3 TensorFlow Lite Micro</title>
      <dc:creator>ZedIoT</dc:creator>
      <pubDate>Tue, 22 Sep 2026 12:30:00 +0000</pubDate>
      <link>https://dev.to/zediot/from-wake-word-detection-to-edge-intelligence-the-technical-potential-of-esp32-s3-tensorflow-lite-30hd</link>
      <guid>https://dev.to/zediot/from-wake-word-detection-to-edge-intelligence-the-technical-potential-of-esp32-s3-tensorflow-lite-30hd</guid>
      <description>&lt;p&gt;When a smart speaker, vacuum robot, or wearable adds "wake word" support, it almost always ships with a dedicated voice chip — an ASR5505, a BD3751, or an XMOS XVF. Those chips are genuinely good at one thing: listening for a fixed keyword at ultra-low power. The moment you need a custom wake word, a second language, or a model you can update after the product leaves the factory, that convenience turns into a wall.&lt;/p&gt;

&lt;p&gt;That's the gap ESP32-S3 + TensorFlow Lite Micro (TFLM) fills. Instead of a chip that hard-codes "listen for these three words," you get a general-purpose MCU running a model you can retrain, requantize, and push out over the air. This article walks through what that actually looks like end to end — the hardware, the audio pipeline, the model, and the deployment loop — not just a "hello world" wake word demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Why Run AI on MCUs? ESP32-S3 as an Edge AI Platform
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1.1 The stagnation of traditional wake word systems
&lt;/h3&gt;

&lt;p&gt;Wake word detection is now a baseline feature, not a premium one. But the dedicated voice chips behind most of it share the same three limits:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Wake words are fixed&lt;/strong&gt; and can't be changed dynamically.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Firmware and model updates depend on the chip vendor.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No flexibility&lt;/strong&gt; for personalized or multi-language wake words.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As voice interaction becomes table stakes, that hardware-level rigidity is the bottleneck.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.2 The rise of MCU + AI frameworks
&lt;/h3&gt;

&lt;p&gt;The alternative is running a lightweight model directly on a general-purpose MCU. ESP32-S3 sits in the sweet spot of compute, power, and cost for workloads where cloud inference is impractical.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ESP32-S3&lt;/strong&gt; has dual-mode Wi-Fi + BLE and built-in &lt;strong&gt;AI vector instructions&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TensorFlow Lite Micro&lt;/strong&gt; is a minimal inference framework for resource-constrained devices.&lt;/li&gt;
&lt;li&gt;Together they run on-device AI — &lt;strong&gt;wake word detection, gesture recognition, sound classification&lt;/strong&gt; — inside a few hundred KB of memory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This means AI no longer depends on the cloud. Devices can sense, analyze, and respond locally, even offline or in low-power environments.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.3 Purpose of this article
&lt;/h3&gt;

&lt;p&gt;The point here is broader than a wake word demo. Three things worth understanding:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;TensorFlow Lite Micro is far more than a wake word tool.&lt;/li&gt;
&lt;li&gt;ESP32-S3 extends AI computation down to the MCU level.&lt;/li&gt;
&lt;li&gt;Deploying models on general-purpose MCUs is becoming the new mainstream for low-power intelligent devices.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  2. Technical Principles: From Wake Word Detection to Edge Perception
&lt;/h2&gt;

&lt;h3&gt;
  
  
  2.1 ESP32-S3 Hardware Overview
&lt;/h3&gt;

&lt;p&gt;The ESP32-S3 is Espressif's current-generation IoT MCU, with meaningful upgrades in compute, AI acceleration, and peripheral expansion over earlier ESP32 parts.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Module&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CPU&lt;/td&gt;
&lt;td&gt;Xtensa LX7 dual-core, up to 240 MHz&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI / DSP Acceleration&lt;/td&gt;
&lt;td&gt;SIMD vector instruction set for convolution and matrix ops&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory&lt;/td&gt;
&lt;td&gt;512 KB SRAM, expandable with external PSRAM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wireless&lt;/td&gt;
&lt;td&gt;Wi-Fi 2.4 GHz + BLE 5.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interfaces&lt;/td&gt;
&lt;td&gt;I2S, SPI, UART, ADC, PWM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical Use Cases&lt;/td&gt;
&lt;td&gt;Offline voice recognition, motion detection, sound analysis, vibration monitoring&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The vector instruction set accelerates CNN and LSTM-style operations, which removes the need for a separate AI co-processor. A single ESP32-S3 can "hear," "detect," and "understand" its environment.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.2 What is TensorFlow Lite Micro (TFLM)?
&lt;/h3&gt;

&lt;p&gt;TFLM is Google's lightweight inference framework for MCUs, DSPs, and other embedded targets. Its core idea: a microcontroller can run a deep learning model even without an OS or dynamic memory allocation.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Small footprint&lt;/td&gt;
&lt;td&gt;Runtime library &amp;lt; 100 KB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No dependencies&lt;/td&gt;
&lt;td&gt;Works without RTOS, malloc, or filesystem&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Highly portable&lt;/td&gt;
&lt;td&gt;Supports ARM, RISC-V, and Xtensa&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quantized models&lt;/td&gt;
&lt;td&gt;Runs int8/uint8 networks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Custom operators&lt;/td&gt;
&lt;td&gt;User-defined ops and lightweight optimizations&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That minimalist design is what makes TFLM a fit for ESP32-S3 — AI capability without sacrificing latency or power.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.3 System workflow
&lt;/h3&gt;

&lt;p&gt;Running TFLM on ESP32-S3 for wake word or sound classification follows this flow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Audio capture&lt;/strong&gt; via I2S from a MEMS microphone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feature extraction&lt;/strong&gt; (MFCC) on the MCU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inference&lt;/strong&gt; through the quantized model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Classification output&lt;/strong&gt; that triggers the wake/action.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This lets you build custom auditory models without vendor-locked algorithms. Example applications:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Custom wake words for smart home devices.&lt;/li&gt;
&lt;li&gt;Mechanical noise classification in industrial equipment.&lt;/li&gt;
&lt;li&gt;Environmental sound analysis in wearables.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2.4 Why this architecture is sustainable
&lt;/h3&gt;

&lt;p&gt;Dedicated voice chips are static; MCU + TFLM systems are evolutionary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Models can be retrained and updated anytime.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Different environments can use different models.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloud training + on-device inference&lt;/strong&gt; form a continuous feedback loop.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Devices stay adaptable long after deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Implementation Path: Building Local Wake Word Detection on ESP32-S3
&lt;/h2&gt;

&lt;p&gt;A complete on-device wake word system has five stages:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Audio capture and preprocessing&lt;/li&gt;
&lt;li&gt;Feature extraction (MFCC)&lt;/li&gt;
&lt;li&gt;Model design and quantization&lt;/li&gt;
&lt;li&gt;Model deployment and inference&lt;/li&gt;
&lt;li&gt;Performance and power evaluation&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3.1 Audio Input and Front-End Processing
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;(1) Hardware Interface&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;ESP32-S3 natively supports the I2S digital audio interface, compatible with common MEMS mics like &lt;strong&gt;INMP441, SPH0645, and MSM261S4030&lt;/strong&gt;. Digital connection avoids analog noise, which matters in small devices.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Parameter&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sampling rate&lt;/td&gt;
&lt;td&gt;16 kHz&lt;/td&gt;
&lt;td&gt;Covers human voice band&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bit depth&lt;/td&gt;
&lt;td&gt;16-bit&lt;/td&gt;
&lt;td&gt;Balances accuracy and bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Channel&lt;/td&gt;
&lt;td&gt;Mono&lt;/td&gt;
&lt;td&gt;Stereo unnecessary for speech&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frame length&lt;/td&gt;
&lt;td&gt;40 ms (640 samples)&lt;/td&gt;
&lt;td&gt;Matches MFCC window&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;ESP-IDF provides a full I2S driver with DMA-based buffering:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="n"&gt;i2s_config_t&lt;/span&gt; &lt;span class="n"&gt;i2s_config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mode&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;I2S_MODE_MASTER&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;I2S_MODE_RX&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sample_rate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;16000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bits_per_sample&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;I2S_BITS_PER_SAMPLE_16BIT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;channel_format&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;I2S_CHANNEL_FMT_ONLY_LEFT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;communication_format&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;I2S_COMM_FORMAT_I2S&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dma_buf_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dma_buf_len&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;(2) Signal Preprocessing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Before feeding data into the model, apply standard conditioning:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;High-pass filtering&lt;/strong&gt; — removes DC bias&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pre-emphasis&lt;/strong&gt; — enhances high-frequency components&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Framing + Hamming window&lt;/strong&gt; — maintains temporal continuity&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;VAD (Voice Activity Detection)&lt;/strong&gt; — reduces inference frequency during silence&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The ESP-DSP library exposes &lt;code&gt;esp_dsp_preemphasis_f32()&lt;/code&gt; and &lt;code&gt;esp_dsp_hamming_window_f32()&lt;/code&gt; to handle these on the MCU.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.2 Feature Extraction: MFCC
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;(1) Why MFCC&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;MFCC (Mel-Frequency Cepstral Coefficients) is the most widely used feature in speech recognition. It transforms waveforms into perceptually meaningful frequency features, reducing input dimensionality while preserving accuracy in low-power environments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;(2) MFCC Calculation Flow&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FFT&lt;/strong&gt; — compute spectral energy of each frame.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mel filter banks&lt;/strong&gt; — map the spectrum to the Mel scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log transform&lt;/strong&gt; — simulate nonlinear human hearing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DCT&lt;/strong&gt; — extract low-dimensional cepstral coefficients (typically 10–13).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;ESP32-S3's DSP instructions accelerate FFT and DCT, hitting ~2–3 ms per frame at 16 kHz.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.3 Model Design and Quantization
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;(1) Model Architecture&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Typical TFLM speech models use compact CNNs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;th&gt;Example Output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Conv2D + ReLU&lt;/td&gt;
&lt;td&gt;Extract time–frequency features&lt;/td&gt;
&lt;td&gt;20×10×16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DepthwiseConv2D&lt;/td&gt;
&lt;td&gt;Reduce dimensionality, local features&lt;/td&gt;
&lt;td&gt;10×5×32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flatten&lt;/td&gt;
&lt;td&gt;Flatten tensor to vector&lt;/td&gt;
&lt;td&gt;1600&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dense + Softmax&lt;/td&gt;
&lt;td&gt;Output classification probabilities&lt;/td&gt;
&lt;td&gt;2 (yes/no)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These models hit high accuracy at a 100–300 KB footprint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;(2) Model Training&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Use the official TensorFlow &lt;strong&gt;Speech Commands&lt;/strong&gt; dataset to train custom wake words like "Hey Lamp" or "Hello Board."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;(3) Model Quantization&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Convert float32 → int8 to fit MCU resources:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;converter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lite&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;TFLiteConverter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_saved_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model_path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;converter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;optimizations&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;tf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;lite&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Optimize&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DEFAULT&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;converter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;target_spec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;supported_types&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;tf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;int8&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;tflite_quant_model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;converter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;convert&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Quantization typically reduces size by &lt;strong&gt;4× with under 2% accuracy loss&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.4 Model Deployment and Inference
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;(1) Embedding the Model&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;TFLM loads models as C arrays:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;xxd &lt;span class="nt"&gt;-i&lt;/span&gt; model.tflite &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; model_data.cc
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;unsigned&lt;/span&gt; &lt;span class="kt"&gt;char&lt;/span&gt; &lt;span class="n"&gt;model_data&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="mh"&gt;0x20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mh"&gt;0x00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mh"&gt;0x00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...};&lt;/span&gt;
&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;model_data_len&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;123456&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;(2) Inference Loop Example&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="cp"&gt;#include&lt;/span&gt; &lt;span class="cpf"&gt;"tensorflow/lite/micro/all_ops_resolver.h"&lt;/span&gt;&lt;span class="cp"&gt;
#include&lt;/span&gt; &lt;span class="cpf"&gt;"tensorflow/lite/micro/micro_interpreter.h"&lt;/span&gt;&lt;span class="cp"&gt;
#include&lt;/span&gt; &lt;span class="cpf"&gt;"model_data.h"&lt;/span&gt;&lt;span class="cp"&gt;
&lt;/span&gt;
&lt;span class="cp"&gt;#define TENSOR_ARENA_SIZE (80 * 1024)
&lt;/span&gt;&lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="kt"&gt;uint8_t&lt;/span&gt; &lt;span class="n"&gt;tensor_arena&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;TENSOR_ARENA_SIZE&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;

&lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;app_main&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;void&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="n"&gt;tflite&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;Model&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tflite&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;GetModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model_data&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="n"&gt;tflite&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;AllOpsResolver&lt;/span&gt; &lt;span class="n"&gt;resolver&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;static&lt;/span&gt; &lt;span class="n"&gt;tflite&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;MicroInterpreter&lt;/span&gt; &lt;span class="n"&gt;interpreter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;resolver&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;tensor_arena&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;TENSOR_ARENA_SIZE&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;interpreter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AllocateTensors&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

    &lt;span class="n"&gt;TfLiteTensor&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;input&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;interpreter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;input&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;GetAudioFeature&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;input&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;int8&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="n"&gt;interpreter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Invoke&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
        &lt;span class="n"&gt;TfLiteTensor&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;interpreter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;uint8&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;printf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Wake word detected!&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This loop reaches real-time inference at &lt;strong&gt;15–20 FPS&lt;/strong&gt; on a 240 MHz ESP32-S3 core.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.5 Performance Metrics and Power Consumption
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Inference latency&lt;/td&gt;
&lt;td&gt;50–60 ms&lt;/td&gt;
&lt;td&gt;Per-frame recognition time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model size&lt;/td&gt;
&lt;td&gt;~240 KB&lt;/td&gt;
&lt;td&gt;After int8 quantization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory usage&lt;/td&gt;
&lt;td&gt;~350 KB&lt;/td&gt;
&lt;td&gt;Including tensors and buffers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU load&lt;/td&gt;
&lt;td&gt;50–60%&lt;/td&gt;
&lt;td&gt;Single-core utilization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Power&lt;/td&gt;
&lt;td&gt;120 mA active / &amp;lt;10 mA standby&lt;/td&gt;
&lt;td&gt;Battery-friendly&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;With low-power listening (periodic sampling + event wake-up), average draw can drop to &lt;strong&gt;30–40 mA&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.6 Optimization Tips
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Use fixed input dimensions&lt;/strong&gt; to prevent memory fragmentation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Apply DMA buffering&lt;/strong&gt; for efficient audio input.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Simplify post-processing&lt;/strong&gt; to only output the top confidence label.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run dual-core parallelism&lt;/strong&gt; — one core for inference, the other for sampling and comms.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. Beyond Wake Word: ESP32-S3 Edge AI Use Cases
&lt;/h2&gt;

&lt;p&gt;Wake word detection is only the entry point. The same hardware can run multiple kinds of perception just by swapping the model — ESP32-S3 + TFLM is a programmable edge-intelligence framework, not a single-purpose voice solution.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.1 Environmental Sound Recognition
&lt;/h3&gt;

&lt;p&gt;In smart home and security, sound recognition extends a system's "hearing":&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Detecting &lt;strong&gt;glass breaking, doorbells, or smoke alarms&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Identifying &lt;strong&gt;pet activity or abnormal noises&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Triggering &lt;strong&gt;local alarms&lt;/strong&gt; from acoustic events.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These models take a one-second MFCC sequence and output classifications like &lt;code&gt;["dog_bark", "alarm", "speech", "background"]&lt;/code&gt;, running at ~8–12 FPS.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.2 Equipment Status and Vibration Detection
&lt;/h3&gt;

&lt;p&gt;Industrial gear often can't stay connected continuously, but its sound and vibration carry diagnostic signal. TFLM models let ESP32-S3 detect a &lt;strong&gt;worn motor, imbalanced fan, or dry-running pump&lt;/strong&gt; on-device.&lt;/p&gt;

&lt;p&gt;Advantages:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;High real-time performance&lt;/strong&gt; — no cloud upload needed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Low power&lt;/strong&gt; — continuous listening under 200 mW.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Strong security&lt;/strong&gt; — only anomaly results are reported, avoiding data leaks and wasted bandwidth.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4.3 Gesture and Motion Recognition
&lt;/h3&gt;

&lt;p&gt;Swap the mic for an IMU (accelerometer + gyroscope) and TFLM runs lightweight motion models for wearables:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gesture operations&lt;/strong&gt; (wrist-raise to wake, hand-wave control)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Posture recognition&lt;/strong&gt; (walking, running, falling)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User behavior modeling&lt;/strong&gt; (usage frequency, movement rhythm)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The dual-core design lets one core handle sensor data while the other runs inference.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.4 Environmental Semantics and Multimodal Fusion
&lt;/h3&gt;

&lt;p&gt;TFLM also supports lightweight multimodal fusion — combining mic, light, temperature, humidity, and IR inputs to infer states like "occupied," "noisy," or "secure." In smart home or commercial settings this enables automatic volume adjustment, occupancy detection, and intrusion alerts.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Hybrid Edge and Cloud Architecture for ESP32-S3 Devices
&lt;/h2&gt;

&lt;p&gt;ESP32-S3 is built for on-device inference, but cloud connectivity can be added selectively for model updates, analytics, and fleet management. TFLM's real strength is closing the loop between cloud training and device inference.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.1 Roles of Local and Cloud Components
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Device (ESP32-S3)&lt;/th&gt;
&lt;th&gt;Cloud (TensorFlow / Server)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Data collection&lt;/td&gt;
&lt;td&gt;Audio and sensor sampling&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Feature extraction&lt;/td&gt;
&lt;td&gt;MFCC / FFT&lt;/td&gt;
&lt;td&gt;Data cleaning and augmentation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model training&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Full TensorFlow training&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model deployment&lt;/td&gt;
&lt;td&gt;OTA update of .tflite files&lt;/td&gt;
&lt;td&gt;Model management and distribution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference&lt;/td&gt;
&lt;td&gt;Real-time TFLM inference&lt;/td&gt;
&lt;td&gt;Event analysis and statistics&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  5.2 OTA Model Update Mechanism
&lt;/h3&gt;

&lt;p&gt;ESP32-S3 supports OTA updates, letting you deliver model files as independent firmware partitions. When noise profiles, accents, or environments change, retrain and redeploy a new model via the cloud — enabling continuous on-device intelligence evolution.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Production Applications of ESP32-S3 Edge AI
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Use Case&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Smart Home&lt;/td&gt;
&lt;td&gt;Offline voice control, ambient sound detection, local security alerts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wearables&lt;/td&gt;
&lt;td&gt;Gesture recognition, fall detection, voice command input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Industrial Monitoring&lt;/td&gt;
&lt;td&gt;Motor vibration analysis, anomaly sound detection, predictive maintenance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retail Terminals&lt;/td&gt;
&lt;td&gt;Voice-controlled ads, customer interaction systems&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agriculture &amp;amp; Security&lt;/td&gt;
&lt;td&gt;Animal activity monitoring, noise tracking, acoustic alerts&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These share three traits: &lt;strong&gt;real-time response&lt;/strong&gt; (no cloud delay), &lt;strong&gt;low power&lt;/strong&gt; (always-on sensing), and &lt;strong&gt;data privacy&lt;/strong&gt; (only events, not raw audio, leave the device).&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Comparative Insights
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Dedicated Voice Chip&lt;/th&gt;
&lt;th&gt;ESP32-S3 + TFLM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Function Scope&lt;/td&gt;
&lt;td&gt;Fixed wake words / commands&lt;/td&gt;
&lt;td&gt;Customizable AI models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flexibility&lt;/td&gt;
&lt;td&gt;Firmware locked&lt;/td&gt;
&lt;td&gt;Retrainable, replaceable models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Algorithm Openness&lt;/td&gt;
&lt;td&gt;Proprietary SDK&lt;/td&gt;
&lt;td&gt;Open-source&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OTA Capability&lt;/td&gt;
&lt;td&gt;Usually unsupported&lt;/td&gt;
&lt;td&gt;Full model hot-swapping&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Application Range&lt;/td&gt;
&lt;td&gt;Voice control in appliances&lt;/td&gt;
&lt;td&gt;Cross-industry edge AI perception&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is the real paradigm shift: instead of buying chips that define functions, developers define capabilities through models. The same MCU can listen, detect, and adapt — through software-defined intelligence.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Summary and Takeaways
&lt;/h2&gt;

&lt;p&gt;ESP32-S3 + TFLM extends well beyond wake word detection. It pushes AI down to the MCU, turning low-power devices into adaptive, intelligent systems.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Engineering&lt;/strong&gt;: efficient on-device inference under tight compute and memory, quantized and deployable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Product&lt;/strong&gt;: updatable models and OTA learning loops extend product life cycles.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Industry&lt;/strong&gt;: edge AI scales down from high-end SoCs to MCUs, bringing affordable intelligence to homes, wearables, and factories.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Wake word detection is just the beginning. As every MCU learns to listen, perceive, and reason locally, edge intelligence becomes a native capability — not an optional feature.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Are you running wake word detection on a dedicated voice chip today, or have you moved it onto a general-purpose MCU like the ESP32-S3? Where did the dedicated-chip approach start to break for your product?&lt;/em&gt;&lt;/p&gt;

</description>
      <category>esp32</category>
      <category>tinyml</category>
      <category>iot</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Best ESP32 Firmware Frameworks in 2026: ESP-IDF, Arduino, ESPHome, or Zephyr?</title>
      <dc:creator>ZedIoT</dc:creator>
      <pubDate>Thu, 17 Sep 2026 12:15:00 +0000</pubDate>
      <link>https://dev.to/zediot/best-esp32-firmware-frameworks-in-2026-esp-idf-arduino-esphome-or-zephyr-52ef</link>
      <guid>https://dev.to/zediot/best-esp32-firmware-frameworks-in-2026-esp-idf-arduino-esphome-or-zephyr-52ef</guid>
      <description>&lt;h1&gt;
  
  
  Best ESP32 Firmware Frameworks in 2026: ESP-IDF, Arduino, ESPHome, or Zephyr?
&lt;/h1&gt;

&lt;p&gt;When people search for the "best ESP32 firmware framework," they're usually asking three different questions at once. First: &lt;strong&gt;what gets a board online fastest?&lt;/strong&gt; Second: &lt;strong&gt;what supports a long-lived production firmware architecture?&lt;/strong&gt; Third: &lt;strong&gt;what is really built for a narrower ecosystem or device-delivery path?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If those questions aren't separated, ESP-IDF, Arduino, ESPHome, and Zephyr get treated as four flat alternatives — even though they sit at &lt;strong&gt;different abstraction levels&lt;/strong&gt; and serve &lt;strong&gt;different project types&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Here's the core conclusion up front: in 2026, if you're building an ESP32 product that needs &lt;strong&gt;long-term maintainability, close access to chip capabilities, faster support for new SoCs, and clear control over OTA, logging, drivers, and runtime boundaries&lt;/strong&gt;, ESP-IDF should still be your default starting point. &lt;strong&gt;Arduino&lt;/strong&gt; is still excellent for fast validation, simpler devices, and library reuse — but it shouldn't remain the unquestioned architectural center of a complex product. &lt;strong&gt;ESPHome&lt;/strong&gt; is a strong choice for Home Assistant-oriented nodes and smart-home endpoints, not for most general custom firmware. And &lt;strong&gt;Zephyr&lt;/strong&gt; is worth the extra complexity only when cross-vendor RTOS unification is a first-class requirement.&lt;/p&gt;

&lt;p&gt;In other words, 2026 isn't about one framework replacing another. It's about different ESP32 workflows converging toward the &lt;strong&gt;same ESP-IDF base&lt;/strong&gt; from different distances above it.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The easiest mistake is not capability mismatch — it's abstraction mismatch
&lt;/h2&gt;

&lt;h3&gt;
  
  
  These options don't all solve the same problem
&lt;/h3&gt;

&lt;p&gt;Teams say they're "comparing ESP32 frameworks," but the objects on the table aren't actually peers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ESP-IDF&lt;/strong&gt; is Espressif's official IoT development framework and the most direct path to chip capabilities.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Arduino ESP32&lt;/strong&gt; is a higher-level programming path built around the Arduino model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ESPHome&lt;/strong&gt; is a declarative device system aimed at Home Assistant and smart-home endpoints.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zephyr&lt;/strong&gt; is a cross-vendor RTOS platform rather than an ESP32-specific product stack.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PlatformIO&lt;/strong&gt; is mainly a build and project environment, not a framework itself.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's why broad popularity comparisons usually fail. The better question is whether your project behaves more like a &lt;strong&gt;production firmware system&lt;/strong&gt;, a &lt;strong&gt;fast prototype&lt;/strong&gt;, a &lt;strong&gt;smart-home appliance node&lt;/strong&gt;, or &lt;strong&gt;one board inside a larger unified RTOS platform&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  By 2026, more ESP32 paths are converging toward the same base layer
&lt;/h3&gt;

&lt;p&gt;One of the biggest shifts by 2026 is that the ecosystem is less fragmented than it first appears:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ESP-IDF&lt;/strong&gt; remains the first landing zone for new Espressif chips, official peripherals, and official documentation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Arduino ESP32 3.x&lt;/strong&gt; is now clearly built around newer ESP-IDF foundations rather than living in a separate universe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ESPHome&lt;/strong&gt; has pushed further toward ESP-IDF as the default ESP32 framework direction, especially in its 2026 generation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That changes the selection model. For many teams, the real question is no longer &lt;em&gt;which isolated ecosystem to join&lt;/em&gt;, but &lt;em&gt;how far above ESP-IDF you want to work&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. What each path is actually good at
&lt;/h2&gt;

&lt;h3&gt;
  
  
  ESP-IDF: the default for product-grade firmware
&lt;/h3&gt;

&lt;p&gt;If your project looks like any of these, ESP-IDF should usually be your first serious default:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The device is expected to &lt;strong&gt;ship and stay maintainable&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;You need &lt;strong&gt;explicit control&lt;/strong&gt; over peripherals, partitions, logs, OTA, power, and failure boundaries&lt;/li&gt;
&lt;li&gt;You want &lt;strong&gt;earlier access to new chips and capabilities&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;You don't want the system boxed in by a convenience abstraction&lt;/li&gt;
&lt;li&gt;The team is building an &lt;strong&gt;embedded product&lt;/strong&gt;, not just a demo node&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The real value of ESP-IDF isn't simply that it's lower level. It's that it &lt;strong&gt;keeps the architecture honest&lt;/strong&gt;: BSP, drivers, protocols, tasks, state, and operations layers can be separated more cleanly; new SoCs and official features appear here first; and when something fails, debugging stays closer to the real system boundary. For a production device, that clarity becomes maintenance-cost control.&lt;/p&gt;

&lt;h3&gt;
  
  
  Arduino: excellent for validation, but don't idealize it
&lt;/h3&gt;

&lt;p&gt;The strengths of the Arduino path are real: &lt;strong&gt;very fast board bring-up, wide library availability, tons of example material, and a low barrier&lt;/strong&gt; for hardware validation and light connected nodes.&lt;/p&gt;

&lt;p&gt;If your short-term goal is any of the following, Arduino is a good choice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Validating a board and peripheral path quickly&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Proving a simple connected node is worth pursuing&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Moving with a team that already thinks in the Arduino model&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shipping something structurally simple&lt;/strong&gt; with limited system boundaries&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But it shouldn't be romanticized into the answer for every long-lived product. Once a project grows into complex OTA and version-governance needs, heavy concurrency, tight memory/power control, earlier SoC-feature adoption, or clearer driver/protocol/business-state boundaries, Arduino as the &lt;em&gt;only&lt;/em&gt; architectural center gets much harder.&lt;/p&gt;

&lt;p&gt;The better 2026 judgment: Arduino is still a very good entry path and validation tool, but for complex products it's best treated as a &lt;strong&gt;high-level gateway into the ESP-IDF world&lt;/strong&gt; or an &lt;strong&gt;early-stage route&lt;/strong&gt; — not an infinitely scalable production architecture.&lt;/p&gt;

&lt;h3&gt;
  
  
  ESPHome: strong for device endpoints, not general firmware
&lt;/h3&gt;

&lt;p&gt;ESPHome's biggest strength isn't that it's lower level. It's that it makes a certain class of devices very fast to deliver: &lt;strong&gt;Home Assistant nodes, voice satellites, sensors, actuators, household bridge devices, and YAML-driven component-based devices&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If the device truly exists to serve Home Assistant or a similar smart-home ecosystem, ESPHome often beats writing everything yourself — OTA, logging, Wi-Fi, API behavior, and entity mapping are already aligned, many peripherals already exist, and time-to-result is extremely high.&lt;/p&gt;

&lt;p&gt;But the tradeoffs matter: it's first a &lt;strong&gt;device-delivery system&lt;/strong&gt;, not a universal firmware platform. Once you need deeper business state machines, unusual protocol bridges, or broader architecture beyond the component model, the abstraction starts to limit you. And if the product isn't fundamentally a smart-home ecosystem node, many of its built-in advantages stop being structural advantages.&lt;/p&gt;

&lt;p&gt;The fair statement isn't that ESPHome is unprofessional — it's that ESPHome is &lt;strong&gt;highly professional for a Home Assistant-oriented device class&lt;/strong&gt;, but it doesn't replace a general custom ESP32 firmware architecture.&lt;/p&gt;

&lt;h3&gt;
  
  
  Zephyr: for platform unification, not single-vendor teams
&lt;/h3&gt;

&lt;p&gt;The reason to choose Zephyr isn't merely that it also supports ESP32. It becomes rational when the project already prioritizes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;One RTOS architecture across multiple MCU vendors&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stronger consistency&lt;/strong&gt; in threads, device models, configuration, and platform conventions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A team already comfortable&lt;/strong&gt; with &lt;code&gt;west&lt;/code&gt;, Kconfig, device trees, and platform governance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In that situation Zephyr is valuable because it answers a &lt;strong&gt;platform-level question&lt;/strong&gt;, not a single-product firmware question. But for teams focused mainly on ESP32 products, the costs are clear: higher cognitive overhead, less direct alignment with Espressif-first troubleshooting, and sometimes slower access to the newest chip-specific features or familiar community knowledge.&lt;/p&gt;

&lt;p&gt;The rule is straightforward: &lt;strong&gt;Zephyr deserves priority only when platform unification itself is a hard business or architectural requirement.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Compare by project boundaries, not by hype
&lt;/h2&gt;

&lt;p&gt;In 2026, "best" should not mean "most talked about." It should mean &lt;strong&gt;least contradictory to the real constraints of your product&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. A practical default ordering for 2026
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The default sequence for most custom firmware teams
&lt;/h3&gt;

&lt;p&gt;For most commercial ESP32 work, a reasonable default sequence is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Start by asking whether ESP-IDF can simply be the base.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;If the current phase is primarily &lt;strong&gt;hardware validation&lt;/strong&gt; and the team is clearly Arduino-shaped, evaluate &lt;strong&gt;Arduino&lt;/strong&gt; next.&lt;/li&gt;
&lt;li&gt;If the device is naturally a &lt;strong&gt;Home Assistant or smart-home endpoint&lt;/strong&gt;, evaluate &lt;strong&gt;ESPHome&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Only prioritize &lt;strong&gt;Zephyr&lt;/strong&gt; when platform consistency matters more than single-chip efficiency.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This ordering isn't saying the other paths are bad. It's saying that by 2026, the most expensive mistake is no longer "starting with a framework that feels harder." The more expensive mistake is &lt;strong&gt;keeping a production product trapped too long inside an abstraction that's either too high-level or too ecosystem-specific for what the device is becoming.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  One practical self-check
&lt;/h3&gt;

&lt;p&gt;If the choice still feels unclear, ask these four questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Will this device &lt;strong&gt;ship and be maintained long term&lt;/strong&gt;?&lt;/li&gt;
&lt;li&gt;Is it fundamentally a &lt;strong&gt;Home Assistant or smart-home endpoint&lt;/strong&gt;?&lt;/li&gt;
&lt;li&gt;Does this product line need to &lt;strong&gt;share one RTOS platform across multiple MCU vendors&lt;/strong&gt;?&lt;/li&gt;
&lt;li&gt;Is this decision for &lt;strong&gt;this month's prototype&lt;/strong&gt;, or for &lt;strong&gt;the next 12 months of maintainable firmware&lt;/strong&gt;?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In many cases, those four answers decide the framework without any popularity ranking at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. When NOT to choose each path
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Don't default to Arduino&lt;/strong&gt; for a product with known future complexity in drivers, OTA, state machines, and operations, promising to "refactor later."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't stretch ESPHome beyond its natural boundary&lt;/strong&gt; if the project isn't fundamentally about delivering a Home Assistant-oriented endpoint.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't reach for Zephyr first&lt;/strong&gt; if your real goal is simply to deliver one stable ESP32 product — the official stack is usually the more direct, lower-risk route.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  6. Final judgment
&lt;/h2&gt;

&lt;p&gt;Choosing an ESP32 framework in 2026 is less about picking one of four isolated universes and more about deciding &lt;strong&gt;how far above ESP-IDF you want to work&lt;/strong&gt;, and whether you truly need either a Home Assistant device abstraction or a cross-MCU platform abstraction.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Building a &lt;strong&gt;long-lived commercial firmware system&lt;/strong&gt; → start with &lt;strong&gt;ESP-IDF&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Building a &lt;strong&gt;prototype or simple device&lt;/strong&gt; → &lt;strong&gt;Arduino&lt;/strong&gt; remains valuable.&lt;/li&gt;
&lt;li&gt;Building a &lt;strong&gt;Home Assistant endpoint&lt;/strong&gt; → &lt;strong&gt;ESPHome&lt;/strong&gt; is a strong fit.&lt;/li&gt;
&lt;li&gt;Building a &lt;strong&gt;multi-vendor RTOS platform&lt;/strong&gt; → &lt;strong&gt;Zephyr&lt;/strong&gt; becomes rational.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The best framework isn't the hottest one. It's the one that makes your product boundaries &lt;strong&gt;least self-contradictory over time&lt;/strong&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Which framework are you currently using for your ESP32 work — and have you ever had to migrate off it as the product grew? I'd love to hear the story in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>esp32</category>
      <category>iot</category>
      <category>arduino</category>
      <category>esphome</category>
    </item>
    <item>
      <title>Building a Reliable ESP32-S3 Voice Satellite: I2S, PDM, and the Audio Pipeline You're Ignoring</title>
      <dc:creator>ZedIoT</dc:creator>
      <pubDate>Tue, 15 Sep 2026 07:45:04 +0000</pubDate>
      <link>https://dev.to/zediot/building-a-reliable-esp32-s3-voice-satellite-i2s-pdm-and-the-audio-pipeline-youre-ignoring-47kh</link>
      <guid>https://dev.to/zediot/building-a-reliable-esp32-s3-voice-satellite-i2s-pdm-and-the-audio-pipeline-youre-ignoring-47kh</guid>
      <description>&lt;h1&gt;
  
  
  Building a Reliable ESP32-S3 Voice Satellite: I2S, PDM, and the Audio Pipeline You're Ignoring
&lt;/h1&gt;

&lt;p&gt;When a Home Assistant voice satellite built on ESP32-S3 misses commands, answers slowly, or cuts off mid-reply, the first instinct is to blame the &lt;strong&gt;wake word model&lt;/strong&gt; or the &lt;strong&gt;microphone sensitivity&lt;/strong&gt;. Those matter — but they're not the whole system.&lt;/p&gt;

&lt;p&gt;The real truth: the user experience of an ESP32-S3 voice node is determined by &lt;strong&gt;microphone capture → I2S/PDM timing → device-side buffering → Wi-Fi upload → the Home Assistant Assist pipeline → TTS return audio → speaker playback&lt;/strong&gt;, working together. If any single boundary stalls, jitters, or competes for CPU, the final symptom is the same: "slow, unreliable, or hard to understand."&lt;/p&gt;

&lt;p&gt;ESPHome's own Voice Assistant documentation warns that audio and voice components consume &lt;strong&gt;significant RAM and CPU&lt;/strong&gt;, and that Bluetooth/BLE components can cause issues when run alongside voice. That warning should be read as an &lt;strong&gt;architecture boundary&lt;/strong&gt;, not a footnote. A voice satellite is not a board with a microphone glued on — it's a continuous real-time audio path squeezed through a constrained MCU, a wireless network, and a home automation platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The real voice path is longer than the YAML file
&lt;/h2&gt;

&lt;p&gt;ESPHome's &lt;code&gt;voice_assistant&lt;/code&gt; component lets an ESP32 send microphone audio to Home Assistant Assist for processing. The Assist pipeline typically includes &lt;strong&gt;wake word detection, speech-to-text, intent recognition, and text-to-speech&lt;/strong&gt;. The split is elegant: the small device handles capture and playback, while Home Assistant handles understanding and action.&lt;/p&gt;

&lt;p&gt;But latency accumulates across that split. A single voice interaction quietly stacks up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Microphone sampling&lt;/strong&gt; and local buffering&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wake or push-to-talk&lt;/strong&gt; activation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wi-Fi upload&lt;/strong&gt; of audio chunks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Home Assistant STT, intent, and TTS&lt;/strong&gt; processing&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Return audio delivery&lt;/strong&gt; and speaker playback&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When a voice assistant feels slow, the cause is rarely one function. It's usually that &lt;strong&gt;capture, network, pipeline, and playback latency were never measured separately&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. I2S and PDM are about clocks and buffers — not just pin names
&lt;/h2&gt;

&lt;p&gt;ESPHome's &lt;code&gt;i2s_audio&lt;/code&gt; component handles sending and receiving audio on ESP32-family chips. A standard &lt;strong&gt;I2S&lt;/strong&gt; bus uses BCLK, LRCLK/WS, and DIN/DOUT, while &lt;strong&gt;PDM&lt;/strong&gt; microphones use a different clock and data pattern. Espressif's ESP32-S3 I2S documentation treats standard I2S, TDM, and PDM as &lt;strong&gt;distinct modes&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For a voice satellite, the I2S-vs-PDM choice should not come down to module price. The stronger questions are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does the &lt;strong&gt;microphone output mode&lt;/strong&gt; match what the ESPHome component supports?&lt;/li&gt;
&lt;li&gt;Do &lt;strong&gt;sample rate, bit width, and channel settings&lt;/strong&gt; match what the Assist pipeline expects?&lt;/li&gt;
&lt;li&gt;Can the device &lt;strong&gt;buffer audio through short Wi-Fi, logging, and playback jitter&lt;/strong&gt;?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One sharp gotcha: ESPHome notes that PDM microphone support is &lt;strong&gt;primarily available on ESP32 and ESP32-S3&lt;/strong&gt;. The same config cannot be blindly moved across ESP32 variants and assumed to behave identically.&lt;/p&gt;

&lt;p&gt;A working I2S/PDM config only proves the device can capture audio. It does &lt;strong&gt;not&lt;/strong&gt; prove the voice stream stays stable under network jitter and playback competition.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. ESP32-S3 is a good voice node — but not an unlimited one
&lt;/h2&gt;

&lt;p&gt;ESP32-S3 fits voice work better than older ESP32 choices because it brings &lt;strong&gt;dual cores, Wi-Fi, BLE 5.0, native USB, and AI vector instructions&lt;/strong&gt; that help with tasks like &lt;strong&gt;Micro Wake Word&lt;/strong&gt;. ESPHome's platform docs single out ESP32-S3 as especially useful for ML applications like Micro Wake Word.&lt;/p&gt;

&lt;p&gt;That still doesn't make it unlimited. A voice satellite is often already running:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Continuous microphone capture&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Wake or button activation&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;API or WebSocket transport&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LED status indication&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Speaker playback&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Logs and remote debugging&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the same node also owns &lt;strong&gt;BLE scanning, complex sensors, display animation, Matter/Thread roles, or high-frequency automations&lt;/strong&gt;, resource competition becomes the real failure mode. ESPHome's audio/voice resource warning should define the node's scope.&lt;/p&gt;

&lt;p&gt;When a node owns voice &lt;em&gt;and&lt;/em&gt; Bluetooth scanning &lt;em&gt;and&lt;/em&gt; UI &lt;em&gt;and&lt;/em&gt; several sensor loops, failure usually shows up first as &lt;strong&gt;audio dropouts or intermittent restarts&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Recommended layering: make each audio boundary observable
&lt;/h2&gt;

&lt;p&gt;The pipeline runs in a strict sequence, and each stage should be observable on its own:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MEMS microphone → I2S/PDM capture → device buffer → Wi-Fi upload → Home Assistant Assist pipeline → TTS return → I2S speaker playback → user response&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The point is simple: &lt;strong&gt;don't debug "bad voice" as one vague problem.&lt;/strong&gt; Each stage should be testable in isolation.&lt;/p&gt;

&lt;p&gt;For example, test the microphone path with &lt;strong&gt;short repeated phrases&lt;/strong&gt; and inspect noise, clipping, and gain &lt;em&gt;before&lt;/em&gt; entering a full conversation. Watch device stability and logs before adding optional components. Use Home Assistant's &lt;strong&gt;pipeline debug tools&lt;/strong&gt; to isolate STT and intent behavior. Test speaker output with a &lt;strong&gt;fixed TTS or prompt sound&lt;/strong&gt; before combining it with the full interaction.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Common bottlenecks and safer fixes
&lt;/h2&gt;

&lt;p&gt;Diagnostic order matters because the voice path is &lt;strong&gt;sequential&lt;/strong&gt;. If capture is weak, a better STT engine still receives poor audio. If the Assist pipeline is slow, raising microphone gain won't make TTS return any sooner.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. A practical debugging sequence
&lt;/h2&gt;

&lt;p&gt;A deployable ESP32-S3 voice node should be tested in this order:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Test raw microphone input first.&lt;/strong&gt; Use fixed short phrases and check noise floor, clipping, volume, and room noise before running the full Assist flow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validate device stability.&lt;/strong&gt; After enabling voice components, disable unnecessary BLE, display, sensor polling, and verbose logs. Confirm the device runs without restart.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test the Assist pipeline separately.&lt;/strong&gt; Use Home Assistant's debug or text pipeline tools to confirm intent recognition works before blaming the satellite.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add TTS playback later.&lt;/strong&gt; Play fixed prompts or fixed TTS first, then validate amplifier, power, and speaker behavior.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Move to the real room last.&lt;/strong&gt; Test distance, background noise, router placement, and multiple speakers in the intended location.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Voice satellite debugging should start with &lt;strong&gt;raw audio and pipeline segmentation&lt;/strong&gt;, not with repeated edits to the full YAML file.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. When a basic ESP32-S3 voice satellite is the wrong tool
&lt;/h2&gt;

&lt;p&gt;ESP32-S3 + ESPHome is a strong fit for &lt;strong&gt;room-level voice entry points, push-to-talk nodes, near-field control, desk satellites, and Home Assistant prototypes&lt;/strong&gt;. But some requirements should not be forced through a basic dev-board design:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Far-field pickup and beamforming&lt;/strong&gt; in a living room&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Noisy kitchens, workshops, or commercial spaces&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fully local STT/TTS&lt;/strong&gt; with response time close to commercial smart speakers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-room conversational behavior&lt;/strong&gt;, echo cancellation, and playback coordination&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Productized hardware&lt;/strong&gt; with enclosure acoustics, certification, and long-term support&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those cases are better served by &lt;strong&gt;dedicated voice hardware, microphone arrays, audio processors&lt;/strong&gt;, or a design where ESP32-S3 acts only as a &lt;strong&gt;button, LED, or near-field capture node&lt;/strong&gt; instead of owning the entire voice experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Conclusion: stabilize the audio path before optimizing intelligence
&lt;/h2&gt;

&lt;p&gt;ESP32-S3 voice satellites are valuable because they're &lt;strong&gt;low cost, customizable, and tightly integrated&lt;/strong&gt; with Home Assistant and ESPHome. They can distribute local smart-home control across rooms and make voice prototypes easy to build.&lt;/p&gt;

&lt;p&gt;Their success condition is &lt;em&gt;not&lt;/em&gt; "the Voice Assistant example compiles." It's that the end-to-end path is &lt;strong&gt;explainable&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Microphone capture is stable&lt;/strong&gt; and not over-amplifying noise&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I2S/PDM timing and buffers&lt;/strong&gt; survive short jitter&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;ESP32-S3 node avoids unrelated heavy tasks&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;Assist pipeline can be debugged independently&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TTS and speaker playback are verified on their own&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without these boundaries, every problem looks like poor recognition. With them, ESP32-S3 becomes a reliable voice satellite — instead of a dev board that only &lt;em&gt;sometimes&lt;/em&gt; understands you.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;What's the hardest part you've hit when tuning your own ESP32 voice node — microphone gain, Wi-Fi jitter, or the Assist pipeline itself? Drop it in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>esp32</category>
      <category>iot</category>
      <category>esphome</category>
      <category>homeassistant</category>
    </item>
    <item>
      <title>Debugging Long-Uptime ESPHome Devices on ESP32</title>
      <dc:creator>ZedIoT</dc:creator>
      <pubDate>Tue, 08 Sep 2026 12:20:00 +0000</pubDate>
      <link>https://dev.to/zediot/debugging-long-uptime-esphome-devices-on-esp32-b3l</link>
      <guid>https://dev.to/zediot/debugging-long-uptime-esphome-devices-on-esp32-b3l</guid>
      <description>&lt;h1&gt;
  
  
  Debugging Long-Uptime ESPHome Devices on ESP32
&lt;/h1&gt;

&lt;p&gt;ESPHome devices that fail after days or weeks usually suffer from accumulated system effects: heap pressure, fragmentation, Wi-Fi instability, blocking components, sensor timing, and stale values.&lt;/p&gt;

&lt;p&gt;Many ESPHome devices look stable right after flashing. Sensors report values, Home Assistant discovers entities, and automations work. The harder failures show up later: the node reboots after several days, the API disconnects, a sensor value freezes, or the only fix seems to be power cycling the device.&lt;/p&gt;

&lt;p&gt;The core conclusion is straightforward: long-uptime ESPHome failures are rarely caused by one bad YAML line. They are usually accumulated system effects across memory behavior, blocking components, Wi-Fi conditions, logging, and sensor timing. If the device does not expose uptime, reset reason, free heap, minimum free heap, fragmentation, Wi-Fi signal, and last valid readings, it is difficult to tell a memory leak from heap fragmentation, a network issue, or a stalled peripheral.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd5i59emlqlzqe55u0aft.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd5i59emlqlzqe55u0aft.webp" alt=" " width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What long-uptime debugging means:&lt;/strong&gt; diagnosing devices that work at first but fail only after days or weeks. The target is not compile errors or a single wiring mistake. The target is reboot patterns, stale values, intermittent disconnections, and runtime health signals.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Why "it ran for one day" is not a stability test
&lt;/h2&gt;

&lt;p&gt;ESP32 and ESPHome prototypes can be misleading. Once the device appears in Home Assistant and updates a few entities, it is tempting to treat the firmware as finished. Long runtime exposes problems that short bench tests miss.&lt;/p&gt;

&lt;p&gt;Common long-uptime failure sources include heap leaks, memory fragmentation, Wi-Fi reconnect storms, blocking I2C/UART calls, log flooding, and stale sensor values.&lt;/p&gt;

&lt;p&gt;The key rule: if an ESPHome node exposes only business sensors and no runtime diagnostics, a failure after several weeks becomes guesswork instead of engineering analysis.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Add diagnostic entities before changing the design
&lt;/h2&gt;

&lt;p&gt;The first response should not be rewriting the YAML. The first response should be making runtime health visible. A useful minimum set is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;uptime&lt;/strong&gt; — so every restart becomes visible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;reset reason&lt;/strong&gt; — so software restarts, watchdogs, brownouts, and power resets are not mixed together.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;free heap&lt;/strong&gt; — to track current memory availability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;minimum free heap&lt;/strong&gt; — to catch low points that disappear after reboot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;fragmentation or maximum block size&lt;/strong&gt; — to expose fragmented heap behavior.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wi-Fi signal&lt;/strong&gt; — to avoid treating radio problems as firmware crashes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;last valid reading&lt;/strong&gt; — to distinguish stale data from fresh data.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;debug&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;update_interval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;60s&lt;/span&gt;

&lt;span class="na"&gt;sensor&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;platform&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;uptime&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Node&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Uptime"&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;platform&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;debug&lt;/span&gt;
    &lt;span class="na"&gt;free&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Heap&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Free"&lt;/span&gt;
    &lt;span class="na"&gt;block&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Heap&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Max&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Block"&lt;/span&gt;
    &lt;span class="na"&gt;loop_time&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Loop&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Time"&lt;/span&gt;

&lt;span class="na"&gt;text_sensor&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;platform&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;debug&lt;/span&gt;
    &lt;span class="na"&gt;reset_reason&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Reset&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Reason"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is not a full production template. It is the debugging boundary: business entities describe the environment, while diagnostic entities describe whether the node itself can still be trusted.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo0nb63vscz7kgzcqkasx.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo0nb63vscz7kgzcqkasx.webp" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Use one diagnostic path to narrow the failure
&lt;/h2&gt;

&lt;p&gt;The decision path starts from the observed anomaly:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Start from the &lt;strong&gt;device anomaly&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Ask: &lt;strong&gt;did uptime reset?&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;If &lt;strong&gt;yes&lt;/strong&gt;, check the &lt;strong&gt;reset reason&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;If &lt;strong&gt;no&lt;/strong&gt;, ask whether &lt;strong&gt;business values are stale&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;From reset reason, correlate &lt;strong&gt;heap, Wi-Fi, and power&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;From stale values, check &lt;strong&gt;blocked components and bus errors&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Both paths converge on a &lt;strong&gt;minimal reproduction&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Then &lt;strong&gt;change one variable and observe for 3–7 days&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The first split matters: did the device really reboot, or did one part of the data path stall? Reboots push the investigation toward reset reason, heap, power, and watchdog behavior. Stale values without a reboot push it toward sensor drivers, I2C or UART behavior, blocking calls, and external services.&lt;/p&gt;

&lt;p&gt;Do not change Wi-Fi, logging, sampling intervals, sensor configuration, and power at the same time. Long-uptime failures already take time to reproduce. Changing several variables at once makes the next result harder to interpret.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Heap debugging is about low points and fragmentation, not only current free memory
&lt;/h2&gt;

&lt;p&gt;Many ESP32 nodes have enough free heap right after boot. After days of runtime, two different problems can appear:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Total free heap gradually drops&lt;/strong&gt;, which can indicate a leak or unbounded cache.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Total free heap looks acceptable, but the largest contiguous block shrinks&lt;/strong&gt;, so larger allocations fail.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is why current free heap is not enough. A minimum-free signal can expose short low-memory events, while fragmentation or largest-block diagnostics can show that memory is available but not available in useful contiguous chunks.&lt;/p&gt;

&lt;p&gt;If an ESP32 node reboots only after reconnects, sensor faults, display refreshes, or bursts of logging, observe heap low points and largest block size before blaming the last visible component.&lt;/p&gt;

&lt;p&gt;A practical narrowing sequence is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Reduce log verbosity so the device is not spending long periods formatting and transmitting logs.&lt;/li&gt;
&lt;li&gt;Temporarily remove nonessential components such as web server, display, Bluetooth scanning, or high-frequency template sensors.&lt;/li&gt;
&lt;li&gt;Increase sensor &lt;code&gt;update_interval&lt;/code&gt; to see whether a specific sampling cadence triggers the failure.&lt;/li&gt;
&lt;li&gt;Remove complex lambda code and string formatting to see whether the heap curve stabilizes.&lt;/li&gt;
&lt;li&gt;Run the same configuration on another board and power supply to separate firmware behavior from hardware variance.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  5. Wi-Fi and API disconnects are not always firmware crashes
&lt;/h2&gt;

&lt;p&gt;An ESPHome device showing offline in Home Assistant does not automatically mean the MCU crashed. Wi-Fi roaming, weak RSSI, router restarts, API connection behavior, mDNS resolution, and network congestion can all look like device failure from the dashboard.&lt;/p&gt;

&lt;p&gt;Ask two questions first:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Did uptime reset?&lt;/strong&gt; If not, the firmware may still be running.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is there serial or local log output?&lt;/strong&gt; If yes, the problem may be the network or API path.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For devices inside metal cabinets, distribution boxes, cold rooms, equipment rooms, or industrial spaces, radio quality is part of device stability. Do not repair a network problem as a firmware problem. Add Wi-Fi signal, connection state, and last publish time first; then decide whether to move the router, change the antenna, use Ethernet, or delegate the critical path to a more reliable gateway.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. When ESPHome is the wrong abstraction
&lt;/h2&gt;

&lt;p&gt;ESPHome is excellent for configurable Home Assistant devices, small sensor gateways, and fast integration work. It becomes less suitable when the node turns into a production controller with complex runtime requirements.&lt;/p&gt;

&lt;p&gt;Be cautious when the project needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;strict real-time control, complex state machines, or safety interlocks.&lt;/li&gt;
&lt;li&gt;local queues, protocol retries, persistent buffering, or multiple coordinated tasks.&lt;/li&gt;
&lt;li&gt;staged OTA, remote log collection, self-recovery, and fleet operations.&lt;/li&gt;
&lt;li&gt;code-level control over memory, tasks, stack behavior, and peripheral failures.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical boundary is this: use ESPHome for observable, configurable, low-friction edge nodes. When the device becomes a production gateway or controller, consider ESP-IDF, custom firmware, or moving the complex logic into an edge gateway or platform service.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://esphome.io/components/debug.html" rel="noopener noreferrer"&gt;ESPHome Debug Component&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://esphome.io/components/sensor/uptime.html" rel="noopener noreferrer"&gt;ESPHome Uptime Sensor&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://esphome.io/components/sensor/wifi_signal.html" rel="noopener noreferrer"&gt;ESPHome WiFi Signal Sensor&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://esphome.io/components/logger.html" rel="noopener noreferrer"&gt;ESPHome Logger Component&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://esphome.io/components/api.html" rel="noopener noreferrer"&gt;ESPHome API Component&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>esphome</category>
      <category>esp32</category>
      <category>homeassistant</category>
      <category>iot</category>
    </item>
    <item>
      <title>Migrating MediaPipe Gesture Recognition to RKNN: From .task to RK3566 Deployment</title>
      <dc:creator>ZedIoT</dc:creator>
      <pubDate>Thu, 03 Sep 2026 12:20:00 +0000</pubDate>
      <link>https://dev.to/zediot/migrating-mediapipe-gesture-recognition-to-rknn-from-task-to-rk3566-deployment-45a4</link>
      <guid>https://dev.to/zediot/migrating-mediapipe-gesture-recognition-to-rknn-from-task-to-rk3566-deployment-45a4</guid>
      <description>&lt;h1&gt;
  
  
  Migrating MediaPipe Gesture Recognition to RKNN: From .task to RK3566 Deployment
&lt;/h1&gt;

&lt;p&gt;A step-by-step guide to convert MediaPipe Gesture Recognition models to RKNN and running real-time inference on the RK3566 NPU, with code and troubleshooting tips.&lt;/p&gt;

&lt;p&gt;MediaPipe gesture recognition is an important human–computer interaction technique in computer vision. Google's MediaPipe offers a complete Gesture Recognizer pipeline, which includes four stages: hand detection, hand landmark detection, embedding generation, and gesture classification. This enables real-time, end-to-end gesture recognition across many applications.&lt;/p&gt;

&lt;p&gt;However, MediaPipe is primarily optimized for PC CPU, NVIDIA GPU, and Android GPU environments. Running the same models directly on embedded SoCs such as Rockchip RK3566 or RK3588 leads to inefficient performance.&lt;/p&gt;

&lt;p&gt;To bring MediaPipe gesture recognition to embedded hardware, each model must be converted to RKNN format. Rockchip provides RKNN Toolkit 2, which converts mainstream deep-learning models (TFLite, ONNX, Caffe, etc.) into the &lt;code&gt;.rknn&lt;/code&gt; format. These models can then run on the NPU for hardware acceleration. With this toolchain, the MediaPipe gesture recognition pipeline can be migrated onto RK3566 for low-power, high-performance edge inference.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model Components in MediaPipe Gesture Recognizer
&lt;/h2&gt;

&lt;p&gt;The MediaPipe Gesture Recognition pipeline is built from four separate TFLite models. The &lt;code&gt;gesture_recognizer.task&lt;/code&gt; file is not a single model. It is a Task Bundle that includes several &lt;code&gt;.tflite&lt;/code&gt; models and configuration files. After unpacking it, you will see:&lt;/p&gt;

&lt;h3&gt;
  
  
  hand_landmarker.task
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;hand_detector.tflite&lt;/code&gt; — Palm/hand detection&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;hand_landmarks_detector.tflite&lt;/code&gt; — 21 hand keypoints&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  hand_gesture_recognizer.task
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;gesture_embedder.tflite&lt;/code&gt; — Converts keypoints into embedding vectors&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;canned_gesture_classifier.tflite&lt;/code&gt; — Classifies hand gestures&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Together, they form this pipeline:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Hand Detection → Landmark Detection → Embedding → Gesture Classification
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Why convert them one by one?
&lt;/h3&gt;

&lt;p&gt;RKNN Toolkit 2 cannot parse &lt;code&gt;.task&lt;/code&gt; files, so the four TFLite models must be extracted and each converted to &lt;code&gt;.rknn&lt;/code&gt; individually.&lt;/p&gt;

&lt;p&gt;During inference, the four RKNN models must be called sequentially, following MediaPipe's original order, to reproduce the complete gesture pipeline.&lt;/p&gt;

&lt;h3&gt;
  
  
  Summary of migration steps
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Unpack &lt;code&gt;.task&lt;/code&gt; → extract &lt;code&gt;.tflite&lt;/code&gt; models&lt;/li&gt;
&lt;li&gt;Convert each TFLite model → &lt;code&gt;.rknn&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Build the pipeline on RK3566 → run full gesture recognition&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2m8isz6rlhgu9qm83bb0.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2m8isz6rlhgu9qm83bb0.webp" alt=" " width="800" height="534"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Model Conversion Workflow
&lt;/h2&gt;

&lt;p&gt;During conversion, each MediaPipe Gesture Recognition model must preserve its original preprocessing rules, or the output will drift. The workflow runs in three stages:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage A — Unpacking (PC):&lt;/strong&gt; extract the four TFLite models (&lt;code&gt;hand_detector.tflite&lt;/code&gt;, &lt;code&gt;hand_landmarks_detector.tflite&lt;/code&gt;, &lt;code&gt;gesture_embedder.tflite&lt;/code&gt;, &lt;code&gt;canned_gesture_classifier.tflite&lt;/code&gt;) from &lt;code&gt;gesture_recognizer.task&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage B — Conversion (PC):&lt;/strong&gt; set &lt;code&gt;rknn.config(target='rk3566', w8a8)&lt;/code&gt;, provide a &lt;code&gt;dataset.txt&lt;/code&gt; for the image models, then run &lt;code&gt;load_tflite → build → export_rknn&lt;/code&gt; for each model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage C — Deployment (RK3566):&lt;/strong&gt; feed the input frame through the four RKNN models sequentially and emit the gesture output.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1. Install RKNN Toolkit 2
&lt;/h3&gt;

&lt;p&gt;Install on Windows/Linux x86_64 (Mac requires a VM/container):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;rknn-toolkit2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Recommended Python version: 3.6–3.10. Version 2.3.2 is widely used.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2. Prepare the Four TFLite Models
&lt;/h3&gt;

&lt;p&gt;Extract these from the task bundle:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;hand_detector.tflite&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;hand_landmarks_detector.tflite&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;gesture_embedder.tflite&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;canned_gesture_classifier.tflite&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Step 3. Write the Conversion Script
&lt;/h3&gt;

&lt;p&gt;Each model inside the MediaPipe task file must be converted from TFLite to RKNN before it can run on the RK3566 NPU.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Important:&lt;/strong&gt; In recent RKNN Toolkit versions, you must call &lt;code&gt;rknn.config()&lt;/code&gt; before &lt;code&gt;load_tflite()&lt;/code&gt;.&lt;/p&gt;

&lt;h4&gt;
  
  
  Template conversion script
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;rknn.api&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;RKNN&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;convert_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tflite_file&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rknn_file&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;is_image_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;rknn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNN&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="c1"&gt;# Configuration
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;is_image_model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;mean_values&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt;
            &lt;span class="n"&gt;std_values&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt;
            &lt;span class="n"&gt;target_platform&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;rk3566&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;quantized_dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;w8a8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;config&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;target_platform&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;rk3566&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;quantized_dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;w8a8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Load TFLite
&lt;/span&gt;    &lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_tflite&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;tflite_file&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Build model
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;is_image_model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;build&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;do_quantization&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dataset&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;dataset.txt&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;build&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;do_quantization&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Export
&lt;/span&gt;    &lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;export_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rknn_file&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;rknn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;release&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# Convert all models
&lt;/span&gt;&lt;span class="nf"&gt;convert_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;hand_detector.tflite&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;hand_detector.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;convert_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;hand_landmarks_detector.tflite&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;hand_landmarks_detector.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;convert_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;gesture_embedder.tflite&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;gesture_embedder.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;convert_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;canned_gesture_classifier.tflite&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;canned_gesture_classifier.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Common Errors and Fixes
&lt;/h4&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;code&gt;E config: Invalid quantized_dtype 'asymmetric_quantized-u8'&lt;/code&gt;&lt;/strong&gt;&lt;br&gt;
Cause: RKNN Toolkit 2.3.2 no longer supports this dtype.&lt;br&gt;
Fix: Use &lt;code&gt;quantized_dtype='w8a8'&lt;/code&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;code&gt;E load_tflite: Please call rknn.config first!&lt;/code&gt;&lt;/strong&gt;&lt;br&gt;
Cause: Incorrect function order.&lt;br&gt;
Fix: Always call &lt;code&gt;config()&lt;/code&gt; before &lt;code&gt;load_tflite()&lt;/code&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;code&gt;E build: Dataset file dataset.txt not found!&lt;/code&gt;&lt;/strong&gt;&lt;br&gt;
Cause: Quantization requires a calibration dataset.&lt;br&gt;
Fix (choose one): prepare &lt;code&gt;dataset.txt&lt;/code&gt; with RGB image paths, or disable quantization (&lt;code&gt;do_quantization=False&lt;/code&gt;).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;RKNN file is empty&lt;/strong&gt;&lt;br&gt;
Cause: &lt;code&gt;build()&lt;/code&gt; failed.&lt;br&gt;
Fix: Check build logs, then fix the dataset or parameters.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Building the Full Gesture Pipeline on RK3566
&lt;/h2&gt;

&lt;p&gt;When running MediaPipe Gesture Recognition on the RK3566 NPU, the four RKNN models must be executed in sequence.&lt;/p&gt;

&lt;h3&gt;
  
  
  Data Flow (RKNN version)
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Frame
  → hand_detector.rknn → box → crop/normalize ROI
  → hand_landmarks_detector.rknn → 21 keypoints → flatten/normalize
  → gesture_embedder.rknn → embedding
  → canned_gesture_classifier.rknn → gesture ID
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Keep the preprocessing steps the same as MediaPipe's (RGB, 0–1 normalization, input sizes; vector models take float32 directly).&lt;/p&gt;

&lt;h2&gt;
  
  
  Single-Image Inference Example (Minimal Working Path)
&lt;/h2&gt;

&lt;p&gt;This example runs MediaPipe Gesture Recognition on RK3566 using four independent RKNN models.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note: Different versions of the hand detector and landmark models may output different formats. Some return multiple boxes and scores, while others output center + size. This example shows a common parsing template. If your model behaves differently, print the output shapes and adjust accordingly.&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;rknn.api&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;RKNN&lt;/span&gt;

&lt;span class="c1"&gt;# ========== Utility functions ==========
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;to_rgb_norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img_bgr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;img&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cvtColor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img_bgr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;COLOR_BGR2RGB&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;img&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;resize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;img&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;astype&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;float32&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mf"&gt;255.0&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;img&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;crop_by_box&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img_bgr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;box_xyxy&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pad&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;w&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;img_bgr&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;box_xyxy&lt;/span&gt;
    &lt;span class="c1"&gt;# Add padding to avoid tight crops
&lt;/span&gt;    &lt;span class="n"&gt;cx&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="n"&gt;cy&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
    &lt;span class="n"&gt;bw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;bh&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;bw&lt;/span&gt; &lt;span class="o"&gt;*=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;pad&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;bh&lt;/span&gt; &lt;span class="o"&gt;*=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;pad&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;x1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cx&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;bw&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cx&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;bw&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;y1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cy&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;bh&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cy&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;bh&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;img_bgr&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;y2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;copy&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;norm_landmarks_to_roi_xy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lm_21x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;roi_xyxy&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;roi_xyxy&lt;/span&gt;
    &lt;span class="n"&gt;rw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;x1&lt;/span&gt;
    &lt;span class="n"&gt;rh&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;
    &lt;span class="c1"&gt;# Convert normalized ROI coordinates (0-1) back to image coordinates
&lt;/span&gt;    &lt;span class="n"&gt;pts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;21&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;x1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;lm_21x2&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;rw&lt;/span&gt;
        &lt;span class="n"&gt;y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;lm_21x2&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;rh&lt;/span&gt;
        &lt;span class="n"&gt;pts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pts&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;float32&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# ========== Load models ==========
&lt;/span&gt;&lt;span class="n"&gt;det&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNN&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;det&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;hand_detector.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;det&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;init_runtime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;lm&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNN&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;lm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;hand_landmarks_detector.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;lm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;init_runtime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;emb&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNN&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;emb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;gesture_embedder.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;emb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;init_runtime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;clf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNN&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;clf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;canned_gesture_classifier.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;clf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;init_runtime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# ========== Run inference on one image ==========
&lt;/span&gt;&lt;span class="n"&gt;img_path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;test.jpg&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
&lt;span class="n"&gt;ori&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;imread&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img_path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;H&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;W&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ori&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="c1"&gt;# 1) Hand detection
&lt;/span&gt;&lt;span class="n"&gt;det_in_size&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;224&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;224&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;det_in&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;to_rgb_norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ori&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;det_in_size&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;det_in&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expand_dims&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;det_in&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# NHWC
&lt;/span&gt;&lt;span class="n"&gt;det_out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;det&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;det_in&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="c1"&gt;# ★★ Print det_out to confirm its real structure, then adjust parsing accordingly ★★
# Assume format is [N, 6]: x1, y1, x2, y2, score, class (normalized coordinates)
&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;det_out&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;boxes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;reshape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# Adjust if needed
&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;        &lt;span class="c1"&gt;# Score threshold
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;No hand detected&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="nf"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Choose the box with the highest confidence
&lt;/span&gt;&lt;span class="n"&gt;best&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;argmax&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;])]&lt;/span&gt;
&lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;best&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x1&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;W&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y1&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;H&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x2&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;W&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;H&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;roi&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;roi_xyxy&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;crop_by_box&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ori&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;pad&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.15&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 2) Landmark detection
&lt;/span&gt;&lt;span class="n"&gt;lm_in_size&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;224&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;224&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;lm_in&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;to_rgb_norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;roi&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lm_in_size&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;lm_in&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expand_dims&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lm_in&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;lm_out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;lm_in&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="c1"&gt;# ★★ Print lm_out to confirm shape ★★
# Common outputs: [1,21,3] or [1,63] — x,y normalized to ROI
&lt;/span&gt;&lt;span class="n"&gt;lm_arr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lm_out&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]).&lt;/span&gt;&lt;span class="nf"&gt;reshape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;21&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;lm_xy_roi&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lm_arr&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;lm_xy&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;norm_landmarks_to_roi_xy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lm_xy_roi&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;roi_xyxy&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 3) Embedding (flatten 21x2 -&amp;gt; 42)
&lt;/span&gt;&lt;span class="n"&gt;vec_42&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lm_xy_roi&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reshape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;astype&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;float32&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;vec_42&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expand_dims&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vec_42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;emb_out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;emb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;vec_42&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;emb_out&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]).&lt;/span&gt;&lt;span class="nf"&gt;astype&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;float32&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;reshape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 4) Classification
&lt;/span&gt;&lt;span class="n"&gt;logits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;clf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expand_dims&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)])[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;probs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;logits&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;reshape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;gid&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;argmax&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Gesture ID:&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;gid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;score:&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;gid&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;

&lt;span class="c1"&gt;# Visualization
&lt;/span&gt;&lt;span class="nf"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;lm_xy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;astype&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;circle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ori&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;rectangle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ori&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;128&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;putText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ori&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;G:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;gid&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;gid&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
            &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;FONT_HERSHEY_SIMPLEX&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.6&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;imwrite&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;result.jpg&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ori&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Release
&lt;/span&gt;&lt;span class="n"&gt;det&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;release&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;lm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;release&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;emb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;release&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;clf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;release&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Saved result.jpg&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two places must be checked by printing actual output before finalizing the parsing logic:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The shape/meaning of &lt;code&gt;det_out&lt;/code&gt; (some models output center + size; some output absolute or normalized coordinates).&lt;/li&gt;
&lt;li&gt;The shape/meaning of &lt;code&gt;lm_out&lt;/code&gt; (some return 63 dims; some include z/visibility).&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Real-Time Camera Demo (OpenCV + RKNN)
&lt;/h2&gt;

&lt;p&gt;This version handles single-hand recognition. To support multiple hands, loop through all valid detection boxes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;rknn.api&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;RKNN&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;to_rgb_norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;img&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cvtColor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;COLOR_BGR2RGB&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;img&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;resize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;astype&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;float32&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mf"&gt;255.0&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...]&lt;/span&gt;  &lt;span class="c1"&gt;# NHWC
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;det&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNN&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;det&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;hand_detector.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;det&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;init_runtime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;lm&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNN&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;lm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;hand_landmarks_detector.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;lm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;init_runtime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;emb&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNN&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;emb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;gesture_embedder.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;emb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;init_runtime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;clf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RKNN&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;clf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_rknn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;canned_gesture_classifier.rknn&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;clf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;init_runtime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;cap&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;VideoCapture&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;det_size&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;224&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;224&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;lm_size&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;224&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;224&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;frame&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cap&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;break&lt;/span&gt;
        &lt;span class="n"&gt;H&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;W&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="n"&gt;t0&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

        &lt;span class="n"&gt;det_out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;det&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="nf"&gt;to_rgb_norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;det_size&lt;/span&gt;&lt;span class="p"&gt;)])&lt;/span&gt;
        &lt;span class="c1"&gt;# ★★ Parse det_out to obtain the best hand bounding box — adjust as above ★★
&lt;/span&gt;        &lt;span class="n"&gt;boxes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;det_out&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]).&lt;/span&gt;&lt;span class="nf"&gt;reshape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;boxes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;putText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;No hand&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;imshow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;RKNN Gesture&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;waitKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;27&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;break&lt;/span&gt;
            &lt;span class="k"&gt;continue&lt;/span&gt;

        &lt;span class="n"&gt;best&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;argmax&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;boxes&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;])]&lt;/span&gt;
        &lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;best&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;W&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;best&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;H&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;best&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;W&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;best&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;H&lt;/span&gt;
        &lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="n"&gt;x1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;W&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;H&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;roi&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;y2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;copy&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

        &lt;span class="n"&gt;lm_out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="nf"&gt;to_rgb_norm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;roi&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lm_size&lt;/span&gt;&lt;span class="p"&gt;)])&lt;/span&gt;
        &lt;span class="c1"&gt;# ★★ Parse lm_out to get the 21x2 normalized coordinates — adjust as above ★★
&lt;/span&gt;        &lt;span class="n"&gt;lm_arr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lm_out&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]).&lt;/span&gt;&lt;span class="nf"&gt;reshape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;21&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;lm_xy_roi&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lm_arr&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

        &lt;span class="c1"&gt;# Convert back to full-image coordinates
&lt;/span&gt;        &lt;span class="n"&gt;rw&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rh&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y2&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;lm_xy&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;zeros&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="mi"&gt;21&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;float32&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;lm_xy&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;x1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;lm_xy_roi&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;rw&lt;/span&gt;
        &lt;span class="n"&gt;lm_xy&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;lm_xy_roi&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;rh&lt;/span&gt;

        &lt;span class="c1"&gt;# embed &amp;amp; classify
&lt;/span&gt;        &lt;span class="n"&gt;vec_42&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lm_xy_roi&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reshape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;astype&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;float32&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;emb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;vec_42&lt;/span&gt;&lt;span class="p"&gt;])[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]).&lt;/span&gt;&lt;span class="nf"&gt;reshape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;astype&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;float32&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;probs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;clf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inference&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;])[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]).&lt;/span&gt;&lt;span class="nf"&gt;reshape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;gid&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;argmax&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

        &lt;span class="c1"&gt;# draw
&lt;/span&gt;        &lt;span class="nf"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;lm_xy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;astype&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;circle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;rectangle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;128&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;fps&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;t0&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;1e-6&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;putText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;G:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;gid&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; p:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;gid&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; FPS:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;fps&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;imshow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;RKNN Gesture&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;waitKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;27&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;break&lt;/span&gt;

    &lt;span class="n"&gt;cap&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;release&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;cv2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;destroyAllWindows&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;det&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;release&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;lm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;release&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;emb&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;release&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;clf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;release&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Practical Optimization Tips
&lt;/h2&gt;

&lt;p&gt;Optimizing preprocessing and quantization improves the accuracy of MediaPipe Gesture Recognition when deployed on RK3566.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Quantization dataset:&lt;/strong&gt; Prepare 100–300 RGB images of hands with varied skin tones, lighting, and backgrounds for the detector and landmark models. This greatly improves INT8 stability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input size:&lt;/strong&gt; Use the exact input resolution defined in the TFLite models (check with Netron).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consistent preprocessing:&lt;/strong&gt; Use the same RGB layout and 0–1 normalization (mean=0, std=255) as in conversion. Avoid mismatches between training and runtime pipelines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ROI affine alignment:&lt;/strong&gt; If the detector outputs rotated boxes or the model expects upright palms, apply optional rotation alignment before cropping.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pipeline optimization:&lt;/strong&gt; Use frame-to-frame tracking to reduce detector frequency (run the detector once every N frames, run landmarks in between).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-hand support:&lt;/strong&gt; For each box above the threshold, run the remaining three stages independently. Limit the maximum number of hands to keep real-time performance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Threading / CPU pinning:&lt;/strong&gt; On RK3566, use multi-threading to separate camera capture, NPU inference, and drawing for smoother performance.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Quick Troubleshooting Table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Issue&lt;/th&gt;
&lt;th&gt;Likely cause&lt;/th&gt;
&lt;th&gt;Fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Wrong gesture output&lt;/td&gt;
&lt;td&gt;Incorrect model execution order&lt;/td&gt;
&lt;td&gt;Run Detection → Landmark → Embedding → Classification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy drift&lt;/td&gt;
&lt;td&gt;Mismatched preprocessing&lt;/td&gt;
&lt;td&gt;Match RGB layout + 0–1 normalization to conversion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;INT8 accuracy loss&lt;/td&gt;
&lt;td&gt;Weak quantization dataset&lt;/td&gt;
&lt;td&gt;Use 100–300 varied hand images&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Slow frame rate&lt;/td&gt;
&lt;td&gt;Detector runs every frame&lt;/td&gt;
&lt;td&gt;Run detector every N frames, landmarks in between&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Empty &lt;code&gt;.rknn&lt;/code&gt; file&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;build()&lt;/code&gt; failed&lt;/td&gt;
&lt;td&gt;Check logs, fix dataset or params&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;MediaPipe's gesture embedder often applies centering, scale normalization, and mirroring to keypoints. Reproduce these steps when needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Thought
&lt;/h2&gt;

&lt;p&gt;By converting the four MediaPipe gesture-recognition submodels into RKNN models and running them sequentially on RK3566, you can build a low-power, real-time gesture recognition system.&lt;/p&gt;

&lt;p&gt;The key is to maintain identical preprocessing, provide a good quantization dataset, and apply real-world engineering optimizations such as lowering detection frequency or using multithreading. This gives a complete path from MediaPipe → RKNN → RK3566 edge deployment — a strong combination of low-cost hardware and high-value CV capability for smart home, retail, education, and more.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. What models need to be converted when deploying hand-related pipelines on RK3566?
&lt;/h3&gt;

&lt;p&gt;You must convert each model used in the pipeline — detection, landmark extraction, embedding, and classification — into &lt;code&gt;.rknn&lt;/code&gt; format so they can run efficiently on the RK3566 NPU.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. What are the most common issues when converting models to RKNN?
&lt;/h3&gt;

&lt;p&gt;Typical issues include incorrect input size, mismatched normalization, unsupported quantization settings, missing calibration images, or incorrect model ordering during execution.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. What is the correct execution order for hand-analysis pipelines on RK3566?
&lt;/h3&gt;

&lt;p&gt;The standard sequence is Detection → Landmark Extraction → Embedding → Classification. Running them in the wrong order causes incorrect outputs or reduced accuracy.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Can RK3566 achieve real-time performance for hand-analysis workloads?
&lt;/h3&gt;

&lt;p&gt;Yes. With proper preprocessing, model conversion, and optimization, RK3566 can achieve real-time performance while maintaining low power consumption, making it well-suited for embedded AI applications.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>python</category>
    </item>
    <item>
      <title>ESP32 High-Density LED Control with RMT, DMA, and WLED</title>
      <dc:creator>ZedIoT</dc:creator>
      <pubDate>Tue, 01 Sep 2026 12:10:00 +0000</pubDate>
      <link>https://dev.to/zediot/esp32-high-density-led-control-with-rmt-dma-and-wled-2p5b</link>
      <guid>https://dev.to/zediot/esp32-high-density-led-control-with-rmt-dma-and-wled-2p5b</guid>
      <description>&lt;h1&gt;
  
  
  High-Density LED Control with ESP32 RMT and WLED
&lt;/h1&gt;

&lt;p&gt;Driving large WS2812 or SK6812 installations with ESP32 and WLED is not just an MCU performance problem. The real constraints are LEDs per output, RMT interrupt or DMA behavior, SRAM usage, power injection, Wi-Fi load, and synchronization strategy together.&lt;/p&gt;

&lt;p&gt;When an ESP32 + WLED project grows from a short decorative strip to hundreds, thousands, or several thousand addressable LEDs, the bottleneck is rarely just "whether the ESP32 is fast enough." High-density LED control is constrained by output segmentation, serial LED timing, RMT interrupt or DMA behavior, SRAM usage, power injection, Wi-Fi load, and synchronization strategy together.&lt;/p&gt;

&lt;p&gt;If you keep adding LEDs to one data line, you will usually see lower frame rate, visible skew, voltage drop, color shift, and occasional flicker before you run out of raw MCU compute. If you only switch to a faster board without splitting outputs, redesigning power, or defining sync boundaries, the 800 kHz one-wire protocol and real installation wiring will still dominate the result.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What "high-density addressable LED controller" means here:&lt;/strong&gt; one or more ESP32/WLED nodes driving hundreds to thousands of WS2812, SK6812, or similar one-wire addressable LEDs. It is not just a strip-light hobby setup; it is a small edge-control system with real-time output, power, networking, and field maintenance boundaries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The key decision line:&lt;/strong&gt; If the target is above roughly 500 LEDs, design around five decisions first — LEDs per output, number of outputs, power injection, sync method, and number of controllers. If the target approaches or exceeds 2000 LEDs, multi-output or multi-controller architecture is usually safer than forcing everything through one long data line.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F62y0bhyzj6e7eo5cg76s.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F62y0bhyzj6e7eo5cg76s.webp" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Why "more pixels" is not linear scaling
&lt;/h2&gt;

&lt;p&gt;Addressable LEDs create a misleading intuition. If 100 LEDs work, it feels like 1000 LEDs should only mean buying more strip. In practice, this is not how the system scales.&lt;/p&gt;

&lt;p&gt;One-wire LED protocols are serialized. The more pixels on one output, the longer it takes to transmit a complete frame. Even if the MCU can calculate the effect, the output line still has to send timing-sensitive data one pixel after another. WLED's multi-strip documentation reflects that reality: it recommends ESP32 for more than one output and describes four outputs as a practical sweet spot. It also gives examples such as 512 LEDs per pin × 4, 800 LEDs per pin × 4, and 1000 LEDs per pin × 4, instead of encouraging one infinitely long strip.&lt;/p&gt;

&lt;p&gt;The first rule of high-density LED control is therefore not "buy a faster chip." It is &lt;strong&gt;reduce the length of each serial output chain&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Why RMT, DMA, and Wi-Fi affect LED stability
&lt;/h2&gt;

&lt;p&gt;ESP32 projects commonly use the RMT peripheral to drive timing-sensitive WS2812-style signals. RMT was originally designed as a remote-control transceiver, but Espressif documents LED strip output as a practical use case. Espressif also notes a critical limitation: on non-ESP32-S3 chips, large LED output can rely heavily on interrupts and ping-pong buffering, so Wi-Fi or Bluetooth interrupt pressure can create timing exceptions.&lt;/p&gt;

&lt;p&gt;This is the source of many flicker problems. The color algorithm is not necessarily wrong. The LED output task is competing with network control, animation calculation, Web UI activity, MQTT traffic, or synchronization work.&lt;/p&gt;

&lt;p&gt;ESP32-S3 matters for this reason. Espressif's RMT FAQ recommends ESP32-S3 for RMT-heavy use because it supports RMT DMA, which moves more of the output workload away from the CPU interrupt path. The point is not that ESP32-S3 is always "faster." The point is: &lt;strong&gt;when LED output competes with Wi-Fi, Bluetooth, audio, or sync tasks, DMA and resource separation matter more than peak clock speed.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The five architecture decisions in a large WLED build
&lt;/h2&gt;

&lt;h3&gt;
  
  
  3.1 LEDs per output set the serial refresh ceiling
&lt;/h3&gt;

&lt;p&gt;Every WS2812/SK6812 output is a serial chain. More pixels per output means lower maximum frame rate and more visible delay in fast effects. If the installation is slow ambient lighting, that may be acceptable. If it is stage lighting, a pixel matrix, or music-reactive output, per-output length must be more conservative.&lt;/p&gt;

&lt;p&gt;When LEDs per output are too high, the first thing you lose is &lt;strong&gt;frame rate and dynamic consistency&lt;/strong&gt;, not static lighting capability.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.2 Output count determines how much work can be split
&lt;/h3&gt;

&lt;p&gt;WLED supports multiple outputs and lets users configure LED type, GPIO, length, and color order at runtime. For ESP32 builds, multiple outputs are not just a wiring convenience. They split one long serial queue into several shorter chains.&lt;/p&gt;

&lt;p&gt;More outputs still have a cost. They increase configuration, power, wiring, sync, and troubleshooting complexity. WLED's own guidance makes four outputs a sensible starting point for many single-controller builds.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.3 RMT or DMA decides whether output timing is fragile
&lt;/h3&gt;

&lt;p&gt;Classic ESP32 can run many WLED installations, but under high LED count, active Wi-Fi, heavy sync traffic, or audio-reactive effects, interrupt latency can become visible. ESP32-S3 RMT DMA reduces that pressure, but it does not remove the need for output segmentation, power design, and memory budgeting.&lt;/p&gt;

&lt;p&gt;If the installation needs both high-density LED output and real-time Wi-Fi control or audio reaction, choosing ESP32-S3 or splitting the load across nodes is usually safer than squeezing a classic ESP32 harder.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.4 Power injection decides whether "it lights" also means "it is correct"
&lt;/h3&gt;

&lt;p&gt;Many LED problems are misdiagnosed as firmware problems. Large strips commonly show yellowing at the far end, voltage drop under full white, local flicker, weak common ground, and undersized power wiring. WLED includes an automatic brightness limiter, but current limiting does not replace correct power capacity, wire gauge, injection points, and grounding.&lt;/p&gt;

&lt;p&gt;When power design is weak, reducing brightness can make the system look stable, but it does not prove the control architecture is reliable.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.5 Multi-controller sync defines the system boundary
&lt;/h3&gt;

&lt;p&gt;As LED count rises, multiple controllers often become more realistic than forcing one controller to own the entire installation. WLED's DDP virtual LED model can attach remote WLED nodes to a controlling instance, or the system can use network-level synchronization. This is useful when the physical installation is spatially distributed, power zones are clear, and one failure should not affect the whole site.&lt;/p&gt;

&lt;p&gt;Multi-controller systems also introduce latency, sync skew, configuration drift, and recovery behavior. They work best as an intentional installation architecture, not as a patch for poor early segmentation.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Recommended architecture: split outputs before scaling controllers
&lt;/h2&gt;

&lt;p&gt;The decision path works backward from the pixel scale and effect target:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Start from &lt;strong&gt;pixel scale and effect target&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Decide &lt;strong&gt;LEDs per output&lt;/strong&gt; and &lt;strong&gt;output count&lt;/strong&gt;, and map &lt;strong&gt;power zones&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;From LEDs per output, choose the &lt;strong&gt;RMT / DMA output path&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;From power zones, design &lt;strong&gt;field wiring and injection&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Both feed into &lt;strong&gt;WLED control and sync&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Only then decide between &lt;strong&gt;single or multiple controllers&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The key point is to work backward from pixel scale and effect target, then decide LEDs per output, output count, and power zones. Only after those boundaries are clear should you decide whether one controller is enough. Starting with one development board and attaching all strips to it usually mixes timing, power, and maintenance risk into one hard-to-debug system.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Practical guidance by project scale
&lt;/h2&gt;

&lt;p&gt;These numbers are not hard limits. They are architecture signals. The higher the pixel count, the more the system should be broken into small, testable boundaries.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Under 500 LEDs:&lt;/strong&gt; get power and wiring right first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Between 500 and 2000 LEDs:&lt;/strong&gt; prioritize multi-output segmentation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Above 2000 LEDs:&lt;/strong&gt; evaluate ESP32-S3, RMT DMA, multi-controller design, and network sync early.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;At every scale:&lt;/strong&gt; power injection and field labeling are not finishing details.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A reliable installation lets each zone be powered, limited, diagnosed, and recovered on its own; whole-site sync is a coordination layer, not the only thing keeping the installation alive.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Pre-delivery checklist for high-density WLED systems
&lt;/h2&gt;

&lt;p&gt;Before handoff, validate at least these points:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Each output's LED count, GPIO, color order, and physical wiring match the configuration.&lt;/li&gt;
&lt;li&gt;Every power injection point stays within safe voltage and temperature under typical and high-brightness effects.&lt;/li&gt;
&lt;li&gt;Wi-Fi control, Web UI, MQTT, sync, or audio reaction do not cause flicker when active together.&lt;/li&gt;
&lt;li&gt;One output disconnect, one controller reboot, or a short network interruption has a clear recovery behavior.&lt;/li&gt;
&lt;li&gt;Field maintenance staff can identify every output and power zone from labels or configuration records.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A high-density LED installation is reliable only when it remains explainable under network load, high brightness, partial power loss, and maintenance handoff. First-light success is not enough.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. When ESP32 + WLED should not be forced into the whole job
&lt;/h2&gt;

&lt;p&gt;ESP32 + WLED is excellent for small and medium decorative lighting, home automation, cabinets, local ambient lighting, and maintainable multi-zone installations. But some cases should not be forced through one ESP32 + WLED controller:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Large stage or video-wall systems that require strict frame synchronization.&lt;/li&gt;
&lt;li&gt;Very high pixel counts with high refresh-rate effects.&lt;/li&gt;
&lt;li&gt;Industrial installations that require long-distance noise immunity and centralized operations.&lt;/li&gt;
&lt;li&gt;Systems that need wired networking, redundant control, or strict fault isolation.&lt;/li&gt;
&lt;li&gt;Projects where maintenance teams cannot work from GPIO, zone, and power-injection documentation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those systems may be better served by dedicated LED controllers, Art-Net/sACN infrastructure, Ethernet-distributed nodes, or WLED as a local zone controller rather than the whole-site master.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Conclusion: design boundaries before choosing the board
&lt;/h2&gt;

&lt;p&gt;ESP32 + WLED is valuable because it is fast to deploy, mature, configurable, and practical for real spaces. In high-density projects, however, the decisive question is not "can ESP32 light this many pixels?" The decisive question is &lt;strong&gt;whether output, power, timing, and sync have been separated into testable boundaries.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the lighting system must run for a long time and be maintained by someone else, it is not just a strip-light project. It is a small edge-control system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://kno.wled.ge/advanced/multi-strip/" rel="noopener noreferrer"&gt;WLED Multi-strip Support&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://kno.wled.ge/features/settings/" rel="noopener noreferrer"&gt;WLED Settings: LED outputs and brightness limiter&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://kno.wled.ge/interfaces/udp-realtime/" rel="noopener noreferrer"&gt;WLED Virtual LEDs via DDP&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.espressif.com/projects/esp-faq/en/latest/software-framework/peripherals/rmt.html" rel="noopener noreferrer"&gt;Espressif ESP-FAQ: Remote Control Transceiver (RMT)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.espressif.com/projects/esp-idf/en/latest/esp32/api-reference/peripherals/rmt.html" rel="noopener noreferrer"&gt;ESP-IDF Programming Guide: RMT peripheral&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>esp32</category>
      <category>iot</category>
      <category>esphome</category>
    </item>
    <item>
      <title>ESP32 Energy Metering with HLW8032, BL0942, and ESPHome</title>
      <dc:creator>ZedIoT</dc:creator>
      <pubDate>Thu, 27 Aug 2026 12:40:00 +0000</pubDate>
      <link>https://dev.to/zediot/esp32-energy-metering-with-hlw8032-bl0942-and-esphome-15dm</link>
      <guid>https://dev.to/zediot/esp32-energy-metering-with-hlw8032-bl0942-and-esphome-15dm</guid>
      <description>&lt;h1&gt;
  
  
  ESP32 Energy Metering with HLW8032, BL0942, and ESPHome
&lt;/h1&gt;

&lt;p&gt;ESP32 energy metering with HLW8032, BL0942, and ESPHome is not just about reading voltage, current, power, and energy. This article explains how to design the UART boundary, reporting cadence, calibration, entity model, and diagnostics as one stable data path.&lt;/p&gt;

&lt;p&gt;Many ESP32 energy metering projects start well. You connect an HLW8032 or BL0942 module, enable the matching ESPHome component, and Home Assistant quickly shows voltage, current, power, and energy. But reading values is not the same as building an energy metering node that can run reliably over time.&lt;/p&gt;

&lt;p&gt;The core conclusion is this: the hard part of ESP32 energy metering is not whether HLW8032 or BL0942 can be read. The hard part is designing the metering chip, UART, Wi-Fi behavior, ESPHome entities, calibration, and diagnostics as one stable data path. If the project focuses only on sensor YAML, it can later fail on transient loads, serial conflicts, unstable sampling, Wi-Fi reconnects, Home Assistant database growth, and calibration drift.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;In this article, an ESP32 energy metering node means an edge device where ESP32 reads voltage, current, power, and energy from a metering chip such as HLW8032 or BL0942, then exposes those values through ESPHome to Home Assistant or another upper-layer platform. It is suitable for device energy monitoring, trend observation, and low-risk operational diagnostics. It should not be treated as a billing-grade meter or an electrical protection device.&lt;/p&gt;

&lt;p&gt;If the goal is to monitor the energy behavior of one appliance, one small circuit, or one commercial device inside Home Assistant, ESP32 + ESPHome + HLW8032/BL0942 is a fast and low-cost path. If the goal is billing, electrical protection, high-accuracy compliance measurement, or safety interlocking, use certified meters, protection devices, or industrial acquisition hardware instead of stretching an ESPHome node beyond its boundary.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa660l2tks2ifhosptcyi.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa660l2tks2ifhosptcyi.webp" alt=" " width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Why energy metering is more fragile than ordinary sensing
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1.1 Metering data is not like temperature or humidity data
&lt;/h3&gt;

&lt;p&gt;Energy metering often looks like a normal sensor integration, but the data behaves differently. Temperature and humidity usually change slowly. Energy metering has to deal with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;transient startup and shutdown behavior&lt;/li&gt;
&lt;li&gt;switching supplies, compressors, motors, and heaters&lt;/li&gt;
&lt;li&gt;relationships between voltage, current, power, power factor, and accumulated energy&lt;/li&gt;
&lt;li&gt;tradeoffs between sampling, filtering, calibration, and reporting frequency&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is why the first question should not be only whether ESPHome has a component for the chip. The better question is whether the data path produces stable values, whether abnormal values are diagnosable, and whether the upper platform can store and use the data over time.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.2 HLW8032 and BL0942 are metering front ends, not complete product architectures
&lt;/h3&gt;

&lt;p&gt;HLW8032 and BL0942 typically provide metering data over UART. ESPHome has official components for these chips, which makes integration much easier. But a component does not automatically solve the product architecture.&lt;/p&gt;

&lt;p&gt;A complete node still needs answers for questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is UART ownership fixed, or can it collide with logging, debugging, or other peripherals?&lt;/li&gt;
&lt;li&gt;Is the update interval aligned with the load behavior?&lt;/li&gt;
&lt;li&gt;Where do calibration values come from, and can they be checked in the field?&lt;/li&gt;
&lt;li&gt;What happens when Wi-Fi reconnects or Home Assistant is unavailable?&lt;/li&gt;
&lt;li&gt;How should accumulated energy, instant power, and abnormal state be modeled separately?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If these questions are not answered early, a working reading can hide long-term reliability risk.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. A more reliable ESP32 energy metering stack
&lt;/h2&gt;

&lt;p&gt;It helps to think in five layers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Physical wiring and electrical safety&lt;/strong&gt; — isolation, grounding, and touch protection.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metering front end&lt;/strong&gt; — HLW8032/BL0942 sampling and UART transport.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ESPHome integration&lt;/strong&gt; — component, update interval, and entity mapping.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Calibration and validation&lt;/strong&gt; — known-load verification and calibration records.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Diagnostics and operations&lt;/strong&gt; — communication state, abnormal-data guards, and retention.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The point is simple: an energy metering node is not just a chip integration. It is a trustable data path from sampling to operations. When any layer takes on the wrong job, troubleshooting becomes harder later.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Five common mistakes
&lt;/h2&gt;

&lt;h3&gt;
  
  
  3.1 Leaving the UART boundary flexible for too long
&lt;/h3&gt;

&lt;p&gt;HLW8032 and BL0942 both depend on a serial communication path. ESP32 has more UART flexibility than ESP8266, but projects still fail when serial resources are treated casually:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;debug logging and the metering chip share a serial path&lt;/li&gt;
&lt;li&gt;RS485, a display, or another serial peripheral is added later&lt;/li&gt;
&lt;li&gt;boot logs, level shifting, and wiring order are not constrained for field use&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A more reliable design fixes UART ownership from the first hardware and YAML version. Keep the metering chip's pins, baud rate, wiring, and debug strategy explicit. Energy metering nodes are not good places for loose field rewiring, because occasional serial noise can turn into permanent uncertainty.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.2 Reporting as fast as possible
&lt;/h3&gt;

&lt;p&gt;Energy metering is not always better when it is faster. Excessive reporting creates three problems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;higher Wi-Fi and ESPHome API load&lt;/li&gt;
&lt;li&gt;larger Home Assistant recorder storage&lt;/li&gt;
&lt;li&gt;more false interpretation of motor starts, relay switching, or power supply transients&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A better design separates use cases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;instant power can update more often, but should avoid meaningless jitter&lt;/li&gt;
&lt;li&gt;accumulated energy can update less frequently&lt;/li&gt;
&lt;li&gt;anomaly detection should use duration, thresholds, and device state, not one spike&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;If the node is used to tell whether equipment is running, whether energy behavior is abnormal, or whether a device is in standby, stable and explainable reporting is more valuable than a refresh rate that only looks real-time.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  3.3 Treating calibration as a one-time YAML value
&lt;/h3&gt;

&lt;p&gt;Default module readings are usually only a starting point. Real calibration depends on shunts, current transformers, module batches, load type, and installation.&lt;/p&gt;

&lt;p&gt;In practice, a better workflow is to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;verify with a known load&lt;/li&gt;
&lt;li&gt;calibrate voltage, current, power, and energy intentionally&lt;/li&gt;
&lt;li&gt;record the date, load condition, and configuration version&lt;/li&gt;
&lt;li&gt;avoid reusing one coefficient set across different hardware batches without checking&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without calibration notes, a later "8% high power reading" is hard to interpret. It could be hardware drift, configuration error, or a real load change.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.4 Exposing too many Home Assistant entities
&lt;/h3&gt;

&lt;p&gt;It is tempting to expose every available field. That looks rich at first, but it creates long-term cost:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;users see too many unstable or hard-to-explain entities&lt;/li&gt;
&lt;li&gt;database retention and automation logic become harder to maintain&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A cleaner model separates entities into three groups:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;core entities:&lt;/strong&gt; voltage, current, power, accumulated energy&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;diagnostic entities:&lt;/strong&gt; communication state, last update time, error count, node RSSI&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;business entities:&lt;/strong&gt; equipment running state, standby detection, energy band, anomaly flag&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Raw readings should support operational judgment, not dump hardware detail into the upper layer.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.5 Not designing for offline and abnormal data
&lt;/h3&gt;

&lt;p&gt;Once an energy metering node is installed near real equipment, it will face weak Wi-Fi, power loss, load shutdown, chip communication failure, and sudden readings. Without an abnormal-data strategy, Home Assistant often shows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a device that looks like it suddenly used too much power&lt;/li&gt;
&lt;li&gt;accumulated energy jumps&lt;/li&gt;
&lt;li&gt;automation triggered by one transient value&lt;/li&gt;
&lt;li&gt;unclear responsibility between equipment failure and node failure&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Better strategies include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;separate diagnostic entities for communication status and last update time&lt;/li&gt;
&lt;li&gt;guards or labels for impossible values&lt;/li&gt;
&lt;li&gt;duration thresholds for anomaly detection&lt;/li&gt;
&lt;li&gt;different states for "equipment has no load" and "metering node unavailable"&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. Choosing between HLW8032 and BL0942
&lt;/h2&gt;

&lt;p&gt;For most ESPHome projects, the choice should depend less on the chip name and more on module availability, wiring, documentation quality, and accuracy expectations.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The HLW8032 versus BL0942 choice is rarely the largest factor in whether an ESPHome energy monitor succeeds. For most projects, reliability depends more on module quality, electrical safety, calibration workflow, UART ownership, and reporting strategy.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fun07dllrzg0ifphv4orl.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fun07dllrzg0ifphv4orl.webp" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  5. A production-minded ESPHome design direction
&lt;/h2&gt;

&lt;p&gt;The example below is not a complete drop-in configuration. It shows the design direction:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;uart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;metering_uart&lt;/span&gt;
  &lt;span class="na"&gt;tx_pin&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;GPIO17&lt;/span&gt;
  &lt;span class="na"&gt;rx_pin&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;GPIO16&lt;/span&gt;
  &lt;span class="na"&gt;baud_rate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;4800&lt;/span&gt;

&lt;span class="na"&gt;sensor&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;platform&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;hlw8032&lt;/span&gt;
    &lt;span class="na"&gt;uart_id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;metering_uart&lt;/span&gt;
    &lt;span class="na"&gt;voltage&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Meter&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Voltage"&lt;/span&gt;
    &lt;span class="na"&gt;current&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Meter&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Current"&lt;/span&gt;
    &lt;span class="na"&gt;power&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Meter&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Power"&lt;/span&gt;
    &lt;span class="na"&gt;energy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Meter&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Energy"&lt;/span&gt;
    &lt;span class="na"&gt;update_interval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;10s&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a real deployment, add:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;calibration parameters and calibration notes&lt;/li&gt;
&lt;li&gt;filtering or value guards for readings&lt;/li&gt;
&lt;li&gt;diagnostic entities such as Wi-Fi signal, uptime, and restart reason&lt;/li&gt;
&lt;li&gt;Home Assistant recorder retention and exclusion strategy&lt;/li&gt;
&lt;li&gt;enclosure, isolation, safe wiring, and touch-protection rules&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The configuration is only the entry point. Reliability comes from constraints across the full data path.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. When not to use ESP32 + ESPHome for energy metering
&lt;/h2&gt;

&lt;p&gt;Do not stretch this stack into these cases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Billing:&lt;/strong&gt; billing needs compliance, sealing, metering class, and an audit trail.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Electrical protection:&lt;/strong&gt; overcurrent, leakage, and short-circuit protection belong to dedicated protection devices.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hard real-time control:&lt;/strong&gt; protection and critical interlocks should not depend on Wi-Fi and Home Assistant.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;High-noise industrial cabinets:&lt;/strong&gt; poor isolation, grounding, and power quality can overwhelm a lightweight ESP32 node.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Many circuits with high refresh rates:&lt;/strong&gt; multi-circuit acquisition is often better handled by professional meters or gateways.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;ESP32 energy metering nodes are strong for visualization, trends, auxiliary diagnostics, and low-risk automation. They are not the right boundary for billing, safety protection, or hard real-time control. Stating that boundary makes the architecture more credible, not weaker.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  7. Conclusion
&lt;/h2&gt;

&lt;p&gt;ESP32, HLW8032, BL0942, and ESPHome can quickly produce a node that shows energy data in Home Assistant. But the engineering value is not that the numbers appear on a dashboard. The real value is whether the data path remains stable over time, whether readings are explainable, whether abnormal states are diagnosable, and whether the upper platform is not overloaded with raw entities.&lt;/p&gt;

&lt;p&gt;If the goal is to understand whether one device is running or whether its energy trend looks abnormal, ESP32 + ESPHome is a strong option. If the goal is billing, electrical protection, or hard real-time control, the ESP32 node should stay in its proper role: a lightweight edge monitoring node, not the final authority in the electrical system.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;ESPHome HLW8032 Sensor&lt;/li&gt;
&lt;li&gt;ESPHome BL0942 Sensor&lt;/li&gt;
&lt;li&gt;ESPHome UART Bus&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>esp32</category>
      <category>iot</category>
      <category>esphome</category>
    </item>
    <item>
      <title>ESP32-C3 vs S3 vs C6: Firmware, TinyML, Matter, and Production Tradeoffs</title>
      <dc:creator>ZedIoT</dc:creator>
      <pubDate>Tue, 25 Aug 2026 12:40:00 +0000</pubDate>
      <link>https://dev.to/zediot/esp32-c3-vs-s3-vs-c6-firmware-tinyml-matter-and-production-tradeoffs-3o0c</link>
      <guid>https://dev.to/zediot/esp32-c3-vs-s3-vs-c6-firmware-tinyml-matter-and-production-tradeoffs-3o0c</guid>
      <description>&lt;h1&gt;
  
  
  ESP32-C3 vs S3 vs C6: Firmware, TinyML, Matter, and Production Tradeoffs
&lt;/h1&gt;

&lt;p&gt;Compare ESP32-C3, S3, and C6 for CPU, memory, USB, wireless, TinyML, Matter, OTA, and production validation before choosing a custom firmware platform.&lt;/p&gt;

&lt;p&gt;For a connected sensor or compact controller, ESP32-C3 is usually the lowest-risk starting point. If the device must combine a display, camera, audio, USB, or local inference, validate ESP32-S3 first. If the roadmap explicitly requires on-chip 802.15.4, particularly Matter over Thread or Zigbee, validate ESP32-C6 first. This is not a ranking by age or headline clock speed. Each chip defines a different system boundary.&lt;/p&gt;

&lt;p&gt;The decision should be made against peak memory, concurrent peripherals, radio requirements, OTA rollback, and the maintenance path after security features are enabled. A successful prototype only proves that the happy path ran once. A production selection needs measurable margin under the worst workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short answer
&lt;/h2&gt;

&lt;p&gt;Two distinctions prevent expensive mistakes. C3 and C6 include USB Serial/JTAG, but that is not the general USB OTG capability offered by S3. Also, Matter does not automatically require C6: Matter can run over Wi-Fi. C6 becomes the clear route when Thread or another on-chip 802.15.4 use case is part of the product contract.&lt;/p&gt;

&lt;h2&gt;
  
  
  Freeze the workload before comparing chips
&lt;/h2&gt;

&lt;p&gt;Turn the product brief into a measurable workload sheet:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Connectivity:&lt;/strong&gt; simultaneous Wi-Fi/BLE, TLS sessions, MQTT reconnects, local discovery, or Thread/Zigbee.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data path:&lt;/strong&gt; sensor rate, audio frames, image size, ring buffers, offline queue, and retained logs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interaction:&lt;/strong&gt; display refresh, touch, USB, camera, wake word, and the maximum user-visible response time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maintenance:&lt;/strong&gt; A/B OTA, rollback, crash capture, field diagnostics, Secure Boot, Flash Encryption, and key rotation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;"MQTT works" is not a memory test. Peak pressure may occur when TLS reconnect, OTA download, log writes, and sensor acquisition overlap. A system can report adequate total free heap yet still fail a large contiguous allocation. Test the combined condition instead of estimating each subsystem in isolation.&lt;/p&gt;

&lt;p&gt;The selection flow ends up looking like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Freeze the workload.&lt;/li&gt;
&lt;li&gt;Is on-chip 802.15.4 required? → Yes: validate C6 first.&lt;/li&gt;
&lt;li&gt;Does it need USB OTG, display, camera, audio, or heavier inference? → Yes: validate S3 first.&lt;/li&gt;
&lt;li&gt;Otherwise → start validation with C3.&lt;/li&gt;
&lt;li&gt;Run worst-case production tests. If resource, RF, OTA, and security margins pass, freeze the chip and module; otherwise, re-examine the workload.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Convert specifications into firmware consequences
&lt;/h2&gt;

&lt;p&gt;Espressif documents ESP32-C3 as a single-core RISC-V device up to 160 MHz with 400 KB of on-chip SRAM, 2.4 GHz Wi-Fi 4, and Bluetooth 5 LE. ESP32-S3 has two Xtensa LX7 cores up to 240 MHz, 512 KB of on-chip SRAM, vector instructions, LCD/camera support, and USB OTG. ESP32-C6 differentiates itself with Wi-Fi 6, BLE, IEEE 802.15.4, and high-performance plus low-power RISC-V cores. Package, flash, PSRAM, and pin availability still depend on the selected SoC revision and module.&lt;/p&gt;

&lt;p&gt;Those specifications change architecture decisions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;C3's&lt;/strong&gt; single core is sufficient for many nodes, but the network stack, application tasks, and interrupt service compete more directly. Task priorities, non-blocking drivers, and reconnect-time latency must be deliberate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;S3's&lt;/strong&gt; second core and vector support create room for richer edge workloads, but do not remove memory and bandwidth limits. A framebuffer, camera DMA, audio buffers, and PSRAM traffic can contend at the same time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;C6&lt;/strong&gt; is primarily a protocol-roadmap decision, not a replacement for S3 multimedia. Thread/Zigbee and Wi-Fi/BLE coexistence bring RF scheduling, certification, and stack-resource work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use the official ESP32-C3 datasheet, ESP32-S3 datasheet, and ESP32-C6 datasheet as the baseline. Record the chip revision, ESP-IDF version, module, and differences between the development kit and production PCB with the decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build a firmware budget, not a flash-size guess
&lt;/h2&gt;

&lt;p&gt;Create a resource budget before schematic freeze and make CI report the same measures on every build.&lt;/p&gt;

&lt;p&gt;PSRAM is not unlimited memory. It depends on the S3 module/variant and is commonly useful for large buffers, framebuffers, or model data. It should not blindly absorb every real-time allocation. Espressif's LCD documentation notes that framebuffers, CPU activity, and EDMA can share PSRAM bandwidth and become starved. Measure display, network, and local processing concurrently.&lt;/p&gt;

&lt;p&gt;Tie that budget to a repeatable peak-load scenario. A display device may look comfortable on a static page, then encounter DNS and TLS reconnect, OTA metadata download, screen refresh, sensor acquisition, and offline-queue writes at once. Record minimum free heap, largest free block, task-stack high-water marks, watchdog events, dropped frames, and business-response latency for a fixed workload. Replay it after ESP-IDF, TLS, model, or partition changes. If C3 retains stable margin, moving to S3 does not automatically improve the product; if the failure is a non-separable memory peak or scheduling conflict, isolated micro-optimisations may only defer the architecture decision.&lt;/p&gt;

&lt;p&gt;Freeze a chip together with its module, partition table, ESP-IDF baseline, and security configuration. Modules based on the same SoC can differ in flash, PSRAM, antenna arrangement, and usable pins, while a development board may hide power or programming constraints with external components. The design record should therefore name the module, substitution rules, strapping pins, antenna clearance, peak supply assumptions, and download/JTAG path. This makes a module substitution or SDK upgrade trigger the right validation instead of being treated as an equivalent "same ESP32" change.&lt;/p&gt;

&lt;h2&gt;
  
  
  TinyML: S3 is a natural candidate, not an automatic pass
&lt;/h2&gt;

&lt;p&gt;For wake words, vibration classification, small vision features, or compact detection models, S3's dual cores, vector instructions, and optional PSRAM often make it the practical first candidate. ESP-DL also treats quantisation as central on memory-constrained devices. But loading a model is not the acceptance criterion. Freeze and measure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;input shape, supported operators, INT8/INT16 method, and representative calibration data;&lt;/li&gt;
&lt;li&gt;peak tensor arena, weights, preprocessing, and business buffers at the same time;&lt;/li&gt;
&lt;li&gt;end-to-end latency including acquisition, preprocessing, inference, postprocessing, and transmission;&lt;/li&gt;
&lt;li&gt;p95/p99 latency and watchdog behaviour while Wi-Fi, display, or audio is active;&lt;/li&gt;
&lt;li&gt;accuracy, false positives, false negatives, and an explicit uncertain/manual-review path.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;C3 can execute sufficiently small models, so it should not be excluded by name. Conversely, an unsupported operator set, large image pipeline, or Linux-class runtime can exceed S3's sensible boundary. The right answer may be an MCU plus a dedicated accelerator or application processor.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fost8sn2blkw31lungxoy.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fost8sn2blkw31lungxoy.webp" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Matter and Thread: identify the network bearer first
&lt;/h2&gt;

&lt;p&gt;"Support Matter" is not yet a complete requirement. Is it Matter over Wi-Fi or Matter over Thread? Does the product also need Zigbee? Is the device an end device, router, bridge, or part of a border-router system? How do commissioning, local control, and cloud control degrade independently?&lt;/p&gt;

&lt;p&gt;For on-chip Thread or Zigbee, C6's IEEE 802.15.4 radio is a direct advantage. For Matter over Wi-Fi, C3, S3, and C6 can all be candidates depending on memory, peripherals, and the certification plan. Protocol availability does not equal a certifiable product: antenna design, RF coexistence, credentials, device attestation, commissioning UX, and stack-version control remain production work.&lt;/p&gt;

&lt;p&gt;A gateway or bridge can also accumulate too many roles. Combining 802.15.4, Wi-Fi backhaul, model translation, OTA, and local rules on one MCU expands the fault domain. A radio coprocessor separated from the primary controller can sometimes produce a cleaner upgrade and certification boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prove the selection with a failure matrix
&lt;/h2&gt;

&lt;p&gt;Before freezing the chip, run combinations that represent field failure rather than a feature demo.&lt;/p&gt;

&lt;p&gt;Do not postpone the security lifecycle. Secure Boot, Flash Encryption, and eFuse decisions can change JTAG, download, and repair access. Rehearse key injection, recovery, and RMA on a pre-production batch.&lt;/p&gt;

&lt;p&gt;The failure matrix must also separate a silicon limit from an integration defect. A control timeout during weak-signal reconnect could indicate CPU contention, but it could also come from a driver holding a lock, synchronous logging, or an incorrect backoff policy. Replacing C3 with S3 may hide the symptom without fixing the failure mode. Associate each result with reset reason, heap low-water mark, stack high-water mark, state-machine timing, and OTA rollback reason; upgrade the chip only after the evidence shows that the required workload still crosses the resource or peripheral boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  When none of these chips is the right boundary
&lt;/h2&gt;

&lt;p&gt;High-resolution multi-stream video, complex Linux applications, containers, browser-class UI, large-model inference, or substantial local storage may already be outside a sensible MCU boundary. Consider a Linux SoC, an MCU/MPU split, or a dedicated accelerator. Preserving a one-chip BOM by sacrificing observability, rollback, and performance margin usually moves cost into field maintenance.&lt;/p&gt;

&lt;p&gt;Do not upgrade an established C3 product merely because S3 or C6 exposes more features. A mature C3 design may already have a stable BSP, fixture, certification, and supply chain. Migration reopens drivers, RF, power, factory testing, and OTA risk. It is justified when a measured new workload crosses the current boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recommendation
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Choose &lt;strong&gt;ESP32-C3&lt;/strong&gt; when the product is a focused Wi-Fi/BLE sensing or control node and worst-case heap, latency, and OTA have passed.&lt;/li&gt;
&lt;li&gt;Choose &lt;strong&gt;ESP32-S3&lt;/strong&gt; when USB OTG, display, camera, audio, or TinyML is the core workload and PSRAM/bandwidth/real-time margin is demonstrated.&lt;/li&gt;
&lt;li&gt;Choose &lt;strong&gt;ESP32-C6&lt;/strong&gt; when Thread, Zigbee, or a Wi-Fi 6 roadmap is explicit and coexistence, certification, and OTA resource costs are in the plan.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the requirements still cannot be converted into a module, partition table, driver boundary, and validation matrix, another comparison table will not close the gap.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This guide uses public Espressif documentation and does not include a controlled, cross-chip benchmark on identical boards, firmware, and lab conditions. It therefore makes no universal promise about power, BOM, RF, TinyML latency, or certification. Re-test all numbers on the selected module, PCB, ESP-IDF release, and production configuration.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  References
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;ESP32-C3 Datasheet&lt;/li&gt;
&lt;li&gt;ESP32-S3 Datasheet&lt;/li&gt;
&lt;li&gt;ESP32-C6 Datasheet&lt;/li&gt;
&lt;li&gt;ESP-IDF Chip Series Comparison&lt;/li&gt;
&lt;li&gt;ESP32-C3 USB Serial/JTAG&lt;/li&gt;
&lt;li&gt;ESP32-S3 LCD and PSRAM Bandwidth Notes&lt;/li&gt;
&lt;li&gt;ESP-DL Introduction&lt;/li&gt;
&lt;li&gt;ESP-DL Quantisation Specification&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>esp32</category>
      <category>iot</category>
      <category>tinyml</category>
    </item>
    <item>
      <title>Edge AI Device OTA: Staged Rollouts, Rollbacks, and Remote Recovery</title>
      <dc:creator>ZedIoT</dc:creator>
      <pubDate>Tue, 25 Aug 2026 10:06:15 +0000</pubDate>
      <link>https://dev.to/zediot/edge-ai-device-ota-staged-rollouts-rollbacks-and-remote-recovery-38bo</link>
      <guid>https://dev.to/zediot/edge-ai-device-ota-staged-rollouts-rollbacks-and-remote-recovery-38bo</guid>
      <description>&lt;h1&gt;
  
  
  Edge AI Device OTA: Staged Rollouts, Rollbacks, and Remote Recovery
&lt;/h1&gt;

&lt;p&gt;A practical guide to designing over-the-air updates for edge AI fleets — why firmware, model, and config must ship separately, and how to build rollback and recovery paths from day one.&lt;/p&gt;

&lt;p&gt;The hard part of Edge AI OTA is not pushing a new package. It is designing staged rollout, rollback, and remote recovery for devices whose firmware, model, and configuration all change independently.&lt;/p&gt;

&lt;p&gt;When teams talk about Edge AI deployment, they usually start with the model: can it run on-device, how fast is inference, and what is the power profile. But once devices are deployed in volume, the first serious failure often comes from the release path itself. One device gets the new firmware but not the new model. Another applies a config change before the model file finishes downloading. A third reboots into a bad state and loses the only recovery channel you had.&lt;/p&gt;

&lt;p&gt;The core conclusion is simple: Edge AI OTA should not be treated as "shipping one new package." It should be treated as a layered operations system that releases firmware, model, and configuration separately, validates health during staged rollout, and can roll back deterministically. If you keep shipping them as one bound update, fleet scale will expose recoverability problems before it exposes product problems.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What Edge AI OTA means here:&lt;/strong&gt; the coordinated remote release of firmware, model artifacts, configuration, dependencies, and health rules. It is more than package delivery or firmware flashing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When it's required:&lt;/strong&gt; if an edge AI device will run continuously, receive model updates, or operate in places where onsite support is expensive, OTA must include staged rollout, automatic rollback, and remote recovery from day one. Without those three layers, every new release increases operational risk.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  1. Why Edge AI OTA breaks differently from standard IoT OTA
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1.1 In normal IoT, a failed update usually breaks a feature; in Edge AI, it can break the whole runtime chain
&lt;/h3&gt;

&lt;p&gt;For a simple telemetry or control device, a failed update often means the device stays on the old version or one function becomes unavailable. Edge AI devices are different because at least three classes of change evolve together:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Firmware or system runtime changes&lt;/li&gt;
&lt;li&gt;Model artifact changes&lt;/li&gt;
&lt;li&gt;Configuration changes such as thresholds, feature flags, and resource mappings&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those layers are not naturally synchronized. If the platform does not model their dependencies explicitly, the fleet quickly starts to exhibit failure patterns like these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a new model arrives, but the old firmware cannot support its preprocessing path&lt;/li&gt;
&lt;li&gt;firmware upgrades successfully, but the configuration never switches, so inference services fail to start&lt;/li&gt;
&lt;li&gt;configuration activates first, and the device points to a model that is not fully downloaded&lt;/li&gt;
&lt;li&gt;the device enters a reboot loop, while the platform only reports that the package was delivered&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Once the target is an ESP32 camera, an RK3566 vision box, a gateway with an NPU, or a field industrial terminal, the update process is no longer a simple binary replacement. It becomes dependency management for a live runtime.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.2 From ESP32 to RK3566, release complexity does not scale linearly
&lt;/h3&gt;

&lt;p&gt;Many teams try to manage MCU-class devices and Linux edge boxes with the same mental model: OTA means replacing the software package. That may survive a PoC, but it does not survive fleet operations.&lt;/p&gt;

&lt;p&gt;The reason is that the two device classes have very different boundaries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ESP32 devices usually have tighter memory, more rigid partitions, smaller update artifacts, and weaker observability&lt;/li&gt;
&lt;li&gt;RK3566 class Linux devices can carry much larger models and dependencies, but they introduce service orchestration, disk space management, driver compatibility, and multi-process runtime issues&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the release platform does not adapt rollout policy to device capability and instead insists on one generic OTA path for every node, the first thing to collapse is not release success rate. It is recovery quality and troubleshooting speed.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. What a production-safe Edge AI OTA system must separate
&lt;/h2&gt;

&lt;h3&gt;
  
  
  2.1 Do not bind firmware, model, and configuration into one version number
&lt;/h3&gt;

&lt;p&gt;This is the first habit worth fixing in any Edge AI release pipeline. A single bundled version may look simpler, but it makes root cause analysis and rollback significantly worse.&lt;/p&gt;

&lt;p&gt;A safer structure tracks at least three version planes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Firmware Version:&lt;/strong&gt; drivers, acquisition stack, inference runtime, device management agent&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model Version:&lt;/strong&gt; model weights, quantized artifacts, label maps, pre/post-processing assets&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Config Version:&lt;/strong&gt; thresholds, sampling policy, upload cadence, model selection rules, feature flags&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Why this separation matters:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;firmware rollback and model rollback do not have the same cost or blast radius&lt;/li&gt;
&lt;li&gt;model swaps should not always require a firmware restart&lt;/li&gt;
&lt;li&gt;configuration mistakes usually deserve a fast logical revert, not a full firmware rollback&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;If an Edge AI platform cannot track firmware, model, and configuration independently, it will struggle to do low-risk staged rollout and will struggle even more to identify which layer actually failed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  2.2 A release system must answer one operational question first: what exactly is being released
&lt;/h3&gt;

&lt;p&gt;A production release object should make these points explicit:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;which device groups, customers, regions, or sites are targeted&lt;/li&gt;
&lt;li&gt;whether the change affects firmware, model, configuration, or a combination&lt;/li&gt;
&lt;li&gt;whether a minimum prerequisite version must already be present&lt;/li&gt;
&lt;li&gt;what success means for this release&lt;/li&gt;
&lt;li&gt;which layer should roll back first when health degrades&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without a modeled release object, staged rollout turns into "we picked a few devices to test" and rollback turns into "we pushed the old package again and hoped for the best."&lt;/p&gt;

&lt;p&gt;A useful release model connects these five parts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Release Plan
├── Target Ring        → Canary → 10% Fleet → Region/Customer → Full Rollout
├── Version Set        → Firmware Version + Model Version + Config Version
├── Health Rules       → what "success" is measured against
└── Rollback Policy    → which layer reverts first on failure
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The release plan links a target ring, a version set, health rules, and a rollback policy together so every release is explicit about what ships, to whom, and how it gets pulled back.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.3 Staged rollout is not about shipping to fewer devices first; it is about testing recovery first
&lt;/h3&gt;

&lt;p&gt;Teams often reduce staged rollout to a quantity problem: first 1%, then 10%, then full deployment. That is incomplete. In Edge AI, staged rollout has to validate three things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;whether the upgraded device starts correctly&lt;/li&gt;
&lt;li&gt;whether inference quality and resource behavior remain stable&lt;/li&gt;
&lt;li&gt;whether the platform can detect failure and recover automatically&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the staged phase only checks that the package was delivered, not whether inference, health telemetry, logs, and rollback paths all work, full rollout still carries the same operational risk.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. How to design rollout, rollback, and remote recovery from ESP32 to RK3566
&lt;/h2&gt;

&lt;h3&gt;
  
  
  3.1 ESP32 needs the smallest and most deterministic rollback path
&lt;/h3&gt;

&lt;p&gt;ESP32-class devices are defined by tighter resources, broad physical distribution, and weaker observability. For them, the most valuable OTA feature is not richness. It is survivability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recommended patterns:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;use explicit dual-partition or A/B firmware strategy&lt;/li&gt;
&lt;li&gt;keep model artifacts smaller or layered externally instead of tying every model change to firmware&lt;/li&gt;
&lt;li&gt;require boot health checks after update, such as management-agent connectivity, sensor initialization, or inference thread liveness&lt;/li&gt;
&lt;li&gt;roll back automatically within a bounded time window if those checks fail&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Patterns to avoid:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;replacing firmware, model, and configuration in one large update&lt;/li&gt;
&lt;li&gt;treating "device came online" as enough evidence of release success&lt;/li&gt;
&lt;li&gt;depending entirely on human intervention for rollback&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3.2 RK3566 needs service lifecycle separation more than it needs whole-image replacement
&lt;/h3&gt;

&lt;p&gt;RK3566-class Linux devices often run multiple services at once: camera ingestion, decoding, inference, upload, and remote management. In these systems, the most common failures happen not during the file transfer but after release, when service dependencies become misaligned.&lt;/p&gt;

&lt;p&gt;A safer strategy usually looks like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;manage system, application, and model layers separately&lt;/li&gt;
&lt;li&gt;switch models through manifests, symlinks, or service config rather than replacing the whole system every time&lt;/li&gt;
&lt;li&gt;use post-update checks for service health, disk headroom, NPU readiness, and sample inference replay&lt;/li&gt;
&lt;li&gt;prefer process-level or container-level release over whole-image replacement unless kernel or driver updates require it&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3.3 Automatic rollback must be driven by health signals, not by timeout alone
&lt;/h3&gt;

&lt;p&gt;Many OTA platforms use only one rollback trigger: the device did not come back online in time. That is not enough for Edge AI. A device may be online while inference is already broken.&lt;/p&gt;

&lt;p&gt;Better rollback signals include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;whether the model service started successfully&lt;/li&gt;
&lt;li&gt;whether inference latency exceeds a safe threshold&lt;/li&gt;
&lt;li&gt;whether memory, storage, or temperature enters an abnormal range&lt;/li&gt;
&lt;li&gt;whether critical inputs such as camera, sensor, or encoder streams disappeared&lt;/li&gt;
&lt;li&gt;whether the device still reports version and health summary consistently&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The difference is worth stating plainly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Standard IoT OTA asks whether the device came back online. Edge AI OTA asks whether the device came back online with a healthy inference path. If the platform watches only connectivity, it will misclassify many real failures as successful releases.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  3.4 Remote recovery must be designed before the outage, not after it
&lt;/h3&gt;

&lt;p&gt;At scale, the most expensive part of a bad release is often not the failure itself. It is the requirement to send people onsite.&lt;/p&gt;

&lt;p&gt;That is why Edge AI devices should always preserve a remote recovery path, for example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a minimal management agent separated from the main application stack&lt;/li&gt;
&lt;li&gt;an independent safe mode or recovery partition&lt;/li&gt;
&lt;li&gt;the ability to pause auto-updates, freeze a bad version, and return to a stable model&lt;/li&gt;
&lt;li&gt;the ability to stop rollout immediately by device group, region, or customer segment&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the team discovers during an incident that the management agent broke alongside the main workload, the failure is no longer just a release problem. It is an architecture problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. A practical rollout cadence for Edge AI fleets
&lt;/h2&gt;

&lt;h3&gt;
  
  
  4.1 The right sequence is not "build and push"; it is "validate health, then expand"
&lt;/h3&gt;

&lt;p&gt;A safer rollout rhythm usually looks like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;validate version dependencies on internal devices&lt;/li&gt;
&lt;li&gt;validate upgrade, inference, and rollback chains on a small canary ring&lt;/li&gt;
&lt;li&gt;expand by region, customer, or hardware family&lt;/li&gt;
&lt;li&gt;watch a stability window before full rollout&lt;/li&gt;
&lt;li&gt;keep freeze and rollback windows after rollout instead of deleting the previous version immediately&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The release state flow follows this path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Build Release
 → Internal Validation
   → Canary Rollout
     → Health Pass?
         ├─ Yes → Expand by Ring → Stable Window Passed? → Yes → Full Rollout
         │                                        └─ No → Auto Rollback
         └─ No → Auto Rollback
                                   Auto Rollback → Freeze Version / Investigate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every step either promotes to the next ring on a passed health check, or falls back to an auto rollback that freezes the version for investigation.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.2 When a full staged rollout system may be overkill
&lt;/h3&gt;

&lt;p&gt;Not every Edge AI project needs a complex release orchestration system on day one. A lighter path may be enough when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the fleet is small and easy to service onsite&lt;/li&gt;
&lt;li&gt;models rarely change after deployment&lt;/li&gt;
&lt;li&gt;the device does not carry critical business risk and failed updates are cheap to fix manually&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Even then, version tracking and basic rollback should remain in scope. The moment the project starts to scale, those become the first missing capabilities.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If an Edge AI system updates rarely, operates in small numbers, and remains easy to maintain onsite, a full staged rollout platform may not be the first investment to make. But once the fleet is expected to scale, rollback and remote recovery stop being optional.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  5. Conclusion: the real question is not how to push an update, but how to pull a bad release back
&lt;/h2&gt;

&lt;p&gt;For Edge AI devices, scale is determined less by the first successful deployment than by whether every later update can still be controlled safely. ESP32 and RK3566 have different runtime boundaries, but they obey the same operational rule: releases must be designed as a system that can stage, verify, roll back, and recover instead of a file transfer step.&lt;/p&gt;

&lt;p&gt;So if you are building Edge AI OTA, the highest-value investments are not the ones that make deployment slightly faster. They are the ones that make recovery predictable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;version separation:&lt;/strong&gt; track and release firmware, model, and configuration independently&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;staged validation:&lt;/strong&gt; promote only when health and recovery paths prove out&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;rollback and recovery:&lt;/strong&gt; make sure the platform can regain control after a failed release&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Only when those three layers exist does Edge AI OTA move from "can upgrade" to "can operate for the long run."&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Why should firmware, model, and configuration ship as separate versions in Edge AI OTA?
&lt;/h3&gt;

&lt;p&gt;Because they have different failure costs and rollback blast radius. Separating them lets you roll back a model without a firmware restart, or revert a bad config logically instead of doing a full firmware rollback, and makes root cause analysis much faster.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. What health signals should drive automatic rollback?
&lt;/h3&gt;

&lt;p&gt;Beyond a connectivity timeout, watch whether the model service started, whether inference latency stays within a safe threshold, whether memory/storage/temperature are in a normal range, whether critical inputs (camera, sensor, encoder) still exist, and whether the device keeps reporting version and health summaries consistently.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. How should ESP32 and RK3566 devices be released differently?
&lt;/h3&gt;

&lt;p&gt;ESP32-class devices need the smallest, most deterministic rollback path — dual-partition A/B, boot health checks, and bounded-time auto rollback. RK3566-class Linux devices need service lifecycle separation: switch models via manifests or service config, with post-update checks for service health, disk headroom, and NPU readiness.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. When is a full staged rollout system overkill?
&lt;/h3&gt;

&lt;p&gt;When the fleet is small and easy to service onsite, models rarely change after deployment, and the device carries low business risk with cheap-to-fix failures. Even then, basic version tracking and rollback should stay in place for when the project starts to scale.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>iot</category>
      <category>devops</category>
      <category>tutorial</category>
    </item>
  </channel>
</rss>
