<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: liesliy</title>
    <description>The latest articles on DEV Community by liesliy (@liesliy).</description>
    <link>https://dev.to/liesliy</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4057379%2F01a5a8ba-545c-4533-bd77-5f940e75f65c.png</url>
      <title>DEV Community: liesliy</title>
      <link>https://dev.to/liesliy</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/liesliy"/>
    <language>en</language>
    <item>
      <title>How Two GitHub Issues Shaped RDA v0.9.14: Action Units, Per-Joint Analysis, and Smarter Freeze Detection</title>
      <dc:creator>liesliy</dc:creator>
      <pubDate>Sun, 20 Sep 2026 08:05:23 +0000</pubDate>
      <link>https://dev.to/liesliy/how-two-github-issues-shaped-rda-v0914-action-units-per-joint-analysis-and-smarter-freeze-545f</link>
      <guid>https://dev.to/liesliy/how-two-github-issues-shaped-rda-v0914-action-units-per-joint-analysis-and-smarter-freeze-545f</guid>
      <description>&lt;p&gt;If you're working on robot learning, you already know the pain: a public dataset looks fine on paper, but after burning weeks of GPU time, your policy won't converge. The problem isn't the model — it's the data.&lt;/p&gt;




&lt;p&gt;That's why we built RDA (Robot Data Audit), an open-source tool for auditing LeRobot-format datasets. It runs a four-layer diagnostic pipeline: L1 Integrity Gate for hard verdicts (PASS/REVIEW/EXCLUDE), L2 Trajectory Diagnostics for behavioral observations like action spikes and idle ratios, L3 Dataset Profiling for cross-episode distribution and coverage analysis, and L4 Dataset Summary for aggregated reporting. The design philosophy is &lt;strong&gt;separating measurement from judgment&lt;/strong&gt; — L1 makes hard calls, while L2/L3 lay out the full picture so you can decide.&lt;/p&gt;

&lt;p&gt;Open source tools evolve through conversations. Not the polite kind — the ones where someone digs into your code, runs it on real data, and comes back with evidence that something is wrong.&lt;/p&gt;

&lt;p&gt;Last week, a developer filed two issues on our GitHub. One on LeRobot (huggingface/lerobot#4650), one on LIBERO (Lifelong-Robot-Learning/LIBERO#148). Both were technically precise, both came with his own reproducible analysis, and both led directly to changes in RDA v0.9.14.&lt;/p&gt;

&lt;p&gt;Here's what happened and what we built.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;The Action Unit Problem&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The developer's first issue pointed out that we had the action unit wrong for the svla_so101_pickplace dataset. We said the values were in "raw STS3215 encoder steps" (0-4096 range). He checked the LeRobot driver layer and found it was already converting raw steps to physical degrees — so the [-100, 100] range was in degrees, not steps.&lt;/p&gt;

&lt;p&gt;We went back and verified. He was right. The driver handles the conversion.&lt;/p&gt;

&lt;p&gt;This seemed like a simple documentation fix at first. But it exposed a deeper problem: if you don't know what unit your action data is in, you can't meaningfully compare action discontinuity across datasets. One dataset uses degrees, another uses radians, another uses normalized [-1, 1], another uses raw encoder steps. The same spike looks completely different depending on the unit.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;Proving Scale Invariance&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The developer's intuition was that MAD-based z-scores should be invariant under scale transforms. This is a powerful claim — it would mean the spike detection results don't depend on what unit you measure in, as long as the transforms are linear.&lt;/p&gt;

&lt;p&gt;We ran the test properly. We took the same svla action data and applied four different scale transforms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Degrees (original)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Radians (degrees × π/180)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Raw steps (rescaled to 0-4096 range)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Normalized (rescaled to [-1, 1])&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The result: spike_count was 260 in all four cases. Median idle ratio identical across all four.&lt;/p&gt;

&lt;p&gt;This confirmed that MAD-based z-score detection is inherently scale-invariant. Whether your actions are in degrees or radians or encoder steps, the spike detection produces the same result. We formalized this as invariant INV-011 and added a guardian test (test_inv011_scale_invariance) to ensure no future code change breaks this property.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;Per-Joint Spike Breakdown: Now at the Top Level&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The developer didn't just point out the unit issue — he also shared his own per-joint action discontinuity analysis for the svla dataset. He computed spike counts for each joint individually: 347, 1080, 1279, 926, 497, 26.&lt;/p&gt;

&lt;p&gt;This was exactly the kind of granular insight that matters for dataset quality analysis. The problem was that RDA already computed per-joint data, but it was buried inside the details layer — you had to dig for it.&lt;/p&gt;

&lt;p&gt;In v0.9.14, we moved by_joint to the top-level measurement output. It's now sorted by spike_count descending, with a configurable top_k_joints parameter (default 0 means show all). The old details.by_joint is preserved for backward compatibility.&lt;/p&gt;

&lt;p&gt;Running v0.9.14 on the full svla dataset gives per-joint totals of 351, 1075, 1179, 1227, 717, 25 — very close to the developer's independent calculation. The small deltas come from threshold tuning on our side, but the structure matches. That kind of independent reproduction and cross-validation is exactly what makes open source work.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;Video Freeze: State Cross-Validation&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The second issue (LIBERO#148) was about video freeze detection. The developer made two suggestions that both ended up in v0.9.14.&lt;/p&gt;

&lt;p&gt;The first and more impactful one: add a --freeze-motion-source flag that lets the freeze detector cross-validate against state motion data, not just action data.&lt;/p&gt;

&lt;p&gt;Previously, RDA's video freeze detection checked whether the robot's actions indicated motion during a visually frozen segment. But some datasets have unreliable action data, or no action data at all, while state data (joint positions) might be more trustworthy.&lt;/p&gt;

&lt;p&gt;The new flag has three modes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;action&lt;/strong&gt; (default): Pure vision-based, unchanged from prior behavior. All existing golden-set results reproduce exactly.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;state&lt;/strong&gt;: Cross-validates each visually-frozen segment against observation.state motion. If state motion exceeds the threshold during the frozen segment, the freeze verdict stays (robot commanded stop but video didn't update — real artifact). If state is also still, the segment downgrades to REVIEW (whole-machine stall, likely a data collection pause rather than a video glitch).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;auto&lt;/strong&gt;: Uses state cross-validation when state data is available, falls back to action mode otherwise.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The metric now reports freeze_motion_source (which mode actually ran) and state_cross_validated_segments (how many segments got downgraded) in the JSON output, so the decision path is transparent.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;Thresholds in Frame Equivalents&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The second suggestion from LIBERO#148 was simpler but genuinely useful: annotate thresholds in frame counts, not just seconds.&lt;/p&gt;

&lt;p&gt;When RDA says "minimum freeze duration: 0.5 seconds," you have to do mental arithmetic to figure out how many frames that is at your dataset's frame rate. With the new version, the JSON and text reports annotate thresholds in frame counts for the configured fps. For example, at 30fps, the default 0.5s threshold shows as "≈15 frames." Small change, but it makes the output immediately readable.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;The Bigger Picture&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Four changes, all from one person's two issues. The developer clearly uses RDA seriously, understands the detection algorithms, and cares enough to run his own analysis and share the results. That's the best kind of open source feedback loop.&lt;/p&gt;

&lt;p&gt;The action unit inference, per-joint exposure, scale invariance verification, and state cross-validation aren't just features — they're responses to real usage patterns from someone doing the same kind of dataset auditing work we're trying to support.&lt;/p&gt;

&lt;p&gt;RDA v0.9.14 is on PyPI now: pip install robot-data-audit==0.9.14&lt;/p&gt;

&lt;p&gt;GitHub releases: &lt;a href="https://github.com/liesliy/rda/releases" rel="noopener noreferrer"&gt;https://github.com/liesliy/rda/releases&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you're auditing robot datasets and have feedback, the issue tracker is open. Come with data.&lt;/p&gt;

</description>
      <category>robotics</category>
      <category>data</category>
      <category>tooling</category>
      <category>opensource</category>
    </item>
    <item>
      <title>When We Audited Someone Else's Robot Dataset on GitHub, the Hardware Maker Responded</title>
      <dc:creator>liesliy</dc:creator>
      <pubDate>Fri, 18 Sep 2026 03:16:29 +0000</pubDate>
      <link>https://dev.to/liesliy/when-we-audited-someone-elses-robot-dataset-on-github-the-hardware-maker-responded-1l2o</link>
      <guid>https://dev.to/liesliy/when-we-audited-someone-elses-robot-dataset-on-github-the-hardware-maker-responded-1l2o</guid>
      <description>&lt;p&gt;We ran an automated quality audit on two public xArm robot datasets, posted the report as a GitHub Issue, and got a response that changed how we think about data quality metrics.&lt;/p&gt;

&lt;p&gt;Here's what happened, and what it means for anyone building robot learning pipelines.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;The Setup&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;RDA (Robot Data Audit) is an open-source tool we built for automated quality auditing of LeRobot-format datasets. It runs 19 diagnostic checks across structural integrity, temporal consistency, trajectory quality, and visual health — and produces a layered report in under a minute.&lt;/p&gt;

&lt;p&gt;The latest version, v0.9.12, introduced three new capabilities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Gap detection in timestamp validation (detecting anomalous time intervals that suggest dropped frames)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Robust sampling jitter calculation (separating gap-induced noise from actual clock instability)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Path B timestamp inference for missing frame detection (when frame indices aren't available)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We wanted to test it on real-world data, so we picked two publicly available xArm datasets on HuggingFace: xarm_push_medium and xarm_lift_medium, each with 800 episodes of roughly 20,000 frames.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;What the Audit Found&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The structural integrity checks passed cleanly — no missing frames, no NaN values, no schema inconsistencies. The temporal checks also passed. But the trajectory diagnostics told a more interesting story.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;xarm_push_medium:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;845 action discontinuity spikes across 500 episodes (62.5% affected)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;State space occupancy at just 1.7% median — the arm operates in a very narrow region of its joint space&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;83.3% median idle ratio — only 16.7% of frames show meaningful motion&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;xarm_lift_medium:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;908 extreme acceleration spikes across 439 episodes (54.9% affected)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Similar low occupancy (2.4%) and high idle ratio&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At first glance, these numbers look concerning. Hundreds of spikes? 83% idle? Is this data broken?&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;The Response That Changed Everything&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I posted the audit report as Issue #158 on the xArm-Developer/xArm-Python-SDK repository. Three days later, an engineer from UFACTORY responded.&lt;/p&gt;

&lt;p&gt;The engineer first point was a correction: those datasets weren't collected by UFACTORY. They came from Nicklas Hansen's TD-MPC paper (arXiv:2203.04955) and were later uploaded to the LeRobot platform.&lt;/p&gt;

&lt;p&gt;But the second point was far more valuable:&lt;/p&gt;

&lt;p&gt;"If operated with non-interpolated mode, then the spikes would appear easily."&lt;/p&gt;

&lt;p&gt;And then he shared UFACTORY's standard practice:&lt;/p&gt;

&lt;p&gt;"We normally recommend collecting data with online trajectory planning modes (6 or 7) to utilize the dynamic interpolation and velocity control in our control box. Under those modes, the accelerations would be continuous theoretically."&lt;/p&gt;

&lt;p&gt;This was the missing context. The spikes weren't necessarily a data quality problem — they were a characteristic of how the data was collected.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;Why This Matters for Metric Design&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;UFACTORY's feedback highlighted a fundamental tension in data quality tooling: &lt;strong&gt;metrics without context can be misleading.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The action_discontinuity metric&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;RDA detects action discontinuities by computing the second-order difference of the action array, then applying MAD (Median Absolute Deviation) z-score analysis. Spikes are frames where z &amp;gt; 5.0.&lt;/p&gt;

&lt;p&gt;The MAD approach is robust — unlike standard deviation, it doesn't get pulled off course by the outliers it's trying to detect. But it's also context-blind. It doesn't know whether the data came from:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Direct teleoperation (no smoothing, every micro-movement recorded)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Online trajectory planning with dynamic interpolation (smooth by design)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;A learned policy generating actions autoregressively&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The same spike count means very different things depending on the collection mode. In non-interpolated teleoperation, spikes at grasp/release transitions are expected. In mode 6/7 with interpolation, they'd signal something genuinely wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The idle_ratio metric&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An 83.3% idle ratio sounds terrible. But in RDA's benchmark across 13 public datasets, 11 had median idle ratios above 63%. High idle ratios are normal for manipulation data.&lt;/p&gt;

&lt;p&gt;A push task involves slowly approaching an object, making fine adjustments, waiting for contact — all of which produce tiny action changes that fall below the motion threshold. This isn't noise; it's the task.&lt;/p&gt;

&lt;p&gt;The metric's role is observational: report the number, let the user decide based on their model architecture. Frame-level models like Diffusion Policy might want to trim idle frames. Temporal models might use them as useful context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The velocity_acceleration metric&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The 908 extreme acceleration spikes in xarm_lift_medium made sense once we understood the collection method. Without interpolation, rapid start-stop motions during pick-and-place create genuine acceleration discontinuities. With mode 6/7's built-in velocity control, these would be smoothed out.&lt;/p&gt;

&lt;p&gt;The metric correctly detected the phenomenon. What it couldn't do — yet — was distinguish between "expected artifact of this collection mode" and "actual data quality problem."&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;What RDA Does Next&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This Issue exchange crystallized three improvements for the next RDA release:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The gap_multiplier_used field (v0.9.13)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The gap detection algorithm uses a threshold of median_dt × 2.5. This 2.5 is an internal constant, not exposed to users. But audit reports should be transparent about what thresholds were used. Starting in v0.9.13, the output includes gap_multiplier_used: 2.5 so every report documents exactly how gaps were detected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Control mode metadata&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Future reports should include a "collection method" annotation. When action_discontinuity or velocity_acceleration values are reported, they should be interpreted relative to the expected characteristics of that collection mode. A spike rate of 0.04 might be normal for non-interpolated teleoperation but alarming for interpolated trajectories.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. False Removal Rate research (v0.9.14)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The deeper challenge: how do you distinguish "real problems" from "expected features"? This requires building a labeled dataset of known-good and known-bad data across different collection modes, then measuring how often RDA's diagnostics correctly separate the two.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;The Design Philosophy&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;RDA's architecture separates hard judgments from soft diagnostics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Integrity Gate (L1)&lt;/strong&gt; makes binary calls: missing frames, NaN values, time reversals → EXCLUDE. These are objective, verifiable problems.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Trajectory Diagnostics (L2)&lt;/strong&gt; outputs numbers: spike counts, jitter values, idle ratios. No automatic verdicts.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Dataset Profile (L3)&lt;/strong&gt; provides context: distributions, coverage, temporal structure.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This separation exists because "is this data good?" is not a question a tool can answer alone. It requires domain knowledge, task context, and model architecture considerations. RDA's job is to make the data's characteristics visible — not to pretend it knows more than it does.&lt;/p&gt;

&lt;p&gt;UFACTORY's response validated this approach. If RDA had seen 845 spikes and automatically declared "poor data quality," it would have been wrong. The spikes were a feature of the collection method, not a bug in the data.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;Practical Takeaways&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you're collecting xArm data for robot learning:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Use mode 6 or 7&lt;/strong&gt; (online trajectory planning) when possible. The control box's dynamic interpolation produces naturally smooth accelerations.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;If using non-interpolated mode&lt;/strong&gt;, expect higher spike counts in RDA reports. This is normal, not necessarily problematic.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Compare episodes using spike_rate&lt;/strong&gt; (spikes per step), not spike_count. Episode lengths vary, and longer episodes naturally produce more spikes.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Don't panic about high idle ratios. For push-type tasks, 80%+ idle is typical. Whether to trim those frames depends on your model architecture.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Always check the collection metadata&lt;/strong&gt; alongside audit numbers. Numbers without context are just numbers.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;The Bigger Picture&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Open-source tools improve through open feedback loops. We built RDA to audit datasets. We used it on public data. We shared the results publicly. And a hardware manufacturer gave us exactly the kind of contextual insight that makes the tool better.&lt;/p&gt;

&lt;p&gt;This is how robust data quality tooling gets built — not in isolation, but through the messy, iterative process of applying metrics to real data and learning from domain experts when the numbers don't mean what you thought they meant.&lt;/p&gt;

&lt;p&gt;RDA v0.9.12 is available on PyPI. The GitHub repository is at github.com/liesliy/rda.&lt;/p&gt;

&lt;p&gt;This article is based on a real GitHub Issue exchange between the RDA team and UFACTORY engineers. The audit report was posted as Issue #158 on the xArm-Developer/xArm-Python-SDK repository.&lt;/p&gt;

</description>
      <category>robotics</category>
      <category>opensource</category>
      <category>data</category>
      <category>test</category>
    </item>
    <item>
      <title>Designing Idle Detection for Robot Datasets: Why We Chose Per-Episode Adaptive Thresholds</title>
      <dc:creator>liesliy</dc:creator>
      <pubDate>Thu, 17 Sep 2026 02:19:02 +0000</pubDate>
      <link>https://dev.to/liesliy/designing-idle-detection-for-robot-datasets-why-we-chose-per-episode-adaptive-thresholds-2hpg</link>
      <guid>https://dev.to/liesliy/designing-idle-detection-for-robot-datasets-why-we-chose-per-episode-adaptive-thresholds-2hpg</guid>
      <description>&lt;p&gt;Recently someone raised a thoughtful question on our GitHub Issue:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is the 'minimal state change' threshold in RDA using normalized units or raw STS3215 encoder steps (4096 per turn at 30 fps)? At step-level thresholds, slow fine motion near the brick reads as idle.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This question cuts to the core of idle detection design. Let me explain our reasoning and the trade-offs involved.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;The Problem&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In robot data auditing, idle detection is fundamental but deceptively complex. How do you define "idle" when auditing an episode?&lt;/p&gt;

&lt;p&gt;Intuitively, a robot standing still is idle. But what does "still" mean?&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Sensor noise causing micro-jitters — is that idle?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Slow approach before grasping — is that idle?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Holding pose while waiting for the next command — is that idle?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There's no universal answer.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;Why Not Fixed Thresholds?&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Our first attempt used a fixed threshold: ||Δaction|| &amp;lt; 0.01 means idle.&lt;/p&gt;

&lt;p&gt;Problem: &lt;strong&gt;action scales vary wildly across datasets.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;SO-101 with STS3215 servos: action range 0-4096 (encoder steps)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Franka Panda: normalized [-1, 1]&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;ALOHA: different scale again&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A fixed threshold works for one dataset but fails catastrophically on another.&lt;/p&gt;

&lt;p&gt;Normalization seems like the obvious fix, but normalizing to what? Each dataset has its own action distribution characteristics. Unified normalization loses dataset-specific information.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;The Bimodal Gap Detection Approach&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We realized that &lt;strong&gt;each episode has its own "motion distribution" .&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A typical manipulation episode contains two types of frames:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Idle frames: robot stationary or micro-jittering, small action changes&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Active frames: robot moving, significant action changes&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These two types usually form a &lt;strong&gt;bimodal distribution&lt;/strong&gt; in action change magnitudes.&lt;/p&gt;

&lt;p&gt;Our algorithm:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Compute motion: ||Δaction|| (L2 norm of action first-difference)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Build a histogram (30 bins)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Find the valley between the "low cluster" (idle) and "high cluster" (active)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The valley position becomes the threshold&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If no clear bimodal structure exists (e.g., all frames are slow movements), we fall back to 3 × MAD (Median Absolute Deviation) as a conservative estimate.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;Answering the Issue Question&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RDA's idle threshold operates in raw action space units, not normalized.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For SO-101, this means the threshold is at the encoder step level. But the threshold is &lt;strong&gt;computed per-episode adaptively&lt;/strong&gt;, not a fixed global constant.&lt;/p&gt;

&lt;p&gt;Benefits:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Preserves dataset-specific action scale characteristics&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Each episode adjusts its own threshold based on its motion distribution&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;No cross-dataset normalization assumptions required&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The cost?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Exactly what the Issue identified: slow fine motions can be misclassified as idle.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If a task involves entirely slow, precise operations (e.g., precision assembly), the action change magnitudes are uniformly small, the bimodal distribution may not exist, and the threshold becomes very small — causing many "meaningful but slow" motions to be classified as idle.&lt;/p&gt;

&lt;p&gt;This is a real limitation, not a bug.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;The Design Trade-off&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We considered four approaches:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fixed threshold&lt;/strong&gt; — Simple and interpretable, but not comparable across datasets. A threshold that works for SO-101 fails on Franka.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Normalized threshold&lt;/strong&gt; — Cross-dataset comparable, but loses dataset-specific characteristics. Normalizing to what? Each dataset has its own action distribution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Per-episode adaptive (what we chose)&lt;/strong&gt; — Adapts to each episode's motion distribution. The downside: slow tasks may have their fine motions misclassified as idle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Learned threshold&lt;/strong&gt; — Theoretically optimal, but requires labeled data and generalization is unknown.&lt;/p&gt;

&lt;p&gt;We chose per-episode adaptive because:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;RDA is a diagnostic tool, not a pass/fail gate.&lt;/strong&gt; We output idle_ratio as a measurement, not a judgment. Users interpret the value based on their task characteristics.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Per-episode adaptive works well for most manipulation tasks.&lt;/strong&gt; Typical pick-and-place, insertion tasks have clear bimodal distributions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;For special tasks (entirely slow operations), users should be aware that idle_ratio may be high.&lt;/strong&gt; This itself is valuable diagnostic information.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;What About Slow Tasks?&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Issue questioner's concern about "slow fine motion near the brick" is a real scenario.&lt;/p&gt;

&lt;p&gt;If the task itself is like this, we suggest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Don't just look at the absolute idle_ratio value.&lt;/strong&gt; Compare idle_ratio across episodes of the same task — relative differences are more meaningful.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Look at action discontinuity alongside idle ratio.&lt;/strong&gt; If idle_ratio is high AND action discontinuity is high, the data may have structural issues (e.g., control mode switching during teleoperation). If both are normal, it's likely just the task's inherent characteristic.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;**Consider task-specific threshold tuning. **RDA's idle detection parameters are configurable (mad_multiplier, abs_threshold_floor), adjustable based on actual data.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;Action Discontinuity: Another Perspective&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Beyond idle detection, RDA also detects spikes in the action signal — sudden jumps that may indicate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Control mode switching during teleoperation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Frame drops or interpolation issues in data collection&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;High-frequency controller oscillation&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Spikes are detected on the &lt;strong&gt;second difference&lt;/strong&gt; of the action signal (Δ²a), using MAD-based z-scores: z = 0.6745 × (x - median) / MAD, with a default threshold of |z| &amp;gt; 5.0.&lt;/p&gt;

&lt;p&gt;This is also in raw action units, per-episode adaptive.&lt;/p&gt;

&lt;p&gt;In our G1 dataset audit (300 episodes, 177,811 frames), &lt;strong&gt;299/300 episodes had spikes (3,340 total)&lt;/strong&gt;. Combined with the 65.6% median idle ratio, this suggests the teleoperation data collection process has inherent characteristics worth understanding rather than "fixing."&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;Reflections&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;There's no universal "idle" definition.&lt;/strong&gt;Different tasks, different hardware, different data collection methods — the understanding of "idle" varies. Tools can provide measurements, not judgments.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Adaptive thresholds come at the cost of uncertainty in edge cases.&lt;/strong&gt;Bimodal detection works well on data with clear "still vs. moving" separation. It fails on "entirely slow" data. Users need to understand this boundary.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Open-source community value lies in these real discussions.&lt;/strong&gt;The Issue made us re-examine idle detection's limitations and prompted us to think about how to better communicate applicable scenarios in documentation.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;RDA is currently v0.9.12, and idle detection design is still iterating. If you have datasets with slow, precise operations, we welcome you to test it and discuss on GitHub.&lt;br&gt;
Repository: github.com/liesliy/rda&lt;br&gt;
Related Issue: huggingface/lerobot#4650&lt;/p&gt;

</description>
      <category>robotics</category>
      <category>data</category>
      <category>opensource</category>
      <category>algorithms</category>
    </item>
    <item>
      <title>Detecting Frozen Video in Robot Datasets: Why First/Last Frame Comparison Isn't Enough</title>
      <dc:creator>liesliy</dc:creator>
      <pubDate>Wed, 16 Sep 2026 02:22:17 +0000</pubDate>
      <link>https://dev.to/liesliy/detecting-frozen-video-in-robot-datasets-why-firstlast-frame-comparison-isnt-enough-53ge</link>
      <guid>https://dev.to/liesliy/detecting-frozen-video-in-robot-datasets-why-firstlast-frame-comparison-isnt-enough-53ge</guid>
      <description>&lt;p&gt;While auditing robotics learning datasets, we stumbled onto a problem that seemed simple but turned out to have real depth: detecting frozen video frames.&lt;/p&gt;

&lt;p&gt;The naive approach is obvious — compare the first and last frame of each episode. If they're identical, the video is frozen. This is what some popular tools do. But after testing it on real datasets, we found it misses the cases that actually matter.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;What We Found in LIBERO&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We ran our open-source audit tool RDA against the LIBERO benchmark — 10 manipulation tasks, 379 episodes total. The audit flagged &lt;strong&gt;6 episodes with video freezes (1.6%)&lt;/strong&gt;, all short segments at episode boundaries.&lt;/p&gt;

&lt;p&gt;But here's the thing: the first/last frame comparison method would have caught some of these by coincidence, while also producing false positives on episodes where the robot simply starts and ends in the same position (which is completely normal for manipulation tasks).&lt;/p&gt;

&lt;p&gt;The real issue isn't whether frames are identical — it's whether the video fails to capture motion that should be happening.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;The Action Cross-Reference Approach&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is the key insight: &lt;strong&gt;a frozen video isn't the problem. A frozen video while the robot is supposed to be moving — that's the problem.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;RDA's video freeze detection works in four stages:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Decode to low-res grayscale (64×64)&lt;/strong&gt; — noise reduction matters more than resolution&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Adaptive thresholding&lt;/strong&gt; — compute frame-to-frame pixel differences, use the p10 as a noise floor, set threshold at max(0.10, 0.25 × noise_floor). This automatically adapts to different camera qualities across datasets.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Continuous segment detection&lt;/strong&gt; — only flag sequences of consecutive frozen frames (minimum 0.5s by default), filtering out single-frame noise&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Action cross-validation&lt;/strong&gt; — check the action data during frozen segments. If the robot's commanded actions show movement but the video doesn't change, it's a real freeze. If actions are also near zero, the robot is simply idle.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This distinction eliminates the two failure modes of simpler approaches:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;False positives&lt;/strong&gt;: Robot starts and ends at the same position → first/last frames match → incorrectly flagged as frozen&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;False negatives&lt;/strong&gt;: Video freezes for 5 seconds in the middle of an episode → first/last frames are different → freeze completely missed&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;Validation on G1_WBT&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;To verify the method, we audited the Unitree G1_WBT_Brainco_Pickup_Pillow dataset — 300 episodes, 4 camera views (head stereo + wrist), 177,811 frames at 30fps.&lt;/p&gt;

&lt;p&gt;Results on the 44 episodes with available video:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;0 frozen segments detected&lt;/strong&gt; across all camera streams&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Camera sync&lt;/strong&gt;: temporal offset = 0ms, drift rate = 0ms/min — excellent hardware synchronization&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Episode 18's wrist camera had a visual quality dip (score 0.5/1.0, blurry frames), but RDA correctly classified this as "degraded quality" rather than "frozen" — because the frame pixels were changing, just not sharply&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is exactly where the action cross-reference matters: blurry frames have pixel variation, frozen frames don't. The tool distinguishes between them.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;What Else We Found&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Video freeze turned out to be the least interesting finding. The more impactful issues across these datasets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;LIBERO — Action discontinuities&lt;/strong&gt;: 3,622 spikes across 100% of episodes&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;LIBERO — High idle ratio&lt;/strong&gt;: Only 28.3% of frames show effective motion&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;G1_WBT — Missing video files&lt;/strong&gt;: 278 out of 300 episodes have no video streams&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;G1_WBT — Joint limit violations&lt;/strong&gt;: 7 out of 36 DOF exceed configured limits&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These structural issues are harder to spot than frozen frames and likely have a bigger impact on policy training. Video freeze is visible and easy to understand; action discontinuity and low state-space occupancy are invisible but potentially more harmful.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;The Bigger Picture&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Robot learning datasets are growing fast — Open X-Embodiment, DROID, AgiBot World — but quality assurance hasn't kept up. Most datasets are published with basic sanity checks but no systematic audit of temporal integrity, action continuity, or state-space coverage.&lt;/p&gt;

&lt;p&gt;We built RDA because we kept running into these issues during our own work. It's open source (v0.9.9 on PyPI), supports LeRobot v2/v3 format, and produces structured JSON reports. We've audited 12+ datasets so far and shared findings directly with dataset maintainers through GitHub issues.&lt;/p&gt;

&lt;p&gt;If you're collecting robot data or training policies on open datasets, video freeze is worth checking — but it's probably not the biggest quality issue hiding in your data.&lt;/p&gt;

&lt;p&gt;Tool: RDA (Robot Data Audit)&lt;a href="https://github.com/liesliy/rda" rel="noopener noreferrer"&gt;https://github.com/liesliy/rda&lt;/a&gt; — pip install robot-data-audit&lt;/p&gt;

</description>
      <category>robotics</category>
      <category>data</category>
      <category>testing</category>
      <category>opensource</category>
    </item>
    <item>
      <title>We Audited 13 Public Robot Datasets With One Tool and Zero Tuning. Here's What the Numbers Actually Tell Us.</title>
      <dc:creator>liesliy</dc:creator>
      <pubDate>Tue, 15 Sep 2026 07:35:19 +0000</pubDate>
      <link>https://dev.to/liesliy/we-audited-13-public-robot-datasets-with-one-tool-and-zero-tuning-heres-what-the-numbers-actually-1f41</link>
      <guid>https://dev.to/liesliy/we-audited-13-public-robot-datasets-with-one-tool-and-zero-tuning-heres-what-the-numbers-actually-1f41</guid>
      <description>&lt;p&gt;RDA v0.9.7 was run against 13 LeRobot-format datasets from HuggingFace Hub — 4,940 episodes total, default thresholds, zero per-dataset tuning. The results challenge the assumption that "clean data" and "good training data" are the same thing.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;The Problem&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When most people evaluate robot manipulation datasets, they check one thing: is the data broken? Missing frames? NaN values? Timestamp inversions? If none of those — "data's fine, let's train."&lt;/p&gt;

&lt;p&gt;That logic has a gap.&lt;/p&gt;

&lt;p&gt;We ran RDA (Robot Data Audit) across 13 public LeRobot-format datasets — simulation and real, bimanual and single-arm, scripted and teleoperated — and found that while &lt;strong&gt;all 4,940 episodes pass integrity checks (L1)&lt;/strong&gt; , the behavioral profiles (L2) and training efficiency metrics (L3) vary dramatically. And those variations directly affect training outcomes, yet they're invisible to conventional quality checks.&lt;/p&gt;

&lt;p&gt;This article isn't about ranking datasets. It's about showing what a &lt;strong&gt;multi-layer audit&lt;/strong&gt; reveals when you stop asking "is this data broken?" and start asking "what is this data actually like?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Four-Layer Framework&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;RDA's audit runs every episode through four sequential layers. The design rule is strict: **only hard integrity checks can set an EXCLUDE verdict; diagnostic measurements never do.##&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;L1 — Integrity Gate · Deterministic hard checks (missing / NaN / limit / video-stream) · 9 metrics · ✅ PASS → REVIEW / EXCLUDE&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;L2 — Trajectory Diagnostics · Observational motion &amp;amp; video anomalies · 8 metrics · ❌ findings only&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;L3 — Dataset Profile · Training-data efficiency &amp;amp; coverage · 4 metrics · ❌ findings only&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;L4 — Dataset Summary · Dataset-level P10/P50/P90 aggregation · 📊 report only&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This separation — &lt;strong&gt;measurement vs. judgment&lt;/strong&gt; — is the core design philosophy of RDA v0.9.7. And it wasn't always this way.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;&lt;em&gt;Layer 1: Everything Passes (That's the Starting Line, Not the Finish Line)&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;All 13 datasets, all 4,940 episodes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;✅ missing_dropout — pass (4,940/4,940)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;✅ invalid_values — pass (4,940/4,940)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;✅ schema_consistency — pass (4,940/4,940)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;✅ timestamp_validity — pass (4,940/4,940)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;✅ video_frame_integrity — pass (all video episodes)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;✅ video_timestamp_alignment — pass (all video episodes)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Zero corruption. Zero NaN. Zero timestamp inversions. If your audit stops here, the conclusion is: "all good."&lt;/p&gt;

&lt;p&gt;But it doesn't stop here.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Layer 2: The Behavior Layer — Where the Interesting Stories Are&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Finding 1: Same Robot, Same Lab, 4× Difference in Motion&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;xarm_lift_medium and xarm_push_medium are both collected on the same xArm platform, in the same lab, by the same team. But their behavioral profiles are radically different:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;xarm_lift_medium:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Effective motion ratio (median): 79.2%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Idle ratio (median): 20.8%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Episodes with L2 findings: 72/800 (9%)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Action discontinuity (median): 0 spikes/ep&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;xarm_push_medium:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Effective motion ratio (median): 16.7%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Idle ratio (median): 83.3%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Episodes with L2 findings: 629/800 (78%)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Action discontinuity (median): 1 spike/ep&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Lift is a large-amplitude pick-and-place task — the arm moves most of the time. Push is a small-force nudge-then-watch task — the arm spends most of its time waiting.&lt;/p&gt;

&lt;p&gt;Neither is "bad data." But a loss function trained on a 75%-idle distribution is &lt;strong&gt;structurally biased toward predicting "do nothing."&lt;/strong&gt; If you don't know this about your data, you won't know where to look when training plateaus.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Finding 2: Action Discontinuity Tracks the Controller, Not the Dataset's Reputation&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Action discontinuity measures sudden jumps in the action sequence using MAD-based spike detection. The distribution across datasets reveals something counterintuitive:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;aloha_sim_insertion_scripted — 50/50 (100%) · median 33 spikes/ep · Scripted controller discretization&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;aloha_sim_transfer_cube_scripted — 50/50 (100%) · median 49 spikes/ep · Scripted controller discretization&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;aloha_sim_transfer_cube_human — 47/50 (94%) · median 31 spikes/ep · Teleoperation corrections&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;cmu_stretch — 72/135 (53%) · median 22 spikes/ep · Binary gripper 0↔1 flips&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;jaco_play — 176/1085 (16%) · median 10 spikes/ep · Occasional corrections&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;libero_10 — 0/379 (0%) · median 10 spikes/ep · Smooth teleoperation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;pusht — 0/206 (0%) · median 5 spikes/ep · Smooth simulation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;xarm_lift_medium — 0/800 (0%) · median 0 spikes/ep · Extremely smooth&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The ALOHA scripted datasets have &lt;strong&gt;100% of episodes with action spikes&lt;/strong&gt; — not because the data is corrupted, but because scripted controllers produce discrete command transitions that manifest as step functions in the action space.&lt;/p&gt;

&lt;p&gt;If your policy uses smoothness regularization, &lt;strong&gt;this number decides your curriculum.&lt;/strong&gt; You need different regularization strength for scripted vs. teleoperated data, and that's a Layer 2 insight that L1 alone would never reveal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Finding 3: Video Freeze — A Signal That Needs Context&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In libero_10, RDA detected 6 episodes with video_freeze (stuck frames, detected via inter-frame pixel difference analysis):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Ep 90 — 223 frames · ~0.5s freeze · wrist_image&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ep 198 — 235 frames · ~0.7s freeze · wrist_image&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ep 240 — 289 frames · ~0.8s freeze · wrist_image&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ep 254 — 227 frames · ~0.6s freeze · wrist_image&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ep 255 — 291 frames · ~0.9s freeze · wrist_image&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ep 300 — 237 frames · ~0.7s freeze · wrist_image&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Initial reaction: "camera malfunction, discard these episodes."&lt;/p&gt;

&lt;p&gt;But cross-validation using ffmpeg frame-by-frame extraction from both cameras (agentview + wrist_image) revealed that the frozen frames appear simultaneously in both cameras and are located at episode boundaries (start or end).&lt;/p&gt;

&lt;p&gt;This pattern is consistent with the robot naturally pausing between episode segments — not a camera hardware failure.&lt;/p&gt;

&lt;p&gt;RDA correctly flags these as REVIEW ("inspect before training") rather than EXCLUDE ("data is corrupted"). This distinction is the practical value of layered auditing: &lt;strong&gt;surface the signal, provide context, let the human decide.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Layer 3: The Training Efficiency Layer — The Most Overlooked Dimension&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Finding 4: State Space Occupancy Is Uniformly Low&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;State occupancy measures how much of the discretized state space (10×10 grid) is covered by a dataset's episodes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;pusht — median 39.0% · P10: 28.0% · P90: 51.0%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;libero_10 — median 6.6% · P10: 4.7% · P90: 8.3%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;utokyo_pr2_tabletop — median 5.0% · P10: 4.4% · P90: 5.8%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;droid_100 — median 4.8% · P10: 3.2% · P90: 11.0%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;svla_so101_pickplace — median 4.5% · P10: 4.0% · P90: 5.0%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;aloha_sim_transfer_cube_scripted — median 5.0% · P10: 5.0% · P90: 5.0%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;aloha_sim_insertion_human — median 3.8% · P10: 3.3% · P90: 4.4%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;jaco_play — median 2.8% · P10: 2.2% · P90: 3.4%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;xarm_lift_medium — median 2.4% · P10: 2.1% · P90: 2.5%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;cmu_stretch — median 1.9% · P10: 1.9% · P90: 2.0%&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;xarm_push_medium — median 1.7% · P10: 1.4% · P90: 2.1%&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Outside of pusht (a deliberately designed 2-DOF simulation), most datasets occupy only 1.7%–6.6% of their state space. Episodes within each dataset heavily overlap in state space, suggesting limited exploration diversity.&lt;/p&gt;

&lt;p&gt;This isn't a "quality" judgment. It's a measurement. But if your model needs to generalize to unseen states, &lt;strong&gt;this number tells you whether your current dataset has enough coverage&lt;/strong&gt; — and you'd never know from L1 alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Finding 5: Idle Ratio Varies 4× Across Datasets&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Median idle ratio across all 13 datasets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;20%–30%: xarm_lift (20.8%)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;63%–72%: ALOHA sim (4 datasets), droid_100, jaco_play, libero_10&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;79%–87%: pusht (81.7%), utokyo_pr2 (83.6%), svla_so101 (86.8%)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;11 of 13 datasets have median idle ratio above 63%.&lt;/strong&gt; The implication for training: loss functions on these distributions learn to predict "do nothing" as the default. This isn't a bug — it's a dataset characteristic. But without L3 measurement, you wouldn't know to adjust your training strategy accordingly.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;The Verdict Pipeline Change: Why Findings Don't Auto-Escalate&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Old Problem&lt;/strong&gt;&lt;br&gt;
In v0.5.x, idle_ratio findings above threshold automatically escalated the episode verdict to REVIEW:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;libero_10: 247/379 (65%) REVIEW → Root cause: Low-motion task, not data quality issue&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;pusht: 163/206 (79%) REVIEW → Root cause: Low-motion task, not data quality issue&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cross-validation showed these were false positives — high idle ratio is a natural property of fine-grained manipulation tasks, not a sign of data corruption.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The New Architecture&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;v0.9.7 restructured the verdict pipeline into a three-layer aggregate model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;L1 Hard Checks → directly determine PASS / EXCLUDE&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;L2/L3 Findings → reported as RISK_SIGNAL, never auto-escalate to verdict&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;L4 Summary → aggregated statistics for reporting only&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The Result&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;libero_10: v0.5.x → 132 PASS / 247 REVIEW → v0.9.7 → 373 PASS / 6 REVIEW&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;pusht: v0.5.x → 43 PASS / 163 REVIEW → v0.9.7 → 206 PASS / 0 REVIEW&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;All others: Various REVIEW counts → v0.9.7 → 100% PASS&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;libero_10's remaining 6 REVIEW verdicts come exclusively from video_freeze — an independent, more reliable signal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The separation of measurement and judgment reduced false-positive REVIEWs by 97-100% on low-motion datasets, while preserving genuine defect detection.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Blind Test: Precision 1.000, Recall 0.800 — And an Honest Regression&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Setup&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Dataset: lerobot/pusht (206 episodes)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Injection: 5 defect classes × 10 episodes = 50 defective episodes&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Controls: 156 unmodified episodes&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Mode: --no-video (Fast Audit)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Tuning: Zero tuning, all defaults&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Results&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;empty (stale metadata) — EXCLUDE 10/10 · REVIEW 0 · Strict 10/10 · Broad 10/10&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;NaN in state — EXCLUDE 10/10 · REVIEW 0 · Strict 10/10 · Broad 10/10&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;reversed timestamps — EXCLUDE 10/10 · REVIEW 0 · Strict 10/10 · Broad 10/10&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;frozen episodes — EXCLUDE 0/10 · REVIEW 0/10 · Strict 0/10 · Broad 0/10 ⚠️&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;duplicate frames — EXCLUDE 10/10 · REVIEW 0 · Strict 10/10 · Broad 10/10&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Strict: TP=40, FN=10, FP=0, TN=156 → precision 1.000, recall 0.800&lt;br&gt;
Zero false positives&lt;/strong&gt; — all 156 controls matched clean baseline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Regression We're Not Hiding&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Frozen episode detection is a known regression in v0.9.7. The idle_ratio metric still detects the signal (RISK_SIGNAL), but the new pipeline doesn't auto-escalate it to REVIEW. We're publishing this openly because a tool's credibility comes from transparency, not from claiming perfection.&lt;/p&gt;

&lt;p&gt;This will be addressed in a future version — either by restoring conditional idle_ratio escalation, or by adding a dedicated frozen-episode detector that distinguishes between naturally low-motion tasks and genuinely frozen data.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Design Philosophy: Measure, Don't Judge&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;RDA v0.9.7 makes a deliberate choice: be a measurement tool, &lt;strong&gt;not a judge.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;L1 answers: "Is the data broken?" (hard checks, can EXCLUDE)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;L2 answers: "What does the data look like behaviorally?" (findings only)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;L3 answers: "How efficient is this data for training?" (measurements only)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;L4 answers: "What's the overall distribution?" (summary only)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This choice comes from a simple recognition: &lt;strong&gt;no universal threshold works for every task, every platform, every collection method.&lt;/strong&gt; A 70% idle ratio is normal for a push task and suspicious for a lift task. 33 action spikes per episode is expected for scripted data and alarming for smooth teleoperation.&lt;/p&gt;

&lt;p&gt;RDA surfaces these signals. The judgment — "is this acceptable for my training pipeline?" — stays with you.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;What This Means in Practice&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;"Clean data" ≠ "ready to train." All 13 datasets pass L1. Their L2/L3 profiles differ dramatically. Ignoring the behavioral and efficiency layers means missing information that directly affects training.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Never discuss data quality without task context. Same xArm platform, 4× difference in idle ratio. Task nature, not collection quality, explains the gap.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Use findings, not verdicts, for dataset selection. The RISK_SIGNALs from L2/L3 give you actionable information about motion patterns, exploration coverage, and action smoothness — regardless of whether they trigger a verdict.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Velocity and occupancy are not cross-platform comparable without normalization. RDA classifies velocity as Tier-2 (normalizable) and occupancy as platform-dependent. You need scaling factors before comparing across robots.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;All data is real, reproducible, and comes from public HuggingFace datasets. Full JSON reports and reproduction scripts are available in the repository.&lt;/p&gt;

&lt;p&gt;RDA v0.9.7 on PyPI: pip install robot-data-audit==0.9.7&lt;br&gt;
GitHub: github.com/liesliy/rda&lt;br&gt;
Benchmark data: docs/benchmark.md&lt;/p&gt;

</description>
      <category>robotics</category>
      <category>data</category>
      <category>opensource</category>
      <category>ai</category>
    </item>
    <item>
      <title>I Audited AgiBot's New RL Dataset with RDA</title>
      <dc:creator>liesliy</dc:creator>
      <pubDate>Tue, 01 Sep 2026 09:33:05 +0000</pubDate>
      <link>https://dev.to/liesliy/i-audited-agibots-new-rl-dataset-with-rda-31ig</link>
      <guid>https://dev.to/liesliy/i-audited-agibots-new-rl-dataset-with-rda-31ig</guid>
      <description>&lt;p&gt;AgiBot just released &lt;strong&gt;&lt;a href="https://huggingface.co/datasets/agibot-world/AgiBotWorld2026" rel="noopener noreferrer"&gt;AGIBOT WORLD 2026 Phase 3&lt;/a&gt;&lt;/strong&gt;, a real-robot reinforcement-learning dataset. I used it as an independent test of &lt;a href="https://github.com/liesliy/rda" rel="noopener noreferrer"&gt;RDA&lt;/a&gt;: can a third-party audit tool run against a fresh, large dataset without any tweaks?&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Finding: RDA Spikes Align with AgiBot's Human-Takeover Labels
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3uqc8pyjsksr0vjjx2qq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3uqc8pyjsksr0vjjx2qq.png" alt=" " width="800" height="438"&gt;&lt;/a&gt;&lt;br&gt;
The real-robot HG-DAgger package includes an official &lt;code&gt;intervened&lt;/code&gt; column: frame-by-frame human-takeover labels. This gave us a rare chance to validate RDA.&lt;/p&gt;

&lt;p&gt;I aligned all &lt;strong&gt;1,712&lt;/strong&gt; RDA action-discontinuity spikes against these labels:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;226 spikes (13.2%)&lt;/strong&gt; fall within &lt;strong&gt;±1 frame of a takeover transition&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Random baseline: &lt;strong&gt;~4.3%&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;That is a &lt;strong&gt;3.1× enrichment&lt;/strong&gt; overall (1.6×–5.2× per episode)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most spikes are normal teleop motion — which is why RDA labels them REVIEW, not FAIL. But the clear enrichment at takeover boundaries shows the signal is grounded in real events, not pure noise.&lt;/p&gt;
&lt;h2&gt;
  
  
  Verdict Distribution: 0 EXCLUDE / 0 FAIL
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvzg1unfbnmuiiw4istvq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvzg1unfbnmuiiw4istvq.png" alt=" " width="800" height="616"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across &lt;strong&gt;1,112 episodes&lt;/strong&gt; (5 simulation tasks + one real-robot package):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;th&gt;Count&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;PASS&lt;/td&gt;
&lt;td&gt;31&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;REVIEW&lt;/td&gt;
&lt;td&gt;1,081&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;EXCLUDE&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FAIL&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All 31 PASS episodes come from simulation tasks. Verdict layering works — it is not a blanket REVIEW.&lt;/p&gt;
&lt;h2&gt;
  
  
  What Was Covered
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fth4kos6wolwzc5anc8t8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fth4kos6wolwzc5anc8t8.png" alt=" " width="800" height="417"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I did not settle for a single tiny split. I ran RDA over &lt;strong&gt;all 5 simulation tasks&lt;/strong&gt; and the &lt;strong&gt;smallest real-robot RL package&lt;/strong&gt; (3.57 GB — the others are 17–50 GB each and would not fit on my drive):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dataset&lt;/th&gt;
&lt;th&gt;Episodes&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Simulation × 5 tasks&lt;/td&gt;
&lt;td&gt;1,102&lt;/td&gt;
&lt;td&gt;200–283 episodes per task&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Real robot HG-DAgger&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;23,658 frames, 30 fps&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;
  
  
  Data Engineering Passed the Audit
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3zidwjfbm5ye5xsdq7fq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3zidwjfbm5ye5xsdq7fq.png" alt=" " width="800" height="508"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;4,448&lt;/strong&gt; integrity checks (missing frames, invalid values, schema, timestamps), &lt;strong&gt;0 failures&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Constant &lt;strong&gt;30 fps&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;6 sampled &lt;code&gt;info.json&lt;/code&gt; files, &lt;strong&gt;all LeRobot v2.1&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The data is ready for training. REVIEW items are worth a spot-check, but nothing needs to be cleaned before use.&lt;/p&gt;
&lt;h2&gt;
  
  
  Run It Yourself
&lt;/h2&gt;

&lt;p&gt;Grab the dataset from the &lt;a href="https://huggingface.co/datasets/agibot-world/AgiBotWorld2026" rel="noopener noreferrer"&gt;official HF repo&lt;/a&gt; (simulation splits are a few hundred MB; the audit doesn't read video files, so you can skip the &lt;code&gt;videos/&lt;/code&gt; tarballs), then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;robot-data-audit
rda audit ./your_dataset_dir
rda recommend ./your_dataset_dir &lt;span class="nt"&gt;--offline&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The output is per-episode detail plus an overall verdict, in JSON so it plugs into any pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scope Boundary (Said Honestly)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tested&lt;/strong&gt;: 5/5 simulation tasks (~500 MB) and 1/50 real-robot packages (2%).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tested&lt;/strong&gt;: the remaining 49 real-robot packages (~3.9 TB), ImitationLearning 4.4 TB, and RichInteraction 2.3 TB.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why the format conclusion generalizes&lt;/strong&gt;: every sampled &lt;code&gt;info.json&lt;/code&gt; uses LeRobot v2.1 with the same pipeline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No video files were read&lt;/strong&gt;, so you can skip the &lt;code&gt;videos/&lt;/code&gt; tarballs to save bandwidth.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The full 6 JSON audit reports are on hand for anyone who wants to reproduce or challenge the numbers.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;RDA (Robot Data Audit) is an open-source tool I maintain. PyPI: &lt;code&gt;robot-data-audit&lt;/code&gt;. GitHub: &lt;a href="https://github.com/liesliy/rda" rel="noopener noreferrer"&gt;liesliy/rda&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>robotics</category>
      <category>data</category>
      <category>test</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I Poisoned 50 Episodes of LeRobot's PushT Dataset. The Audit Tool Caught 40 — With Zero False Alarms.</title>
      <dc:creator>liesliy</dc:creator>
      <pubDate>Mon, 31 Aug 2026 03:51:54 +0000</pubDate>
      <link>https://dev.to/liesliy/i-poisoned-50-episodes-of-lerobots-pusht-dataset-the-audit-tool-caught-40-with-zero-false-963</link>
      <guid>https://dev.to/liesliy/i-poisoned-50-episodes-of-lerobots-pusht-dataset-the-audit-tool-caught-40-with-zero-false-963</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;: I took &lt;a href="https://huggingface.co/datasets/lerobot/pusht" rel="noopener noreferrer"&gt;lerobot/pusht&lt;/a&gt; — the same dataset used to train Diffusion Policy — injected 50 defective episodes across 5 defect classes (seed=42), kept 156 episodes untouched as controls, and ran &lt;a href="https://github.com/liesliy/rda" rel="noopener noreferrer"&gt;RDA&lt;/a&gt; (v0.5.4, default thresholds, zero tuning) against it. Result: &lt;strong&gt;precision 1.000, recall 0.800 on a strict criterion, and zero false positives on all 156 clean controls.&lt;/strong&gt; Broad criterion (including review signals): 50/50. Runtime for the full 206-episode audit: &lt;strong&gt;2 seconds&lt;/strong&gt;. All numbers below are from real runs; the honest caveats are published too.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem: benchmarks without ground truth prove nothing
&lt;/h2&gt;

&lt;p&gt;I maintain RDA (Robot Data Audit), an open-source quality-audit tool for robot manipulation datasets. Its README claims 13 metrics, a three-tier verdict system, and a 12-dataset benchmark. But every one of those numbers is an &lt;em&gt;observation&lt;/em&gt; on datasets whose ground truth we don't control. When the tool reports 83% idle ratio on a dataset, is that a real defect or just how that robot behaves? Without answers, "audit" is just a nicer word for "aggregate statistics".&lt;/p&gt;

&lt;p&gt;The standard answer is a &lt;strong&gt;blind test&lt;/strong&gt;: manufacture the ground truth, then see what the tool actually finds. This post documents the first one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why lerobot/pusht
&lt;/h2&gt;

&lt;p&gt;Three reasons, all practical:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;It's famous.&lt;/strong&gt; PushT is the manipulation dataset from the Diffusion Policy paper, hosted by HuggingFace's official LeRobot org. If a data-quality tool wants credibility, this is a dataset its audience already knows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It's standard.&lt;/strong&gt; LeRobot v3.0 format — parquet + mp4 + episode metadata, exactly the format RDA supports natively.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It's small.&lt;/strong&gt; The full dataset is 16 files / 7.7 MB / 206 episodes / 25,650 frames — small enough to audit &lt;em&gt;in full&lt;/em&gt;, no sampling shortcuts.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Download integrity was verified byte-for-byte: 25,650 frames = sum of episode lengths = video frame count; all LFS sha256 hashes matched.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment design
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Step 1 — Baseline.&lt;/strong&gt; Audit the untouched dataset first. Result: 43 PASS / 163 REVIEW / 0 EXCLUDE. The 163 reviews all come from &lt;code&gt;idle_ratio&lt;/code&gt; — PushT is a low-motion task (median idle ratio 82%), the robot spends most of its time repositioning. That's a &lt;em&gt;property&lt;/em&gt; of the data, not a defect, and it matters later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2 — Inject.&lt;/strong&gt; 5 defect classes × 10 episodes each (seed=42), leaving &lt;strong&gt;156 episodes untouched as controls&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;empty&lt;/code&gt;&lt;/strong&gt; — delete all rows from the data parquet, leave &lt;code&gt;meta/episodes&lt;/code&gt; claiming the episode still exists → expected: EXCLUDE&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;nan_state&lt;/code&gt;&lt;/strong&gt; — set &lt;code&gt;observation.state&lt;/code&gt; to NaN on 15 random frames per episode → expected: EXCLUDE&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;timestamp_reverse&lt;/code&gt;&lt;/strong&gt; — reverse the second half of each episode's timestamps → expected: EXCLUDE&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;frozen&lt;/code&gt;&lt;/strong&gt; — freeze the entire episode at first-frame values → expected: REVIEW (statistical signal — shouldn't hard-fail)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;duplicate_frames&lt;/code&gt;&lt;/strong&gt; — duplicate 5 random frames per episode, copies appended at the end → honest probe, see caveats&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Step 3 — Audit blind.&lt;/strong&gt; &lt;code&gt;rda audit&lt;/code&gt; v0.5.4, default thresholds, no tuning, then compare every verdict against the manifest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Confusion matrix — strict criterion (EXCLUDE = flagged):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 flagged   not flagged
defective (50)   TP = 40   FN = 10
control   (156)  FP =  0   TN = 156
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Precision 1.000 · Recall 0.800 · Zero false positives on 156 untouched controls&lt;/strong&gt; — every control episode's verdict is identical to its clean-baseline verdict, so the injection itself didn't perturb anything else.&lt;/p&gt;

&lt;p&gt;Broad criterion (EXCLUDE + REVIEW = flagged): &lt;strong&gt;50/50, recall 1.000.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Per class (strict / broad / detector):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;empty&lt;/strong&gt; — 10/10 strict · 10/10 broad · &lt;code&gt;_zero_frame_guard&lt;/code&gt;, a regression probe for a P0 bug we fixed in v0.4.12 (zero-frame episodes used to silently PASS)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;nan_state&lt;/strong&gt; — 10/10 strict · 10/10 broad · &lt;code&gt;invalid_values&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;timestamp_reverse&lt;/strong&gt; — 10/10 strict · 10/10 broad · &lt;code&gt;timestamp_validity&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;frozen&lt;/strong&gt; — 0/10 strict · &lt;strong&gt;10/10 broad&lt;/strong&gt; · &lt;code&gt;idle_ratio&lt;/code&gt; (effective motion 0%) → REVIEW, by design&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;duplicate_frames&lt;/strong&gt; — 10/10 strict · 10/10 broad · &lt;code&gt;timestamp_validity&lt;/code&gt;, only because of placement; see caveats&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Runtime: &lt;strong&gt;2 seconds&lt;/strong&gt; for 206 episodes / 24,488 frames on a consumer laptop.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest caveats (the part most tool posts skip)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;duplicate_frames&lt;/code&gt; detection is an artifact of placement.&lt;/strong&gt; The copies landed at the episode tail, creating 2–4 negative timestamp deltas each. A &lt;em&gt;mid-stream&lt;/em&gt; duplicate with monotone timestamps would &lt;strong&gt;not&lt;/strong&gt; be flagged by 0.5.4. I publish this because the class name sounds scarier than what was measured.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frozen episodes get REVIEW, not EXCLUDE.&lt;/strong&gt; RDA treats statistical anomalies as review signals, not hard failures. If you believe frozen arms must hard-fail, that's a one-line policy change — the thresholds are open, argue with us in an issue.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;pusht is an unusually easy payload&lt;/strong&gt;: 2-D state, 96×96 video, no multi-stream timestamps. &lt;code&gt;sensor_synchronization&lt;/code&gt; and &lt;code&gt;joint_limit&lt;/code&gt; were N/A the whole run — this blind test doesn't exercise them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;121 of the 156 "clean" controls also got REVIEW&lt;/strong&gt; (all from &lt;code&gt;idle_ratio&lt;/code&gt;). This is exactly why the tool reports distributions instead of bare pass rates — on a task like PushT, a bare pass rate would look terrible for reasons that have nothing to do with data quality.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why this matters beyond one tool
&lt;/h2&gt;

&lt;p&gt;Robot learning is hitting a data-quality wall: more labs are collecting manipulation data, and "the training set had silently corrupt episodes" is the kind of failure you discover &lt;em&gt;after&lt;/em&gt; the policy fails to transfer. Blind-testing your audit tooling against a dataset your audience knows — with published confusion matrices and published weaknesses — is a much better baseline than vendor benchmarks with no ground truth. The full method, manifest schema, and per-episode verdicts are in the repo; the injection tool will ship as &lt;code&gt;rda blindtest&lt;/code&gt; in a future release.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;robot-data-audit&lt;span class="o"&gt;==&lt;/span&gt;0.5.4
&lt;span class="c"&gt;# download lerobot/pusht (v3.0), apply the 5×10 injection with seed=42&lt;/span&gt;
rda audit &amp;lt;blindtest_dataset&amp;gt; &lt;span class="nt"&gt;--format&lt;/span&gt; json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Full experiment doc (tables, per-class detectors, caveats): &lt;a href="https://github.com/liesliy/rda/blob/main/docs/blind_test_20260831.md" rel="noopener noreferrer"&gt;docs/blind_test_20260831.md&lt;/a&gt; · Benchmark: &lt;a href="https://github.com/liesliy/rda/blob/main/docs/benchmark.md" rel="noopener noreferrer"&gt;docs/benchmark.md&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;RDA is MIT-licensed, runs fully locally — your data never leaves the machine.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;If you maintain a robot dataset and want it audited (or blind-tested) as a public benchmark entry, open an issue on &lt;a href="https://github.com/liesliy/rda" rel="noopener noreferrer"&gt;the repo&lt;/a&gt; — the benchmark grows one dataset at a time.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>robotics</category>
      <category>database</category>
      <category>opensource</category>
      <category>ai</category>
    </item>
    <item>
      <title>We Need a Unicode for Tactile Data — Here's Why I Built One</title>
      <dc:creator>liesliy</dc:creator>
      <pubDate>Tue, 25 Aug 2026 07:13:42 +0000</pubDate>
      <link>https://dev.to/liesliy/we-need-a-unicode-for-tactile-data-heres-why-i-built-one-3ce8</link>
      <guid>https://dev.to/liesliy/we-need-a-unicode-for-tactile-data-heres-why-i-built-one-3ce8</guid>
      <description>&lt;p&gt;I've spent the last few years working with tactile sensors in robotics. And if there's one thing that drives me crazy, it's this: &lt;em&gt;every sensor speaks its own language, and nobody's translating.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;GelSight gives you image sequences. BioTac streams 19-channel impedance data. PaXini sends you 8×8 taxel grids. If you work with two different sensors in the same project, you end up writing two completely different data pipelines. Switch robots? Throw away your labels and start over.&lt;/p&gt;

&lt;p&gt;This isn't just annoying — it's actively holding back embodied AI. We're trying to train foundation models on tactile data, but the data is fragmented across dozens of incompatible formats. It's like trying to build GPT when every training corpus uses a different encoding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;So what if tactile data had... Unicode?&lt;/strong&gt;&lt;br&gt;
That's the idea behind TLabel — a unified annotation schema for tactile data across all sensor types.&lt;/p&gt;

&lt;p&gt;Think about what Unicode did for text. Before Unicode, you had ASCII, Shift-JIS, Latin-1, GB2312... the same character, different encodings, constant headaches. Unicode didn't replace any of them — it gave them a shared identity. A Chinese character is a Chinese character, regardless of how your system stores it.&lt;/p&gt;

&lt;p&gt;TLabel does the same thing for touch. It defines 14 semantic dimensions that any tactile sensor can map to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Did something touch? (contact)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Where on the sensor? (contact_region, contact_centroid)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How hard? (force_magnitude)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Which direction? (force_vector)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Is it slipping? (slip_event)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;What's the texture? (texture, friction)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Is the object deforming? (deformation_rate)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;What phase of manipulation? (approach, grasp, hold, release)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;And a few more...&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key insight: not every sensor can annotate all 14 dimensions. And that's fine. A simple force sensor gives you contact + force (Level 2). A vision-based sensor like GelSight can give you spatial info too (Level 3). Nobody's forcing you to pretend your sensor can do more than it can.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;tlabel
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. Three words.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why this matters now&lt;/strong&gt;&lt;br&gt;
We're at a weird moment in tactile sensing. The hardware is getting good — there are dozens of capable sensors now. The ML models are getting better. But the data layer is still stuck in the dark ages.&lt;/p&gt;

&lt;p&gt;Every lab has their own scripts. Every paper has its own format. Every startup that tries to build something on top of tactile data spends 3 months just writing data adapters before they can do anything interesting.&lt;/p&gt;

&lt;p&gt;I built TLabel because I was tired of writing the same adapter code for the fifth time. And I think we've reached the point where the robotics community needs this — the same way the web community needed HTTP, the same way the text community needed Unicode.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it actually looks like&lt;/strong&gt;&lt;br&gt;
Here's the real test: can you load data from a GelSight, a BioTac, and a PaXini, and get the same annotation structure back?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;tlabel&lt;/span&gt;

&lt;span class="c1"&gt;# These all return the same TLabelData structure
&lt;/span&gt;&lt;span class="n"&gt;gelsight_data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tlabel&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gel_sight_recording.pkl&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;biotac_data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tlabel&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;biotac_recording.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;paxini_data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tlabel&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;paxini_recording.mat&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Same schema. Same dimensions. Different sensors.
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;gelsight_data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;frames&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;schema_v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contact&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;biotac_data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;frames&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;schema_v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contact&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the whole point. The downstream code — your model training, your quality checks, your data export — doesn't need to know which sensor it came from.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where we are&lt;/strong&gt;&lt;br&gt;
TLabel is at v0.20.1 with 13 sensor adapters already built. It supports export to FTP-1 Zarr (for foundation models) and LeRobot format. It has a CLI, a quality scoring system, and a compliance level framework.&lt;/p&gt;

&lt;p&gt;It's MIT licensed, actively maintained, and we're actively looking for collaborators — especially if you work with sensors we haven't covered yet.&lt;/p&gt;

&lt;p&gt;👉 GitHub: github.com/liesliy/tlabel&lt;br&gt;
 · PyPI: pypi.org/project/tlabel&lt;/p&gt;

&lt;p&gt;If you're working with tactile data and you're tired of writing the same adapter code for the fifth time — give it a shot. Or even just tell us what you think. The whole point of a standard is that it only works if people use it.&lt;/p&gt;

&lt;p&gt;Built by the TouchLabel AI team. Open source, MIT license. Questions? Hit us up on GitHub Issues or drop a comment.&lt;/p&gt;

</description>
      <category>robotics</category>
      <category>opensource</category>
      <category>data</category>
      <category>tooling</category>
    </item>
    <item>
      <title>Robot Training Data Is Messier Than You Think: Auditing 4,959 Episodes with an Open-Source Tool</title>
      <dc:creator>liesliy</dc:creator>
      <pubDate>Fri, 21 Aug 2026 08:48:10 +0000</pubDate>
      <link>https://dev.to/liesliy/robot-training-data-is-messier-than-you-think-auditing-4959-episodes-with-an-open-source-tool-20ke</link>
      <guid>https://dev.to/liesliy/robot-training-data-is-messier-than-you-think-auditing-4959-episodes-with-an-open-source-tool-20ke</guid>
      <description>&lt;p&gt;Robot learning is getting better very quickly.&lt;/p&gt;

&lt;p&gt;We now have better policy architectures, more capable simulation environments, standardized dataset formats such as LeRobot, and an increasing number of public robot manipulation datasets.&lt;/p&gt;

&lt;p&gt;But there is a basic question that is surprisingly difficult to answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How do we know whether a robot training dataset is actually usable?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A dataset can be perfectly readable and still be a poor training dataset.&lt;/p&gt;

&lt;p&gt;It can have valid Parquet files, correct schemas, and complete metadata while containing excessive idle motion, action discontinuities, sampling problems, distribution anomalies, or episodes that deserve human review.&lt;/p&gt;

&lt;p&gt;I built &lt;strong&gt;RDA (Robot Data Audit)&lt;/strong&gt; to explore this problem.&lt;/p&gt;

&lt;p&gt;I recently ran it across &lt;strong&gt;12 local robot datasets and 4,959 episodes&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Here is what I found.&lt;/p&gt;




&lt;h2&gt;
  
  
  "Can I Load It?" Is Not the Same as "Can I Train on It?"
&lt;/h2&gt;

&lt;p&gt;Robot datasets can contain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;camera observations&lt;/li&gt;
&lt;li&gt;joint states&lt;/li&gt;
&lt;li&gt;actions&lt;/li&gt;
&lt;li&gt;timestamps&lt;/li&gt;
&lt;li&gt;task metadata&lt;/li&gt;
&lt;li&gt;robot configuration&lt;/li&gt;
&lt;li&gt;multiple sensor streams&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most data pipelines start with a simple question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can I load the dataset?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's necessary, but it isn't enough.&lt;/p&gt;

&lt;p&gt;I think robot data quality needs to be considered in several layers:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Structural integrity&lt;/td&gt;
&lt;td&gt;Does the data actually exist and load correctly?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Schema &amp;amp; temporal integrity&lt;/td&gt;
&lt;td&gt;Are fields, values, and timestamps valid?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Behavioral quality&lt;/td&gt;
&lt;td&gt;Does the robot motion contain suspicious patterns?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Task relevance&lt;/td&gt;
&lt;td&gt;Are those patterns actually harmful for this task?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;RDA currently focuses primarily on the first three layers and produces structured evidence that can be used to investigate the fourth.&lt;/p&gt;

&lt;p&gt;This distinction matters.&lt;/p&gt;

&lt;p&gt;A statistical anomaly is not automatically a failed demonstration.&lt;/p&gt;

&lt;p&gt;And a dataset that passes structural validation is not automatically good training data.&lt;/p&gt;




&lt;h1&gt;
  
  
  What RDA Actually Measures
&lt;/h1&gt;

&lt;p&gt;RDA is an open-source, local-first auditing tool for LeRobot-format robot manipulation datasets.&lt;/p&gt;

&lt;p&gt;The core audit runs locally and does not require a GPU.&lt;/p&gt;

&lt;p&gt;The current version checks 13 metrics across three layers.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 1 — Structural Integrity
&lt;/h3&gt;

&lt;p&gt;Deterministic checks include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;missing frames&lt;/li&gt;
&lt;li&gt;NaN / Inf values&lt;/li&gt;
&lt;li&gt;schema consistency&lt;/li&gt;
&lt;li&gt;timestamp validity&lt;/li&gt;
&lt;li&gt;joint limits&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Hard structural failures can result in an &lt;code&gt;EXCLUDE&lt;/code&gt; verdict.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 2 — Temporal &amp;amp; Motion Quality
&lt;/h3&gt;

&lt;p&gt;These checks look for statistical risk signals such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;sensor synchronization&lt;/li&gt;
&lt;li&gt;sampling jitter&lt;/li&gt;
&lt;li&gt;velocity / acceleration anomalies&lt;/li&gt;
&lt;li&gt;action discontinuity&lt;/li&gt;
&lt;li&gt;temporal sufficiency&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are generally signals for investigation rather than proof that an episode is unusable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 3 — Dataset Utility
&lt;/h3&gt;

&lt;p&gt;This includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;idle ratio&lt;/li&gt;
&lt;li&gt;distribution characteristics&lt;/li&gt;
&lt;li&gt;state-space coverage&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These metrics are particularly contextual.&lt;/p&gt;

&lt;p&gt;For example, a high idle ratio may be perfectly reasonable for one task and problematic for another.&lt;/p&gt;




&lt;h1&gt;
  
  
  PASS, REVIEW, EXCLUDE — But Don't Confuse Them with Ground Truth
&lt;/h1&gt;

&lt;p&gt;One design decision became particularly important during development.&lt;/p&gt;

&lt;p&gt;RDA separates &lt;strong&gt;workflow verdicts&lt;/strong&gt; from &lt;strong&gt;evidence levels&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The workflow verdict is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;PASS&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;REVIEW&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;EXCLUDE&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The evidence level is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;HARD_FAIL&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;RISK_SIGNAL&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;UNVERIFIABLE&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are deliberately different concepts.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;High action discontinuity → &lt;code&gt;RISK_SIGNAL&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;does &lt;strong&gt;not&lt;/strong&gt; mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The demonstration is corrupted.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Likewise:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Missing frames → &lt;code&gt;HARD_FAIL&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;does not necessarily mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The entire upstream dataset is bad.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The purpose is to keep an audit tool from turning weak statistical evidence into overly confident labels.&lt;/p&gt;




&lt;h1&gt;
  
  
  4,959 Episodes: Almost Half Triggered a Review Signal
&lt;/h1&gt;

&lt;p&gt;I ran RDA across 12 local dataset copies.&lt;/p&gt;

&lt;p&gt;The overall result was:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;th&gt;Episodes&lt;/th&gt;
&lt;th&gt;Percentage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;PASS&lt;/td&gt;
&lt;td&gt;1,711&lt;/td&gt;
&lt;td&gt;34.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;REVIEW&lt;/td&gt;
&lt;td&gt;2,473&lt;/td&gt;
&lt;td&gt;49.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;EXCLUDE&lt;/td&gt;
&lt;td&gt;775&lt;/td&gt;
&lt;td&gt;15.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These numbers are &lt;strong&gt;not&lt;/strong&gt; a confusion matrix.&lt;/p&gt;

&lt;p&gt;There was no independent human ground truth for this benchmark, so they cannot be used to claim precision or recall.&lt;/p&gt;

&lt;p&gt;The interesting observation is simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A large fraction of the tested episodes contained signals worth investigating before blindly feeding them into training.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That alone is useful.&lt;/p&gt;




&lt;h1&gt;
  
  
  Finding #1: Idle Data Is Everywhere
&lt;/h1&gt;

&lt;p&gt;The median idle ratio across the tested datasets ranged from:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;20.8% to 93.3%.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ten of the twelve datasets had median idle ratios above 65%.&lt;/p&gt;

&lt;p&gt;That raises a surprisingly important question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If a large fraction of your training data represents the robot doing very little, what exactly is the policy learning?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It might be useful temporal context.&lt;/p&gt;

&lt;p&gt;Or it might simply teach the model that doing nothing is common.&lt;/p&gt;

&lt;p&gt;The answer depends heavily on the task and model architecture.&lt;/p&gt;

&lt;p&gt;This led to our first optimization experiment.&lt;/p&gt;




&lt;h1&gt;
  
  
  Finding #2: The Same Robot Can Have Very Different Data Characteristics
&lt;/h1&gt;

&lt;p&gt;Two datasets from the same xArm platform showed a large difference.&lt;/p&gt;

&lt;h3&gt;
  
  
  xArm lift
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Median idle ratio: &lt;strong&gt;20.8%&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;767 / 800 episodes received PASS&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  xArm push
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Median idle ratio: &lt;strong&gt;83.3%&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;562 / 800 episodes received REVIEW&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same robot platform.&lt;/p&gt;

&lt;p&gt;Different task.&lt;/p&gt;

&lt;p&gt;The important lesson is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Data quality is not completely independent of task structure.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A high idle ratio is not automatically a collection failure.&lt;/p&gt;

&lt;p&gt;It may simply reflect how the task is performed.&lt;/p&gt;

&lt;p&gt;That makes universal rules like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Delete every episode above 70% idle."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;dangerous.&lt;/p&gt;




&lt;h1&gt;
  
  
  Finding #3: Action Discontinuity May Reveal the Controller
&lt;/h1&gt;

&lt;p&gt;Another interesting pattern appeared in action discontinuity.&lt;/p&gt;

&lt;p&gt;Some datasets showed spikes in essentially every episode.&lt;/p&gt;

&lt;p&gt;Another xArm dataset had only a handful across 800 episodes.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;simulated ALOHA datasets showed frequent spikes&lt;/li&gt;
&lt;li&gt;the SO-100 dataset also showed frequent spikes&lt;/li&gt;
&lt;li&gt;xArm lift had only 6 spike-containing episodes out of 800&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This suggests that some data-quality signals may tell us as much about the &lt;strong&gt;collection or control stack&lt;/strong&gt; as about the demonstrations themselves.&lt;/p&gt;

&lt;p&gt;That could matter for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;smoothness-regularized policies&lt;/li&gt;
&lt;li&gt;sim-to-real transfer&lt;/li&gt;
&lt;li&gt;controller comparisons&lt;/li&gt;
&lt;li&gt;teleoperation system evaluation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But again:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A discontinuity is a signal, not automatically a failure.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Finding #4: A Dataset Can Be Structurally Clean and Still Need Review
&lt;/h1&gt;

&lt;p&gt;One of the most useful observations from the benchmark was the separation between integrity and behavior.&lt;/p&gt;

&lt;p&gt;Outside of the intentionally corrupted ALOHA fixture, the tested datasets were largely clean at the basic integrity layer.&lt;/p&gt;

&lt;p&gt;Yet the behavioral layer still flagged many episodes for review.&lt;/p&gt;

&lt;p&gt;In other words:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"The files are valid" and "the demonstrations are ideal for training" are different questions.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This seems obvious after writing it down.&lt;/p&gt;

&lt;p&gt;In practice, however, many pipelines stop at the first question.&lt;/p&gt;




&lt;h1&gt;
  
  
  Finding #5: Sometimes the Dataset Copy Is the Problem
&lt;/h1&gt;

&lt;p&gt;One of the most interesting cases was LIBERO.&lt;/p&gt;

&lt;p&gt;The local metadata declared:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1,693 episodes&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;But only:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;920 episodes&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;could actually be read from the local Parquet files.&lt;/p&gt;

&lt;p&gt;The remaining:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;773 episodes&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;were missing from the local copy.&lt;/p&gt;

&lt;p&gt;An earlier loader implementation could interpret this kind of metadata/data-layout mismatch as zero-frame episodes.&lt;/p&gt;

&lt;p&gt;The loader was changed to scan the actual Parquet layout and build a fallback episode index.&lt;/p&gt;

&lt;p&gt;The result was more honest:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The local copy is incomplete.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;LIBERO contains 773 empty episodes.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This distinction matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A broken local copy is not necessarily a broken upstream dataset.&lt;/strong&gt;&lt;/p&gt;




&lt;h1&gt;
  
  
  Then I Compared RDA with Other Tools
&lt;/h1&gt;

&lt;p&gt;I also wanted to understand how different robot-data tooling behaves on deliberately corrupted data.&lt;/p&gt;

&lt;p&gt;So I ran a small cross-tool benchmark involving:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;RDA&lt;/li&gt;
&lt;li&gt;trajlens&lt;/li&gt;
&lt;li&gt;ORBIT&lt;/li&gt;
&lt;li&gt;lerobot-doctor&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The test included known injected defects.&lt;/p&gt;

&lt;p&gt;One result looked like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Defect&lt;/th&gt;
&lt;th&gt;RDA&lt;/th&gt;
&lt;th&gt;trajlens&lt;/th&gt;
&lt;th&gt;ORBIT&lt;/th&gt;
&lt;th&gt;lerobot-doctor&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;NaN actions&lt;/td&gt;
&lt;td&gt;EXCLUDE&lt;/td&gt;
&lt;td&gt;WARN&lt;/td&gt;
&lt;td&gt;Crash&lt;/td&gt;
&lt;td&gt;FAIL&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Timestamp reversal&lt;/td&gt;
&lt;td&gt;EXCLUDE&lt;/td&gt;
&lt;td&gt;FAIL&lt;/td&gt;
&lt;td&gt;Crash&lt;/td&gt;
&lt;td&gt;FAIL&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frozen actions&lt;/td&gt;
&lt;td&gt;PASS&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;FAIL&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The important result wasn't that RDA "won."&lt;/p&gt;

&lt;p&gt;It didn't.&lt;/p&gt;

&lt;p&gt;The frozen-action case exposed a real blind spot.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Frozen-Segment Blind Spot
&lt;/h1&gt;

&lt;p&gt;RDA's current &lt;code&gt;idle_ratio&lt;/code&gt; metric is primarily episode-level.&lt;/p&gt;

&lt;p&gt;That means a dataset can have a reasonable overall idle ratio while still containing a long locally frozen segment.&lt;/p&gt;

&lt;p&gt;Another tool had a dedicated consecutive-identical-actions check and caught it.&lt;/p&gt;

&lt;p&gt;This gives RDA a very concrete next feature:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;frozen_segment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A good audit tool should be able to identify its own blind spots.&lt;/p&gt;

&lt;p&gt;This one is now on our priority list.&lt;/p&gt;




&lt;h1&gt;
  
  
  Can Data Auditing Actually Improve Training?
&lt;/h1&gt;

&lt;p&gt;Finding suspicious data is only half the problem.&lt;/p&gt;

&lt;p&gt;The more interesting question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Can audit results tell us what to do with the dataset?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I ran several experiments around idle-data pruning.&lt;/p&gt;

&lt;p&gt;The result was much more nuanced than I initially expected.&lt;/p&gt;




&lt;h1&gt;
  
  
  Moderate Pruning Helped — Aggressive Pruning Hurt
&lt;/h1&gt;

&lt;p&gt;In one experiment:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;Data retained&lt;/th&gt;
&lt;th&gt;MSE change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Baseline&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Keep 50%&lt;/td&gt;
&lt;td&gt;50.1%&lt;/td&gt;
&lt;td&gt;-28.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Keep 40%&lt;/td&gt;
&lt;td&gt;40.1%&lt;/td&gt;
&lt;td&gt;-12.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Keep 30%&lt;/td&gt;
&lt;td&gt;30.1%&lt;/td&gt;
&lt;td&gt;-13.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Keep 20%&lt;/td&gt;
&lt;td&gt;20.0%&lt;/td&gt;
&lt;td&gt;+11.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Keep 10%&lt;/td&gt;
&lt;td&gt;10.0%&lt;/td&gt;
&lt;td&gt;+98.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;There was a clear dose-response pattern.&lt;/p&gt;

&lt;p&gt;Moderate pruning helped in this experiment.&lt;/p&gt;

&lt;p&gt;Aggressive pruning eventually became harmful.&lt;/p&gt;

&lt;p&gt;So the conclusion is not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Remove idle data."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;There may be a useful pruning region, but it depends on how the remaining data is used.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Then We Tested Different Model Architectures
&lt;/h1&gt;

&lt;p&gt;This was probably the most interesting result.&lt;/p&gt;

&lt;p&gt;We compared a frame-wise MLP with a temporal Transformer.&lt;/p&gt;

&lt;p&gt;The result was almost the opposite.&lt;/p&gt;

&lt;h3&gt;
  
  
  MLP
&lt;/h3&gt;

&lt;p&gt;Idle-data pruning:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;-21.6% MSE&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Temporal Transformer
&lt;/h3&gt;

&lt;p&gt;Idle-data pruning:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;+95.5% MSE&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because temporal models need context.&lt;/p&gt;

&lt;p&gt;Suppose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;idle ratio ≈ 70%&lt;/li&gt;
&lt;li&gt;active segments are scattered&lt;/li&gt;
&lt;li&gt;sequence length = 10&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After pruning, many active segments become shorter than the required temporal window.&lt;/p&gt;

&lt;p&gt;The result:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You removed "low-value" frames and accidentally removed the context required to construct training samples.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There is another issue.&lt;/p&gt;

&lt;p&gt;A temporal policy may need to learn the transition:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;idle → movement&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If all idle context is removed, the model may never see how movement starts.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Practical Lesson
&lt;/h1&gt;

&lt;p&gt;This changed how I think about automated dataset optimization.&lt;/p&gt;

&lt;p&gt;A rule like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;idle_ratio &amp;gt; 70% → delete&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;is too simplistic.&lt;/p&gt;

&lt;p&gt;The better question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What model are you training, and what temporal structure does that model need?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The current experiments suggest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;gentle trimming can be reasonable for frame-wise models&lt;/li&gt;
&lt;li&gt;aggressive idle pruning can be harmful&lt;/li&gt;
&lt;li&gt;temporal models need special handling&lt;/li&gt;
&lt;li&gt;dataset/task domain matters&lt;/li&gt;
&lt;li&gt;audit signals should not automatically become deletion rules&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is why RDA's recommendation layer takes the intended model type into account.&lt;/p&gt;




&lt;h1&gt;
  
  
  A Local-First Architecture
&lt;/h1&gt;

&lt;p&gt;Robot datasets can contain commercially sensitive information.&lt;/p&gt;

&lt;p&gt;They may reveal:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;robot configurations&lt;/li&gt;
&lt;li&gt;manipulation strategies&lt;/li&gt;
&lt;li&gt;factory environments&lt;/li&gt;
&lt;li&gt;camera observations&lt;/li&gt;
&lt;li&gt;proprietary tasks&lt;/li&gt;
&lt;li&gt;operator behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;RDA's core audit therefore runs locally.&lt;/p&gt;

&lt;p&gt;It can generate a blind report containing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;anonymized paths&lt;/li&gt;
&lt;li&gt;dataset statistics&lt;/li&gt;
&lt;li&gt;episode/frame counts&lt;/li&gt;
&lt;li&gt;metrics&lt;/li&gt;
&lt;li&gt;verdicts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;without sending the raw trajectories or images.&lt;/p&gt;

&lt;p&gt;The goal is simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Share evidence without sharing the dataset.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  What RDA Does NOT Claim
&lt;/h1&gt;

&lt;p&gt;This is probably the most important section.&lt;/p&gt;

&lt;p&gt;The current experiments do &lt;strong&gt;not&lt;/strong&gt; establish:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;RDA precision&lt;/li&gt;
&lt;li&gt;RDA recall&lt;/li&gt;
&lt;li&gt;universally optimal thresholds&lt;/li&gt;
&lt;li&gt;that every &lt;code&gt;EXCLUDE&lt;/code&gt; episode should be deleted&lt;/li&gt;
&lt;li&gt;that action discontinuity means corruption&lt;/li&gt;
&lt;li&gt;that high idle ratio means useless data&lt;/li&gt;
&lt;li&gt;that RDA improves training success rate&lt;/li&gt;
&lt;li&gt;that RDA certifies compliance with any data-quality standard&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The benchmark is based on real local dataset runs, but it does not have independent ground-truth labels.&lt;/p&gt;

&lt;p&gt;So the next experiment is obvious:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Get real human or customer QC labels and compare them against RDA without exposing the RDA verdict first.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  The Next Experiment: Blind Human Review
&lt;/h1&gt;

&lt;p&gt;The proposed validation workflow is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Select 100–500 episodes from one robot platform.&lt;/li&gt;
&lt;li&gt;Generate anonymized samples.&lt;/li&gt;
&lt;li&gt;Have reviewers independently label them:&lt;/li&gt;
&lt;/ol&gt;

&lt;ul&gt;
&lt;li&gt;KEEP&lt;/li&gt;
&lt;li&gt;REVIEW&lt;/li&gt;
&lt;li&gt;REMOVE

&lt;ol&gt;
&lt;li&gt;Do not show them RDA's verdict.&lt;/li&gt;
&lt;li&gt;Preserve reviewer decisions and notes.&lt;/li&gt;
&lt;li&gt;Reveal RDA results only after labeling.&lt;/li&gt;
&lt;li&gt;Compare the two sets of evidence.&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That would allow us to measure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;overlap between RDA &lt;code&gt;HARD_FAIL&lt;/code&gt; and human &lt;code&gt;REMOVE&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;overlap between RDA &lt;code&gt;RISK_SIGNAL&lt;/code&gt; and human &lt;code&gt;REVIEW&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;additional defects found by RDA&lt;/li&gt;
&lt;li&gt;reviewer time per episode&lt;/li&gt;
&lt;li&gt;review queue reduction&lt;/li&gt;
&lt;li&gt;potential downstream training impact&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Until that experiment is completed, these remain validation targets rather than marketing claims.&lt;/p&gt;




&lt;h1&gt;
  
  
  Try It
&lt;/h1&gt;

&lt;p&gt;RDA is open source and currently supports LeRobot v2.1 and v3.0.&lt;/p&gt;

&lt;p&gt;Install:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;robot-data-audit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Audit a dataset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;rda audit /path/to/lerobot/dataset &lt;span class="nt"&gt;-v&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Generate a blind report:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;rda audit /path/to/lerobot/dataset &lt;span class="nt"&gt;--blind&lt;/span&gt; &lt;span class="nt"&gt;--format&lt;/span&gt; json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Generate optimization recommendations:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;rda recommend /path/to/dataset &lt;span class="nt"&gt;--policy&lt;/span&gt; frame-wise
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;rda recommend /path/to/dataset &lt;span class="nt"&gt;--policy&lt;/span&gt; temporal
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Launch the optional UI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;robot-data-audit[ui]
rda ui
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/liesliy/rda" rel="noopener noreferrer"&gt;https://github.com/liesliy/rda&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PyPI:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://pypi.org/project/robot-data-audit/" rel="noopener noreferrer"&gt;https://pypi.org/project/robot-data-audit/&lt;/a&gt;&lt;/p&gt;




&lt;h1&gt;
  
  
  What I Want to Test Next
&lt;/h1&gt;

&lt;p&gt;The next step isn't another synthetic benchmark.&lt;/p&gt;

&lt;p&gt;I want to test RDA against &lt;strong&gt;real robot data with independent QC labels&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In particular, I'm looking for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;real-world teleoperation datasets&lt;/li&gt;
&lt;li&gt;datasets with existing QC labels&lt;/li&gt;
&lt;li&gt;datasets containing known failure cases&lt;/li&gt;
&lt;li&gt;datasets used for sim-to-real experiments&lt;/li&gt;
&lt;li&gt;robot companies or research groups willing to run a local audit&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The raw data does not need to leave your environment.&lt;/p&gt;

&lt;p&gt;A small validation set — even &lt;strong&gt;100–500 episodes from one robot and 1–3 tasks&lt;/strong&gt; — would already be extremely useful.&lt;/p&gt;

&lt;p&gt;If you work with robot manipulation data and are willing to challenge these results, I'd love to compare notes.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Takeaway
&lt;/h1&gt;

&lt;p&gt;After auditing thousands of robot episodes, my conclusion is not that robot datasets are "bad."&lt;/p&gt;

&lt;p&gt;It is more subtle:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Robot datasets contain many properties that are invisible to a simple "can I load this file?" check.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And those properties can matter differently depending on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the robot&lt;/li&gt;
&lt;li&gt;the task&lt;/li&gt;
&lt;li&gt;the controller&lt;/li&gt;
&lt;li&gt;the dataset format&lt;/li&gt;
&lt;li&gt;the policy architecture&lt;/li&gt;
&lt;li&gt;the intended use of the data&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal of RDA is not to produce another mysterious quality score.&lt;/p&gt;

&lt;p&gt;It is to make the path from:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;measurement → diagnosis → human review → optimization&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;more reproducible.&lt;/p&gt;

&lt;p&gt;And eventually, hopefully, to make robot data quality something we can discuss with evidence rather than intuition.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before the GPU bill arrives, audit the data.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>robotics</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>dataset</category>
    </item>
    <item>
      <title># Public Robot Dataset Health Check — as measured by RDA v0.5.2</title>
      <dc:creator>liesliy</dc:creator>
      <pubDate>Wed, 19 Aug 2026 03:30:56 +0000</pubDate>
      <link>https://dev.to/liesliy/-public-robot-dataset-health-check-as-measured-by-rda-v052-1ge8</link>
      <guid>https://dev.to/liesliy/-public-robot-dataset-health-check-as-measured-by-rda-v052-1ge8</guid>
      <description>&lt;p&gt;We ran RDA across &lt;strong&gt;11 public LeRobot-format datasets — 4,909 episodes&lt;/strong&gt; in total, spanning sim and real, scripted and human teleop, research arms and $100 hobby hardware. Summary:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;  
  &lt;thead&gt;  
    &lt;tr&gt;  
      &lt;th&gt;Dataset&lt;/th&gt;  
      &lt;th&gt;Type&lt;/th&gt;  
      &lt;th&gt;Eps&lt;/th&gt;  
      &lt;th&gt;Verdicts (P/R/E)&lt;/th&gt;  
      &lt;th&gt;Action spikes&lt;/th&gt;  
      &lt;th&gt;Idle median&lt;/th&gt;  
      &lt;th&gt;Idle p5&lt;/th&gt;  
    &lt;/tr&gt;  
  &lt;/thead&gt;  
  &lt;tbody&gt;  
    &lt;tr&gt;  
      &lt;td&gt;aloha_sim_insertion_human&lt;/td&gt;  
      &lt;td&gt;sim, human&lt;/td&gt;  
      &lt;td&gt;50&lt;/td&gt;  
      &lt;td&gt;5 / 45 / 0&lt;/td&gt;  
      &lt;td&gt;1,338 (100% eps)&lt;/td&gt;  
      &lt;td&gt;70.7%&lt;/td&gt;  
      &lt;td&gt;65.1%&lt;/td&gt;  
    &lt;/tr&gt;  
    &lt;tr&gt;  
      &lt;td&gt;aloha_sim_transfer_cube_scripted&lt;/td&gt;  
      &lt;td&gt;sim, scripted&lt;/td&gt;  
      &lt;td&gt;50&lt;/td&gt;  
      &lt;td&gt;5 / 45 / 0&lt;/td&gt;  
      &lt;td&gt;2,489 (100% eps)&lt;/td&gt;  
      &lt;td&gt;64.0%&lt;/td&gt;  
      &lt;td&gt;44.2%&lt;/td&gt;  
    &lt;/tr&gt;  
    &lt;tr&gt;  
      &lt;td&gt;aloha_sim_insertion_scripted&lt;/td&gt;  
      &lt;td&gt;sim, scripted&lt;/td&gt;  
      &lt;td&gt;50&lt;/td&gt;  
      &lt;td&gt;1 / 49 / 0&lt;/td&gt;  
      &lt;td&gt;1,633 (100% eps)&lt;/td&gt;  
      &lt;td&gt;63.7%&lt;/td&gt;  
      &lt;td&gt;53.7%&lt;/td&gt;  
    &lt;/tr&gt;  
    &lt;tr&gt;  
      &lt;td&gt;droid_100&lt;/td&gt;  
      &lt;td&gt;real Franka&lt;/td&gt;  
      &lt;td&gt;100&lt;/td&gt;  
      &lt;td&gt;34 / 66 / 0&lt;/td&gt;  
      &lt;td&gt;1,428 (99% eps)&lt;/td&gt;  
      &lt;td&gt;70.7%&lt;/td&gt;  
      &lt;td&gt;55.7%&lt;/td&gt;  
    &lt;/tr&gt;  
    &lt;tr&gt;  
      &lt;td&gt;pusht&lt;/td&gt;  
      &lt;td&gt;sim&lt;/td&gt;  
      &lt;td&gt;206&lt;/td&gt;  
      &lt;td&gt;43 / 163 / 0&lt;/td&gt;  
      &lt;td&gt;1,148 (97% eps)&lt;/td&gt;  
      &lt;td&gt;81.7%&lt;/td&gt;  
      &lt;td&gt;59.5%&lt;/td&gt;  
    &lt;/tr&gt;  
    &lt;tr&gt;  
      &lt;td&gt;HuggingFaceVLA/libero&lt;/td&gt;  
      &lt;td&gt;sim&lt;/td&gt;  
      &lt;td&gt;1,693&lt;/td&gt;  
      &lt;td&gt;0 / 3 / 1,690&lt;/td&gt;  
      &lt;td&gt;34 (3 eps)&lt;/td&gt;  
      &lt;td&gt;76.5% (3 eps)&lt;/td&gt;  
      &lt;td&gt;74.6%&lt;/td&gt;  
    &lt;/tr&gt;  
    &lt;tr&gt;  
      &lt;td&gt;bridge_orig_lerobot (sampled)&lt;/td&gt;  
      &lt;td&gt;real WidowX&lt;/td&gt;  
      &lt;td&gt;25&lt;/td&gt;  
      &lt;td&gt;4 / 21 / 0&lt;/td&gt;  
      &lt;td&gt;91 (84% eps)&lt;/td&gt;  
      &lt;td&gt;&lt;strong&gt;93.3%&lt;/strong&gt;&lt;/td&gt;  
      &lt;td&gt;20.5%&lt;/td&gt;  
    &lt;/tr&gt;  
    &lt;tr&gt;  
      &lt;td&gt;xarm_lift_medium&lt;/td&gt;  
      &lt;td&gt;real xArm&lt;/td&gt;  
      &lt;td&gt;800&lt;/td&gt;  
      &lt;td&gt;
&lt;strong&gt;767&lt;/strong&gt; / 33 / 0&lt;/td&gt;  
      &lt;td&gt;6 (1% eps)&lt;/td&gt;  
      &lt;td&gt;&lt;strong&gt;20.8%&lt;/strong&gt;&lt;/td&gt;  
      &lt;td&gt;12.5%&lt;/td&gt;  
    &lt;/tr&gt;  
    &lt;tr&gt;  
      &lt;td&gt;xarm_push_medium&lt;/td&gt;  
      &lt;td&gt;real xArm&lt;/td&gt;  
      &lt;td&gt;800&lt;/td&gt;  
      &lt;td&gt;238 / 562 / 0&lt;/td&gt;  
      &lt;td&gt;845 (62% eps)&lt;/td&gt;  
      &lt;td&gt;83.3%&lt;/td&gt;  
      &lt;td&gt;16.7%&lt;/td&gt;  
    &lt;/tr&gt;  
    &lt;tr&gt;  
      &lt;td&gt;svla_so101_pickplace&lt;/td&gt;  
      &lt;td&gt;real SO-100&lt;/td&gt;  
      &lt;td&gt;50&lt;/td&gt;  
      &lt;td&gt;5 / 45 / 0&lt;/td&gt;  
      &lt;td&gt;260 (100% eps)&lt;/td&gt;  
      &lt;td&gt;86.7%&lt;/td&gt;  
      &lt;td&gt;59.5%&lt;/td&gt;  
    &lt;/tr&gt;  
    &lt;tr&gt;  
      &lt;td&gt;jaco_play&lt;/td&gt;  
      &lt;td&gt;real Jaco&lt;/td&gt;  
      &lt;td&gt;1,085&lt;/td&gt;  
      &lt;td&gt;390 / 695 / 0&lt;/td&gt;  
      &lt;td&gt;11,958 (77% eps)&lt;/td&gt;  
      &lt;td&gt;74.1%&lt;/td&gt;  
      &lt;td&gt;52.7%&lt;/td&gt;  
    &lt;/tr&gt;  
  &lt;/tbody&gt;  
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;All integrity layers (NaN/Inf, timestamp validity, missing frames, schema) came back clean on all 11. libero note: 1,690 of 1,693 episodes read 0 frames — a dataset-side meta/layout mismatch, now flagged EXCLUDE instead of silently passing. bridge sampled 25 episodes; the other 10 datasets were audited in full. (&lt;code&gt;aloha_sim_transfer_cube_human&lt;/code&gt; was audited earlier at v0.5.1 with matching results: 1/49/0, 1,535 spikes, 71.2% idle — it's now a gated repo, so we couldn't re-pull it.)&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Five patterns worth knowing before you train
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Median idle runs 20.8%–93.3%, and 8 of 11 datasets sit above 65%.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Loss functions trained on a 75%-idle distribution are structurally biased toward predicting "do nothing" unless you weight or curriculum it. Bridge data pushes it to 93%. Measure yours before the GPU bill, not after.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Same robot, same lab, four-fold idle difference.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;code&gt;xarm_lift_medium&lt;/code&gt;: 20.8% median idle, 767/800 episodes PASS. &lt;code&gt;xarm_push_medium&lt;/code&gt;: 83.3% median idle, 562/800 REVIEW. Same xArm platform — the difference is task difficulty (lifting vs. pushing a flat object), not collection sloppiness. High idle isn't always a bug; it's a property you need to know and design around. RDA flags both sides of this honestly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Action discontinuity tracks the controller, not the dataset's reputation.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Sim ALOHA and the SO-100 hobby setup spike in literally 100% of episodes; xArm lift data has 6 spikes across 800 episodes. If your policy uses smoothness regularization or you're doing sim-to-real action statistics, this number decides your curriculum.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Clean integrity ≠ good training data.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Integrity passed 11/11 — zero NaNs, zero timestamp reversals, zero missing frames anywhere. The behavior layer still flagged 45–98% of episodes for review in most datasets. Both layers matter; most pipelines check neither.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Cheap hardware produces the most expensive data.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
The community SO-100 pick-place set: 86.7% median idle plus spikes in every episode. If you're fine-tuning on hobby-robot uploads, this is what you're inheriting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;robot-data-audit
rda audit &amp;lt;any lerobot dataset&amp;gt; &lt;span class="nt"&gt;-v&lt;/span&gt;
rda recommend &amp;lt;dataset&amp;gt; &lt;span class="nt"&gt;--policy&lt;/span&gt; temporal   &lt;span class="c"&gt;# or frame-wise, --lang en&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Tool: &lt;a href="https://github.com/liesliy/rda" rel="noopener noreferrer"&gt;https://github.com/liesliy/rda&lt;/a&gt; · PyPI: &lt;code&gt;robot-data-audit&lt;/code&gt; · UI: &lt;code&gt;rda ui&lt;/code&gt; (EN/中文)&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Caveats: RDA flags statistical anomalies, not ground-truth errors. REVIEW means "look before you train," not "discard." All thresholds are open for debate — that's what the issue tracker is for.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>robotics</category>
      <category>opensource</category>
      <category>python</category>
    </item>
    <item>
      <title>I Open-Sourced a Data Quality Auditor for Robot Datasets — It Found 1,535 Action Spikes in the Official ALOHA Demo Set</title>
      <dc:creator>liesliy</dc:creator>
      <pubDate>Wed, 19 Aug 2026 02:55:47 +0000</pubDate>
      <link>https://dev.to/liesliy/i-open-sourced-a-data-quality-auditor-for-robot-datasets-it-found-1535-action-spikes-in-the-24bj</link>
      <guid>https://dev.to/liesliy/i-open-sourced-a-data-quality-auditor-for-robot-datasets-it-found-1535-action-spikes-in-the-24bj</guid>
      <description>&lt;p&gt;Everyone in embodied AI says "data is the new code." Nobody audits it like code.&lt;/p&gt;

&lt;p&gt;We run linters, static analyzers, and CI gates on our source code. Then we feed 50GB of teleoperation recordings into a policy network and hope for the best. So I built &lt;strong&gt;RDA (Robot Data Audit)&lt;/strong&gt; — an open-source CLI that treats robot datasets the way &lt;code&gt;ruff&lt;/code&gt; treats a Python repo.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;pip install robot-data-audit
rda audit /path/to/lerobot/dataset

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It works natively on LeRobot-format datasets (v2.1 + v3.0), checks every episode across two layers — integrity (NaN actions, timestamp reversals, missing frames) and behavior (action discontinuity, idle ratio, frozen segments) — and outputs a per-episode verdict: PASS / REVIEW / EXCLUDE.&lt;/p&gt;

&lt;h3&gt;
  
  
  The dogfooding surprise
&lt;/h3&gt;

&lt;p&gt;Before releasing it, I ran RDA against the datasets everyone treats as ground truth — starting with &lt;code&gt;aloha_sim_transfer_cube_human&lt;/code&gt;, the official ALOHA simulation demonstration set that ships with the LeRobot ecosystem.&lt;/p&gt;

&lt;p&gt;Result from the behavior layer, across 50 episodes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Action discontinuity spikes: 1,535 (~30 per episode)&lt;/li&gt;
&lt;li&gt;Median idle ratio: 0.71** — the arm is effectively stationary 71% of frames&lt;/li&gt;
&lt;li&gt;Median effective motion ratio: 0.29&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To be clear about what this does and doesn't mean: none of this is corruption. The integrity layer came back clean — no NaNs, no broken timestamps. The spikes are real discontinuities in the action space (large frame-to-frame joint jumps), and the high idle ratio likely reflects grasping/hover phases where the gripper holds still. Neither is necessarily a bug in the dataset.&lt;/p&gt;

&lt;p&gt;But that's exactly the point. If you're benchmarking a policy on this data, or worse, fine-tuning on it, these numbers are &lt;strong&gt;context you didn't have&lt;/strong&gt;. Is 30 spikes per episode normal for this task? Does 71% idle time skew your loss toward predicting "do nothing"? Nobody asks, because nobody measures.&lt;/p&gt;

&lt;h3&gt;
  
  
  The bug I found in my own tool (while writing this post)
&lt;/h3&gt;

&lt;p&gt;Honesty section, because dev.to deserves better than marketing:&lt;/p&gt;

&lt;p&gt;My first run reported all 50 episodes as PASS. Green across the board. Celebration ensued.&lt;/p&gt;

&lt;p&gt;Then I cross-checked the behavior layer output against the verdicts and realized they weren't connected — the metrics were computing 1,535 spikes, and the verdict aggregator was ignoring behavior signals entirely. The tool had the evidence and wasn't reading it. The loudest silence in software is a metric that's computed but never consumed.&lt;/p&gt;

&lt;p&gt;Fixed now: behavior signals feed the verdict through a dataset-utility layer, and metric-level findings carry human-readable reasons. The ALOHA run now correctly flags 49/50 episodes as REVIEW with the specific signals attached.&lt;/p&gt;

&lt;p&gt;If you're building anything with a "signal producer → decision aggregator" architecture, test the wiring, not just the signals. I wrote a negative-control test for it before I trusted my own tool again.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why this matters more than it sounds
&lt;/h3&gt;

&lt;p&gt;Robot learning teams are drowning in data collection — teleop sessions, sim rollouts, fleet logs — with almost no tooling for "is this batch usable before I burn GPU hours on it." An episode with a frozen sensor or a corrupted timestamp doesn't fail loudly. It trains quietly.&lt;/p&gt;

&lt;p&gt;RDA's philosophy: audit before train. Cheap checks first (seconds per episode, pure numpy/pandas), verdicts you can gate in CI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;rda audit ./my_dataset &lt;span class="nt"&gt;--format&lt;/span&gt; json &lt;span class="nt"&gt;-o&lt;/span&gt; report.json
&lt;span class="c"&gt;# fail the pipeline if any episode comes back EXCLUDE&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There's also a built-in UI (&lt;code&gt;rda ui&lt;/code&gt;) for browsing verdicts without spelunking JSON — and as of v0.5.2 it's fully bilingual: one toggle switches the entire dashboard, backend recommendation copy included, between English and 中文.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxlav0kxbt13hu9yr024h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxlav0kxbt13hu9yr024h.png" alt=" " width="799" height="440"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What shipped since the first post
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;rda recommend&lt;/code&gt;&lt;/strong&gt; — model-aware optimization advice. Tell it whether you're training a frame-wise model (MLP/BC) or a temporal one (ACT/Diffusion Policy), and it gives different answers for the same data — including an explicit DO_NOT_PRUNE guard for temporal models when valid-window ratio collapses. Every suggestion carries its experimental evidence: pruning cost our seq=10 temporal baseline +296% MSE, while trimmed ALOHA/PushT improved frame-wise baselines by 11–35%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Privacy-first architecture&lt;/strong&gt; — metrics compute locally; only &amp;lt;1KB of aggregates reach the rules API. &lt;code&gt;rda audit&lt;/code&gt; stays 100% offline, always.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LeRobot v2.1 support&lt;/strong&gt; — bridge-style layouts now load natively (first run on bridge data: median idle ratio 93.3%. Real robots spend a lot of time deciding.).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A bug class worth naming: silent PASS.&lt;/strong&gt; 1,690 zero-frame episodes in a popular dataset were passing because "no evidence of problems" was treated as "no problems." Now zero-frame episodes are explicit EXCLUDEs with a diagnosis.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5zin3ec01e19yyxto69v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5zin3ec01e19yyxto69v.png" alt=" " width="799" height="440"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What's next
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;More behavior metrics (jitter, cycle anomalies, calibration drift)&lt;/li&gt;
&lt;li&gt;Trend dashboards across successive audits (already in the UI's History page — feedback wanted)&lt;/li&gt;
&lt;li&gt;Export-to-clean: one-click filtered dataset copy from surviving episodes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The project is early and hungry for real-world datasets to chew on. If you have a LeRobot-format dataset (v2.1 or v3.0), run &lt;code&gt;rda audit&lt;/code&gt; on it and tell me what turns up — especially if it's boring. Boring results from real data are how a tool earns trust.&lt;/p&gt;

&lt;p&gt;Issues, PRs, and "your idle-ratio threshold is wrong, here's why" comments all welcome.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;(RDA is MIT-licensed. I also do paid data-quality deep dives and pipeline integration for teams that want the audit without the homework.)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Tool: &lt;a href="https://github.com/liesliy/rda" rel="noopener noreferrer"&gt;https://github.com/liesliy/rda&lt;/a&gt; · PyPI: &lt;code&gt;robot-data-audit&lt;/code&gt; · UI: &lt;code&gt;rda ui&lt;/code&gt; (EN/中文)&lt;/p&gt;

</description>
      <category>robotics</category>
      <category>ai</category>
      <category>opensource</category>
      <category>python</category>
    </item>
    <item>
      <title>tlabel convert: One CLI to Bridge 9 Tactile Dataset Formats</title>
      <dc:creator>liesliy</dc:creator>
      <pubDate>Tue, 11 Aug 2026 04:08:43 +0000</pubDate>
      <link>https://dev.to/liesliy/tlabel-convert-one-cli-to-bridge-9-tactile-dataset-formats-400k</link>
      <guid>https://dev.to/liesliy/tlabel-convert-one-cli-to-bridge-9-tactile-dataset-formats-400k</guid>
      <description>&lt;p&gt;How a single command can unify GelSight, PaXini, Daimon, ToucHD, and 5 other tactile sensor formats into training-ready data.&lt;/p&gt;




&lt;p&gt;If you work in tactile robotics research, you've been here before:&lt;br&gt;
A collaborator sends you a dataset collected with a PaXini PXCap force array. Your pipeline expects GelSight .pkl files. Your colleague's LeRobot training code needs Zarr. Someone else is publishing results on a Daimon DM-TacClaw in .parquet format.&lt;/p&gt;

&lt;p&gt;Three sensors. Three formats. Three days of writing ad-hoc parsing scripts that you'll delete next week.&lt;/p&gt;

&lt;p&gt;Tactile sensing is having a moment — multiple billion-dollar funding rounds in embodied AI have poured attention (and capital) into the field in 2026. But while hardware is advancing fast, data interoperability is still a mess. Every sensor vendor ships data in a proprietary format, and there's no common lingua franca for tactile manipulation datasets.&lt;/p&gt;

&lt;p&gt;TLabel is an open-source project that tackles exactly this problem. And with the maturation of its CLI and adapter architecture through v0.18.x, the workflow for converting between tactile dataset formats has gotten dramatically simpler.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What TLabel Actually Does&lt;/strong&gt;&lt;br&gt;
Before diving into commands, let's set expectations. TLabel is a data pipeline standardization layer. It does not:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Interface with hardware or collect data&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Run inference or train models&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Replace your training pipeline&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It does:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Define a 14-dimension semantic annotation schema (covering contact, force, slip, texture, deformation, and more)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Provide adapter implementations that translate sensor-specific formats into that schema&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Export annotated data into training-ready formats (LeRobot, FTP-1 Zarr, JSON, CSV)&lt;br&gt;
Think of it as the Unicode for tactile data — one standard schema, every sensor.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The Adapter Landscape&lt;/strong&gt;&lt;br&gt;
TLabel currently ships with 12 built-in adapters covering 9 dataset formats and 3 real-time sensor interfaces:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dataset Adapters (Offline Data Loading)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&amp;lt;&amp;gt; — GelSight Mini / DIGIT, visuo-tactile, .pkl, L3&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&amp;lt;&amp;gt; — PaXini PXCap, force array, .h5, L2&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&amp;lt;&amp;gt; — Daimon DM-TacClaw, multimodal, .parquet, L3&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&amp;lt;&amp;gt; — ToucHD, visuo-tactile, .hdf5, L3&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&amp;lt;&amp;gt; — UniVTAC, visuo-tactile, .hdf5, L3&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&amp;lt;&amp;gt; — VTouch, visuo-tactile, .h5, L3&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&amp;lt;&amp;gt; — YCB-Slide, visuo-tactile, .npy, L3&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&amp;lt;&amp;gt; — TacQuad (AnyTouch), multi-sensor, directory, L3&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&amp;lt;&amp;gt; — TLabel native, meta format, .json, L1–L4&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Real-Time Sensor Adapters&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&amp;lt;&amp;gt; — PaXini GEN3, force array, SDK connection, L2&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&amp;lt;&amp;gt; — Daimon DM-Tac, visuo-tactile, USB / .avi, L3&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&amp;lt;&amp;gt; — PaXini PX6D, 6-axis force, placeholder, L2&lt;br&gt;
That covers the majority of tactile sensors used in manipulation research today.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Getting Started&lt;/strong&gt;&lt;br&gt;
Install is straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;tlabel

&lt;span class="c"&gt;# Or with sensor-specific extras:&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;tlabel[gelsight]     &lt;span class="c"&gt;# GelSight / DIGIT (.pkl)&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;tlabel[paxini]       &lt;span class="c"&gt;# PaXini PXCap (.h5)&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;tlabel[daimon]       &lt;span class="c"&gt;# Daimon DM-TacClaw (.parquet)&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;tlabel[ftp1]         &lt;span class="c"&gt;# FTP-1 export (zarr)&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;tlabel[all]          &lt;span class="c"&gt;# Everything&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Exploring Adapters from the CLI&lt;/strong&gt;&lt;br&gt;
Once installed, you can inspect what's available without writing any code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# List all registered adapters&lt;/span&gt;
tlabel list
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This prints all dataset and real-time adapters with their type, native format, and compliance level. It's the first command I run when working with a new dataset.&lt;/p&gt;

&lt;p&gt;For details on a specific adapter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;tlabel info gelsight
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This shows the adapter's capability declaration — which of the 14 semantic dimensions it can annotate, its compliance level, and any format-specific notes. For example, GelSight outputs force_vector (L3) but not temperature (L4), while a simple resistive sensor might only declare L1 fields like contact and slip_event.&lt;br&gt;
This capability declaration system is one of TLabel's key design decisions. Rather than forcing every sensor to produce all 14 dimensions (which would mean fabricating data it can't actually measure), each adapter honestly declares what it can and cannot provide.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Validating Your Data&lt;/strong&gt;&lt;br&gt;
Before converting, it's worth checking that your data passes schema validation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;tlabel validate data.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This runs a compliance check against the 14-dimension Schema V2 and reports any issues. Catching format problems early saves debugging time downstream.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Converting Between Formats&lt;/strong&gt;&lt;br&gt;
Here's where things get practical. TLabel provides two paths for format conversion: CLI for quick checks and the Python API for full control.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Conversion via CLI&lt;/strong&gt;&lt;br&gt;
For simple cases — say you have a single GelSight .pkl file and want a JSON summary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;tlabel &lt;span class="nb"&gt;export&lt;/span&gt; &lt;span class="nt"&gt;--input&lt;/span&gt; grasp_data.pkl &lt;span class="nt"&gt;--format&lt;/span&gt; json &lt;span class="nt"&gt;--output&lt;/span&gt; annotations.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This reads the GelSight data through its adapter, applies the Schema V2 annotation, and writes a clean JSON file with all 14 semantic dimensions (at the appropriate compliance level for the sensor).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Full Conversion Pipeline via Python&lt;/strong&gt;&lt;br&gt;
For the heavy lifting — converting entire datasets into training-ready formats — the Python API gives you the most control:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;tlabel&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;tlabel.converters&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;tlabel_to_lerobot&lt;/span&gt;

&lt;span class="c1"&gt;# Load data from any supported sensor
&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tlabel&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;path/to/paxini_data.h5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Inspect what you got
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;span class="c1"&gt;# -&amp;gt; {'num_frames': 500, 'sensor': 'paxini', 'compliance_level': 'L2', ...}
&lt;/span&gt;
&lt;span class="c1"&gt;# Export to JSON/CSV for analysis
&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;export&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Export to FTP-1 Zarr for foundation model training
&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;export_ftp1&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output.zarr&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Convert to LeRobot episode format
&lt;/span&gt;&lt;span class="nf"&gt;tlabel_to_lerobot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;annotations.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lerobot_episode/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here's a concrete example that walks through a realistic workflow — loading PaXini force array data, validating annotations, and exporting to both LeRobot and FTP-1 formats:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;tlabel&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;tlabel.converters&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;tlabel_to_lerobot&lt;/span&gt;

&lt;span class="c1"&gt;# Step 1: Load PaXini PXCap data
&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tlabel&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;experiment_01.h5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Step 2: Validate schema compliance
&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;validate_annotations&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="c1"&gt;# Reports any missing or malformed fields
&lt;/span&gt;
&lt;span class="c1"&gt;# Step 3: Auto-annotate events from signal patterns
&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;annotate_events_auto&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="c1"&gt;# Detects: contact_onset, contact_loss, slip events, force spikes
&lt;/span&gt;
&lt;span class="c1"&gt;# Step 4: Export
&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;export&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;experiment_01.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;           &lt;span class="c1"&gt;# Analysis
&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;export_ftp1&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;experiment_01.zarr&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;      &lt;span class="c1"&gt;# Foundation model training
&lt;/span&gt;&lt;span class="nf"&gt;tlabel_to_lerobot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;experiment_01.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;     &lt;span class="c1"&gt;# LeRobot pipeline
&lt;/span&gt;                   &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lerobot_episode/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Batch Processing Multiple Files&lt;/strong&gt;&lt;br&gt;
When you're dealing with an entire experiment directory (which is the typical case — real manipulation datasets have hundreds of episodes), you can batch-load and convert:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;glob&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;tlabel&lt;/span&gt;

&lt;span class="n"&gt;files&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;glob&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;glob&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;raw_data/*.h5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;files&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tlabel&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;validate_annotations&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;export&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;annotated/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sensor&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;_&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;num_frames&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;f.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Converted &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; -&amp;gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sensor&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; (L&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;compliance_level&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Understanding the Architecture&lt;/strong&gt;&lt;br&gt;
The adapter system sits in a three-layer architecture:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌─────────────────────────────────────────────────┐
│ Layer 1: Schema                                 │
│ 14 semantic dimensions + Compliance Level L1-L4 │
├─────────────────────────────────────────────────┤
│ Layer 2: Adapters                               │
│ DataAdapterBase │ SensorAdapterBase             │
├─────────────────────────────────────────────────┤
│ Layer 3: Downstream                             │
│ Feature derivation · Export · Augmentation      │
│ FTP-1 · LeRobot · RLDS · ROS2                  │
└─────────────────────────────────────────────────┘

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Layer 1 is the schema itself — 14 semantic dimensions covering spatial perception (contact, centroid, region), mechanics (force magnitude, force vector, torque), dynamics (slip event, slip velocity, manipulation phase), surface properties (texture class), and meta-perceptions (deformation, temperature, confidence, compliance level).&lt;br&gt;
Layer 2 is where adapters live. There are two base classes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;DataAdapterBase — for offline dataset files (sublcass this to add support for a new sensor format, takes ~30 minutes)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;SensorAdapterBase — for real-time hardware connections (streaming data from a live sensor)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both produce output conforming to the same Schema V2, which means Layer 3 downstream tools work identically regardless of which sensor the data came from.&lt;/p&gt;

&lt;p&gt;Layer 3 handles everything after annotation: feature derivation, data augmentation, and export into training frameworks. This is where the LeRobot converter, FTP-1 Zarr exporter, and RLDS bridge live.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Compliance Level System&lt;/strong&gt;&lt;br&gt;
One concept worth explaining in more detail: Compliance Levels (L1–L4). This is TLabel's answer to the question "what if my sensor can't measure temperature or 6-axis force?"&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;L1 — Basic Tactile: contact, centroid, slip, confidence. Examples: Single-point resistive, proximity sensors&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;L2 — Force-Aware: L1 + force_magnitude. Examples: PaXini, YCB-Slide, GelSight&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;L3 — Full-Vector: L2 + force_vector. Examples: ToucHD, calibrated DM-TAC&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;L4 — Rich-Semantic: L3 + all optional fields. Examples: BioTac, next-gen multimodal sensors&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key insight: a PaXini at L2 and a ToucHD at L3 both produce valid TLabel output. They just populate different subsets of the 14 dimensions. Downstream code can check compliance_level to decide what it can and cannot use, rather than writing sensor-specific branches.&lt;br&gt;
This is what makes cross-sensor comparison possible. You can train a model on GelSight data (L3) and evaluate it on PaXini data (L2), knowing exactly which fields are comparable and which aren't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What About My Sensor?&lt;/strong&gt;&lt;br&gt;
TLabel is designed for extensibility. If your sensor isn't supported yet, adding an adapter takes about 30 minutes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Fork the adapter template from contrib/adapter-template/&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Subclass DataAdapterBase (for datasets) or SensorAdapterBase (for hardware)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Implement the required methods and declare your compliance level&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Submit a PR or publish as a standalone package&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The project also supports register_external_adapter() and entry_points auto-discovery, so third-party adapters can be published independently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How This Fits Into the Embodied AI Pipeline&lt;/strong&gt;&lt;br&gt;
The broader context matters. Foundation models for robotics — like those being built on top of LeRobot, Open X-Embodiment, and similar frameworks — need diverse, standardized training data. But tactile data has been a bottleneck:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Collection is sensor-specific (hardware-dependent)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Annotation has been ad-hoc (no standard schema)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Training expects uniform input formats&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;TLabel deliberately addresses step 2 and 3 only. It's a standardization layer that sits between your raw sensor data and your training pipeline. By providing a common schema with honest capability declarations, it enables:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Cross-sensor training: Mix data from different sensors in the same training batch&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Capability-aware models: Train models that know what information is available at each compliance level&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Reproducible research: Compare results across labs using different hardware&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The project is also actively contributing tactile data format support upstream to LeRobot via PR #4032, which signals growing recognition that tactile data standards are needed in the broader robotics ecosystem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick Reference&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;See all supported adapters → tlabel list&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Get adapter details → tlabel info gelsight&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Validate a data file → tlabel validate data.json&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Export to JSON → data.export("out.json")&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Export to FTP-1 Zarr → data.export_ftp1("out.zarr")&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Convert to LeRobot format → tlabel_to_lerobot(src, dst)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Load any sensor data → tlabel.load("file")&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Try a demo (no files needed) → tlabel.demo("gelsight")&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Wrapping Up&lt;/strong&gt;&lt;br&gt;
Tactile data interoperability isn't glamorous work — it's infrastructure. But it's the kind of infrastructure that determines whether the field can scale beyond lab-specific pipelines to shared, reproducible, cross-sensor research.&lt;br&gt;
TLabel's adapter architecture and CLI tools won't solve every data problem in tactile robotics. But they provide a concrete, working answer to the question: "How do I get data from sensor X into format Y without writing a custom parser?"&lt;br&gt;
That question used to take a weekend. Now it takes one line.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Links:&lt;/strong&gt;&lt;br&gt;
GitHub: &lt;a href="https://github.com/liesliy/tlabel" rel="noopener noreferrer"&gt;https://github.com/liesliy/tlabel&lt;/a&gt;&lt;br&gt;
PyPI: &lt;a href="https://pypi.org/project/tlabel/" rel="noopener noreferrer"&gt;https://pypi.org/project/tlabel/&lt;/a&gt;&lt;br&gt;
TLabel Paper: PDF on GitHub&lt;br&gt;
Schema V2 Spec: docs/tlabel-format.md&lt;br&gt;
Adapter Template: contrib/adapter-template&lt;br&gt;
Contributing Guide: CONTRIBUTING.md&lt;br&gt;
LeRobot PR #4032: huggingface/lerobot#4032&lt;br&gt;
Previous Dev.to post: TLabel: Unifying Tactile Data Annotation for Robotics&lt;br&gt;
Open X-Embodiment: robotics-transformer-x.github.io&lt;br&gt;
OpenTouch: opentouch.ai&lt;/p&gt;

&lt;p&gt;TL;DR — TLabel is the Unicode for tactile data: one standard schema, every sensor. Install with pip install tlabel, run tlabel list to see what's supported.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>robotics</category>
      <category>data</category>
    </item>
  </channel>
</rss>
