Creating a Robot Sensor Data Recording and Replay System
Every robot learning project eventually needs the same infrastructure: a way to record everything the robot sees and does, and a way to play it back later — for debugging, dataset inspection, offline evaluation, or regenerating labels without touching the physical robot again. This tutorial builds a general-purpose recording and replay system you can drop into any robotics pipeline.
Why You Need This
Recording raw sensor streams alongside your teleoperation or autonomous runs pays off in ways that aren't obvious until you need it:
- Debugging policy failures without needing the robot present
- Re-deriving labels (e.g., re-running object detection or segmentation on old camera frames)
- Regression testing — replaying the exact same episode against a new model version
- Sharing datasets with collaborators who don't have access to the hardware
Architecture
[Sensors/Robot State] --> [Recorder] --> [Storage: rosbag / hdf5 / mcap]
|
v
[Replay Engine]
|
-----------------------------------------
| | |
[Visualization] [Re-labeling] [Offline Eval]
Step 1: Define What to Record
Decide on your sensor set up front — changing it mid-project fragments your dataset. A typical manipulation setup records:
- RGB (and optionally depth) from one or more cameras
- Robot joint positions, velocities, and torques
- End-effector pose
- Gripper state / force
- Any external sensors (F/T sensor, tactile sensors, IMU)
- Timestamps for everything, from a single shared clock
SENSOR_SCHEMA = {
"camera_front_rgb": {"shape": (480, 640, 3), "dtype": "uint8", "hz": 30},
"camera_wrist_rgb": {"shape": (480, 640, 3), "dtype": "uint8", "hz": 30},
"joint_positions": {"shape": (7,), "dtype": "float32", "hz": 100},
"joint_torques": {"shape": (7,), "dtype": "float32", "hz": 100},
"ee_pose": {"shape": (7,), "dtype": "float32", "hz": 100},
"ft_sensor": {"shape": (6,), "dtype": "float32", "hz": 500},
}
Step 2: The Recorder
Sensors often run at different rates. Rather than forcing everything into a single fixed-rate loop, record each stream at its native rate with its own timestamp, and resolve synchronization at replay/read time.
import threading
import queue
import time
class SensorRecorder:
def __init__(self, storage_backend):
self.storage = storage_backend
self.queues = {}
self.threads = []
self.running = False
def register_sensor(self, name, read_fn, hz):
self.queues[name] = queue.Queue()
def _loop():
dt = 1.0 / hz
while self.running:
start = time.time()
data = read_fn()
self.queues[name].put((time.time(), data))
time.sleep(max(0, dt - (time.time() - start)))
self.threads.append(threading.Thread(target=_loop, daemon=True))
def start(self):
self.running = True
for t in self.threads:
t.start()
def stop_and_flush(self, episode_id):
self.running = False
for name, q in self.queues.items():
records = []
while not q.empty():
ts, data = q.get()
records.append((ts, data))
self.storage.write(episode_id, name, records)
Step 3: Storage Backend
For robotics, three formats dominate: rosbag/rosbag2 (if you're in the ROS ecosystem), MCAP (a modern, ROS-agnostic container format gaining popularity), and plain HDF5 (simplest, good for non-ROS pipelines). Here's a minimal HDF5 backend:
import h5py
import numpy as np
class HDF5Storage:
def __init__(self, base_path):
self.base_path = base_path
def write(self, episode_id, sensor_name, records):
path = f"{self.base_path}/episode_{episode_id}.hdf5"
with h5py.File(path, "a") as f:
grp = f.require_group(sensor_name)
timestamps = np.array([r[0] for r in records])
data = np.stack([r[1] for r in records])
grp.create_dataset("timestamps", data=timestamps, compression="gzip")
grp.create_dataset("data", data=data, compression="gzip")
If you expect to scale past a few thousand episodes or need cross-language tooling, MCAP is worth the extra setup — it has strong support for indexed random access and streaming playback.
Step 4: The Replay Engine
Replay needs to reconstruct a synchronized view across sensors that were recorded at different rates. The standard approach is nearest-timestamp alignment to a chosen reference clock (often the slowest sensor, e.g., the camera).
class ReplayEngine:
def __init__(self, hdf5_path):
self.file = h5py.File(hdf5_path, "r")
def get_synced_frame(self, reference_sensor, index):
ref_ts = self.file[reference_sensor]["timestamps"][index]
frame = {reference_sensor: self.file[reference_sensor]["data"][index]}
for sensor_name in self.file.keys():
if sensor_name == reference_sensor:
continue
timestamps = self.file[sensor_name]["timestamps"][:]
nearest_idx = np.argmin(np.abs(timestamps - ref_ts))
frame[sensor_name] = self.file[sensor_name]["data"][nearest_idx]
return frame
def __len__(self):
return len(self.file["camera_front_rgb"]["timestamps"])
Step 5: Playback and Visualization
A simple playback loop lets you scrub through an episode visually — invaluable for spotting labeling errors or teleoperation glitches:
import cv2
def visualize_episode(replay_engine, reference_sensor="camera_front_rgb", fps=30):
for i in range(len(replay_engine)):
frame = replay_engine.get_synced_frame(reference_sensor, i)
img = frame[reference_sensor]
cv2.putText(img, f"t={i}", (10, 20), cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0, 255, 0), 1)
cv2.imshow("Replay", img)
if cv2.waitKey(int(1000 / fps)) & 0xFF == ord('q'):
break
cv2.destroyAllWindows()
Step 6: Replaying Actions Against a Simulator or Policy
Beyond visualization, you can replay recorded actions into a simulator to validate physical plausibility, or feed recorded observations into a trained policy for offline comparison against the original human action:
def offline_policy_comparison(replay_engine, policy, stats):
errors = []
for i in range(len(replay_engine)):
frame = replay_engine.get_synced_frame("camera_front_rgb", i)
pred_action = policy.predict(frame)
true_action = frame.get("action")
if true_action is not None:
errors.append(np.linalg.norm(pred_action - true_action))
return np.mean(errors)
This gives you a fast, hardware-free sanity check before committing to a real robot evaluation.
Practical Tips
- Record a single monotonic clock source across all sensors if possible (e.g., a shared NTP-synced or hardware trigger clock) — this eliminates most synchronization headaches.
- Store raw, unprocessed sensor data. Apply cropping, resizing, or filtering at load time, not at record time — you can't undo lossy preprocessing later.
- Include a manifest file per episode (task name, success flag, sensor list, schema version) so tooling can validate compatibility before loading.
- Compress images inline (e.g., JPEG) if storage is tight, but keep at least one high-fidelity/lossless recording session for cases where compression artifacts matter (e.g., fine texture-based manipulation).
Wrapping Up
Across this series, we've covered the full loop: building a teleoperation system, capturing clean human demonstrations, turning them into a trained imitation learning policy, using VR controllers to make teleoperation more natural, and finally recording/replaying sensor data for debugging and offline evaluation. Together, these form the core data infrastructure behind most modern robot learning pipelines.
Useful Links
Website: www.v-modal.com
SDK Flutter: https://github.com/v-modal/vmodal_sdk_flutter
SDK Android: https://github.com/v-modal/vmodal_sdk_android
Discord: https://discord.gg/K72z28KUx
Top comments (0)