DEV Community

Cover image for Making LLM ratings auditable: quotes, limitations, and "can't judge" instead of zero
Ashish
Ashish

Posted on

Making LLM ratings auditable: quotes, limitations, and "can't judge" instead of zero

From Auditable LLMs to Auditable Vision Rehab: Why Transparency Wins

A recent dev.to post sparked a necessary conversation about LLM-generated ratings. The author proposed a radical ruleset: fix the evidence window, show the actual quotes, admit limitations, and provide a "can't judge" button instead of inventing a score. The core objective is to stop confident-sounding noise from masquerading as signal. In the world of AI, we call this auditability.

As I read this, I realized that vision rehabilitation for amblyopia (lazy eye) has historically operated in the opposite direction. For decades, the industry has been a "black box." Patients—especially adults who were often told it was too late for them—were subjected to opaque patches, red-blue glasses, and rigid clinic schedules. There was no dashboard, no choice of content, and zero visibility into what the dominant eye was seeing versus the lazy eye. It was a system of "trust the process," which is the medical equivalent of a hallucinating LLM.

This is why the development of Amblyotube by Seven Sports is so compelling from a practitioner's perspective. It applies the same logic of transparency and user-control to the challenge of binocular vision.

The Architecture of Control

Amblyotube, built for the Meta Quest, leverages the hardware's fundamental design—two independent eyepieces—to implement dichoptic vision training. Instead of a static medical device, it functions as a transparent layer over the content the user already enjoys. By using any YouTube video as the primary stimulus, the app removes the friction of "boring" therapy, which is the primary reason teenagers (13+) typically abandon traditional patching.

But the real innovation lies in the auditability of the visual experience. The user is not a passive recipient; they are the operator of their own training session. The system provides several key technical levers:

1. The Dominant Eye Shader: Rather than a binary "on/off" patch, this tool provides digital occlusion. Users can adjust blur, contrast, brightness, and opacity. This allows for a gradual transition from full patching to partial shading, ensuring the dominant eye remains active while the lazy eye is forced to engage.

2. AI-Driven MFBF (Lazy Eye Sharpener): To prevent the lazy eye from simply ignoring the image, the software uses AI-driven processing to identify human figures within the video stream. It then applies a sharpening effect exclusively to the lazy eye's feed. This creates a high-contrast signal that the brain cannot easily ignore.

3. The Magenta Focus Cue: One of the hardest parts of amblyopia is "fusion"—getting the brain to merge two different images into one. Amblyotube solves this with a moving magenta circular cue visible only to the lazy eye (while the dominant eye sees a neutral grey). Magenta was chosen because it rarely occurs in natural video, making it a distinct, high-visibility signal that helps the brain lock onto the target and promote neuroplasticity.

Rendering Pipeline & Latency Budget

Under the hood the app runs a stereoscopic render pass at 72 fps (Quest 2/3 native) with a per‑eye resolution of 1832 × 1920. The YouTube stream is decoded on the GPU via the MediaCodec surface, then split into two textures. A lightweight compute shader applies the dominant‑eye occlusion parameters in < 0.3 ms, while a separate TensorFlow Lite model (MobileNet‑V2, 0.5 M parameters) runs inference at 30 fps to detect human silhouettes for the MFBF sharpening pass. The total frame‑time budget stays under 13 ms, leaving headroom for the compositor and avoiding motion‑to‑photon latency spikes that would break immersion.

On‑Device AI Inference & Privacy

All inference happens on‑device; no video frames leave the headset. The model is quantized to INT8, reducing memory footprint to ~1.2 MB and keeping power draw under 1.5 W. This design choice satisfies both GDPR‑style data‑minimization requirements and the practical need for offline operation — users can train on a plane or in a clinic without Wi‑Fi.

Visual Accents & Adaptive Pulse Controls

To further refine the "lock‑on" capability, the software introduces Visual Accents. These include a yellow‑green glow (Highlight) and a red silhouette (Outline) around human figures. They aren't static; they can be set to a "breathing" rhythm (Hz) via pulse controls. This rhythmic oscillation prevents neural adaptation, keeping the user's attention focused on the target. The pulse frequency is exposed as a slider (0.5–4 Hz) and persisted in the user profile, enabling a simple A/B test: the practitioner can compare session logs at 1 Hz vs. 2 Hz and correlate with stereo‑acuity gains.

Extensibility Hooks

Amblyotube exposes a minimal JSON‑over‑WebSocket API for researchers who want to inject custom stimuli (e.g., synthetic gratings, Vernier targets) or pull per‑frame eye‑tracking telemetry (when the Quest Pro eye‑tracking module is present). The schema is versioned, and the app ships with a sandboxed TypeScript snippet that developers can drop into a React Native wrapper for rapid prototyping.

Engineering Engagement

From an engineering standpoint, this is far superior to traditional anaglyph glasses, which distort the entire color palette of the scene. Amblyotube maintains natural visual colors, making the experience sustainable for the recommended 30 to 40‑minute sessions. The combination of per‑eye shader control, on‑device AI, and a transparent data path turns a historically opaque therapy into an auditable, user‑driven loop — exactly the kind of system we demand from our LLMs.

The Human Result

When we move away from the "black box" and give the user the tools to see—and adjust—their own progress, the psychological barrier to rehabilitation drops. Amblyotube isn't a medical cure or a replacement for professional therapy, but it is a powerful assistive tool for visual coordination and attention. It transforms a chore into a recreational activity.

By treating vision training like an auditable system—where the input is controlled, the stimulus is targeted, and the experience is transparent—we can help more people regain the depth perception and binocular vision they thought they had lost.

Start your coordination practice here: https://www.meta.com/en-gb/experiences/amblyotube/25906906972338493/

Top comments (0)