DEV Community

Cover image for An AI “mind-reading” tool can reconstruct what you’re looking at from a brain scan
Ashish
Ashish

Posted on

An AI “mind-reading” tool can reconstruct what you’re looking at from a brain scan

Rendering a Different Image to Each Eye in VR — and Why We Built a Product Around It

This week, an AI tool made headlines for reconstructing what a person is looking at from their brain scans. It's a fascinating result, but it hinges on a simple architectural fact: vision is constructed in the brain from two separate input streams.

As VR developers, we control the final stage of that pipeline. And it turns out that controlling what each eye sees independently is a powerful lever for neuroplasticity.

The Dichoptic Trick

Every VR headset renders two views—one per eye—to create stereoscopic depth. Normally, these views are almost identical, just offset by the inter-pupillary distance. But there is no technical requirement for them to be the same. You can render arbitrarily different content to each eye, and the brain will attempt to fuse them into a single image.

This is called dichoptic presentation. It is the backbone of modern vision therapy for amblyopia (lazy eye), a condition where the brain suppresses input from one eye. If you dim the strong eye and sharpen the weak one, you nudge the brain back into using both.

On Quest, we implement this with a single-pass instanced stereo pipeline. The vertex shader runs once per instance (two instances, one per eye), while the fragment shader executes per-pixel for each view. This lets us keep the geometry pass unified and branch only in the pixel shader where we apply per-eye color grading, blur, or contrast curves. A simplified fragment snippet looks like:

// unity_StereoEyeIndex is 0 for left, 1 for right
float eye = float(unity_StereoEyeIndex);
float3 col = tex2D(_MainTex, uv).rgb;

// Example: stronger eye gets desaturated, weaker eye gets sharpened
col = lerp(Desaturate(col), Sharpen(col), step(0.5, eye) * _Intensity);
return float4(col, 1.0);
Enter fullscreen mode Exit fullscreen mode

Because the heavy lifting stays in a single draw call, we stay comfortably under the 13 ms frame budget at 72 Hz even on Quest 2’s XR2 Gen 1.

Turning a Clinical Concept into a Product

We built Amblyotube at Seven Sports to solve the 'boredom' problem in vision training. It's a Meta Quest app that streams YouTube-style video content while applying per-eye differences as the coordination exercise.

Here are a few technical hurdles we encountered:

  • Per-Eye Compositing: We had to ensure the video and UI layers could be treated independently per eye without breaking depth cues. The video plane sits at a fixed virtual distance (~2 m) to avoid vergence-accommodation conflict. UI elements (progress bar, subtitles) are rendered as world-space quads at the same depth so they receive the same dichoptic treatment. Balancing contrast and focus levels per eye required significant iteration to find the 'sweet spot' for the user. We expose a three-parameter curve (contrast, saturation, Gaussian blur radius) per eye, stored in a ScriptableObject so clinicians or users can load presets without code changes.

  • The Comfort Budget: Deliberately creating visual asymmetry fights against every VR comfort guideline in the book. We had to carefully tune the intensity of the effects so users could complete 30-40 minute sessions without eye strain. Our comfort metric is simple: zero dropped frames and a sustained 72 fps. We profile with adb logcat -s VrApi and watch AppSpaceWarp and CPU/GPU headroom. Fixed foveated rendering (FFR) at level 2 (medium) saves ~15% GPU on Quest 2, giving us margin for the extra per-eye shader work. We also clamp the maximum inter-eye contrast delta to 0.6 in linear space — beyond that, users report diplopia within five minutes.

  • Content as a Constraint: We found that high-engagement content keeps the user's gaze steady. This stability is critical because it ensures the visual stimulus is consistent, making the exercise more effective. Our video player uses ExoPlayer with a custom MediaSource that injects per-frame metadata (timestamp, scene-cut flags) into a ring buffer consumed by the Unity main thread. When a scene cut occurs, we briefly ramp the dichoptic intensity down over 800 ms to avoid a fusional break. The streaming stack supports HLS and DASH with Widevine L1; we transcode to 4K/30 fps HEVC at 18 Mbps peak, which the XR2 decoder handles in hardware.

The Intersection of VR and Neuroplasticity

I believe there is a massive, underexplored space for consumer VR software that focuses on the brain's plasticity. The headset isn't just a display; it's a controlled, binocular input device pointed at the most adaptable system in the human body.

Amblyotube is designed as recreational and educational software—it's not a medical device or a cure. But by making the exercises entertaining, we're seeing that people are actually sticking with them. Early anonymized telemetry (opt-in, GDPR-compliant) shows median session length of 28 minutes and a 60% week-4 retention for users who complete the onboarding calibration. That’s not a clinical endpoint, but it’s a usage signal that the "boring" barrier is lowering.

Next steps on our roadmap: integrate Quest Pro eye-tracking to drive dynamic difficulty — if fixation stability drops, we automatically reduce the dichoptic delta. Also exploring a WebXR companion so clinicians can prescribe specific video playlists and review adherence dashboards without sideloading.

I'm happy to discuss the implementation or the vision behind this in the comments. You can find the app on the Meta Store:
https://www.meta.com/en-gb/experiences/amblyotube/25906906972338493/

Top comments (0)