DEV Community

KROU4
KROU4

Posted on

I turned the PS5 camera in my drawer into a depth-sensing webcam

The PlayStation 5 HD Camera spent most of its life in my drawer. One day I looked at what's inside: two 1080p sensors behind a bridge chip, on USB 3. That's a stereo camera. A stereo camera can measure how far everything is, which means the background blur everyone wants on video calls could come from real depth instead of a neural network guessing where your hair ends.

So I wrote a driver. This is what I ran into on the way.

The camera doesn't boot by itself

Plug the camera into a PC and nothing useful appears. It enumerates as a bare loader (USB 05A9:0580) and waits for firmware, every single time it's plugged in. After the firmware upload it disconnects and comes back as a standard UVC camera (05A9:058C), which Windows, Linux and macOS already know how to drive.

The firmware belongs to Sony, so I can't ship it. The installer downloads the original image from public copies, checks its SHA-256 and applies my changes on the user's computer: a 90-byte patch, published as JSON in the repo.

60 fps and a USB budget

Stock firmware gives 1080p at 30 fps. The patch unlocks 60. Depth needs both sensors at once, though, and that's where USB says no: two full 1080p60 streams need about 498 MB/s, and this camera's link carries about 393.

The fix lives in the bridge chip. The patch adds a mode where the main sensor sends the full 1920x1080 frame and the second sensor sends a 960x540 copy, downscaled in hardware. About 320 MB/s, and it fits. Depth is computed at 640x360 anyway, so the smaller second view loses nothing, and apps still get the main sensor's full 1080p at full field of view.

[image: both sensors side by side]

Depth without AI

The depth pipeline runs in GPU compute shaders, Direct3D 11 on Windows:

  1. Census transform. Every pixel gets a 62-bit signature: is each neighbour in a 9x7 window brighter than the centre? Census compares order, not brightness, so the two sensors' different gains don't matter.
  2. Semi-global matching. Matching costs for 64 disparities, aggregated along four directions, with a smaller penalty for depth jumps where the image has an edge.
  3. Clean-up. A uniqueness test, a left-right consistency check, filling holes from the eight directions around them, and smoothing over time so a still head doesn't shimmer.
  4. A guided filter brings the 640x360 depth back to full resolution along the edges of the picture.
  5. A silhouette in depth. Ears and hair sit behind the face by about the depth of a head; the wall behind sits much farther. A path cost from the face's depth keeps the whole head sharp and the wall blurred.
  6. Bokeh. A scatter-as-gather blur at half resolution, where bright points become visible discs.

On an RTX 3060 Ti all of this takes about 9 ms per 1080p frame.

[image: depth map]

Living inside Windows' camera pipeline

I didn't want a separate virtual camera or an app that has to keep running. Windows has a supported way to process frames inside a camera itself: a Device MFT, which Windows Camera Frame Server loads into that camera's pipeline. It runs in user mode, so there's no kernel driver to sign. Every app sees one camera, "PS5 Camera", and the blur is switched in Windows' own background effects settings, next to Windows Studio Effects.

The Linux build runs the same HLSL shaders compiled to SPIR-V on Vulkan and writes to a v4l2loopback device. On the same GPU the output matches the Windows build frame by frame.

Where it stands

  • Windows 11: 1080p60 and the depth blur.
  • Linux: 1080p 30/60; the blur is experimental.
  • macOS: 1080p 30/60; no blur, because camera extensions there need an Apple developer signature.

It's free and GPL-3.0. If you have the camera, the most useful thing you can send me is how it runs on your GPU, especially integrated graphics.

Not affiliated with Sony. PlayStation and PS5 are trademarks of Sony Interactive Entertainment Inc.

Top comments (1)

Collapse
 
kyisaiah47 profile image
kyisaiah47 •

Does the Device MFT receive both sensor frames in one callback, or synchronize separate capture callbacks before dispatching the shaders?