Stimulating the visual cortex makes a person see small spots of light called phosphenes. A visual prosthesis built on that has to decide, for every image, which electrodes get how much current. I wrote that step in PyTorch, with a differentiable simulator on the other end that shows what the stimulation would look like. It runs on a laptop CPU, with 100, 1,000 or 10,000 electrodes.
This is a simulation. I did not stimulate anything. The only literature numbers in it are the electrode thresholds, though the cortical map and the simulator's design come from papers too. Here is how it is put together, and the evaluation, which surprised me twice.
Placing electrodes and phosphenes
Vision is distorted on the cortex: the center of the visual field takes far more area than the periphery. The complex logarithm w = log(z + a) approximates that map (Schwartz, Vision Research 1980). I put electrodes on a jittered grid in cortical coordinates, one patch per hemisphere, and mapped them back to the visual field. The value a = 0.75 degrees is my own choice.
z = np.exp(uu + 1j * vv) - A_DEG # right visual field, in degrees
keep = z.real >= 0.0 # drop points that cross the vertical meridian
sigma = size_scale * du * (ecc + A_DEG) # phosphene size grows with eccentricity
Phosphene size is the cortical spread divided by magnification, so it scales with eccentricity plus a. A search over grid sizes gives exactly 100, 1,000 or 10,000 electrodes.
From current to percept
Each electrode has a threshold. Brightness is a sigmoid around it, and the percept is the saturated sum of Gaussian blobs:
def brightness(self, current_ua):
return torch.sigmoid((current_ua - self.thr) / (SLOPE * self.thr))
def forward(self, current_ua):
total = self.brightness(current_ua) @ self.basis # basis: (electrodes, pixels)
return 1 - torch.exp(-total)
Fernández et al. (2021) reported single-electrode thresholds of 66.8 ± 36.5 µA with a 96-electrode Utah array. I drew thresholds from a lognormal with that mean and SD, clipped to 15 to 140 µA. After clipping the mean is 65.2 and the SD 31.2 µA, so the SD is lower than the paper's. Brightness slope, phosphene size and the 150 µA current limit are my assumptions. The simulator has no timing and no interaction between neighbouring electrodes, and it is a minimal reimplementation inspired by Dynaphos (van der Grinten et al., eLife 2024), not a port.
Three encoders
The encoder decides a target brightness for each electrode, and a mapper turns that into current using estimated thresholds:
def brightness_to_current(b, thr_hat):
b = b.clamp(0.02, 0.98)
logit = torch.log(b / (1 - b))
return (thr_hat * (1 + SLOPE * logit)).clamp(0.0, I_MAX_UA)
- Simple sampling: each electrode gets the phosphene-weighted mean of the image
- Edges: the same, after a Sobel filter
- Learned: a CNN with 107,017 parameters. It looks at the image at three scales and samples its feature maps at each phosphene's position in the visual field. A small head then combines those features with eccentricity, phosphene size, the estimated threshold and the current budget
Because the learned encoder samples features per electrode, one set of weights serves all three electrode counts. A mean-current limit is enforced by scaling all currents down together.
Calibration
A real patient's thresholds are unknown. I simulated the clinic's procedure: for each electrode, a binary search with 7 steps, asking "did you see a spot?" three times per step and taking the majority, against a noisy psychometric response. That is 21 questions per electrode, 21,000 for a 1,000-electrode array. The estimate lands within about 6% of the true threshold, against about 51% when every electrode is assumed to sit at the population mean.
Training
I trained end to end through the simulator for 3,000 steps on synthetic scenes (shapes, lines, letters), taking about 10 minutes on a laptop CPU. Each step picks an electrode count, a layout and threshold set, a calibration state (40% uncalibrated, otherwise noisy estimates up to 35%) and a current limit of none, 40, 25 or 15 µA. The loss is mean squared error plus 0.5 × (1 − SSIM) against a slightly blurred scene, and a small penalty on mean current.
Evaluation
I scored 120 held-out scenes with a simulated patient whose layout jitter and thresholds were not seen in training. Two metrics: SSIM against the blurred scene (large regions), and the correlation between the edge maps of percept and scene (fine detail).
| Electrodes | Sampling, uncalibrated | Sampling, calibrated | Learned, calibrated |
|---|---|---|---|
| 100 | 0.030 / 0.351 | 0.077 / 0.471 | 0.274 / 0.603 |
| 1,000 | 0.085 / 0.229 | 0.270 / 0.452 | 0.540 / 0.568 |
| 10,000 | 0.150 / 0.157 | 0.470 / 0.431 | 0.658 / 0.512 |
Each cell is edge correlation / SSIM. Two results stand out:
- Calibration is worth about as much as learning. At 1,000 electrodes, simple sampling goes from 0.085 to 0.270 with calibration, and the learned encoder adds another 0.27
- With a 25 µA mean-current limit, the learned encoder at 1,000 electrodes keeps an edge correlation of 0.512, against 0.307 for simple sampling
The two surprises came from looking at things, not from the table. A comparison figure showed a bright vertical stripe down the middle of the simple-sampling percept. The log map overshoots the vertical meridian near the fovea, so left and right electrodes overlapped. Dropping those points and retraining removed it.
The second was SSIM. It did not improve with electrode count and even fell (0.603, 0.568, 0.512 for the learned encoder), which contradicted what the images showed. SSIM against a blurred scene mostly rewards large regions, so a smooth percept from 100 electrodes scores well. I added the edge correlation, which does rise (0.274, 0.540, 0.658). SSIM and MSE are also the training objective, so they favor the learned encoder by construction. Edge correlation was not optimized directly.
Limits
The evaluation uses synthetic scenes, and the simulated patient comes from the same model that trained the encoder, so transfer to a different patient model is unshown. Phosphenes are fixed to the gaze in a real prosthesis, which would need eye tracking that I left out. These are image metrics, not a test with a person.
I ran the demo with my laptop camera, locally and on an EC2 instance behind CloudFront. In both, the image got better as I raised the electrode count. I deleted the AWS resources afterwards.
This covers the output side. For the input side, my earlier post rebuilds a brain-to-text decoder from public neural data.
References
- Schwartz, Vision Research 1980: https://www.sciencedirect.com/science/article/abs/pii/0042698980900905
- Fernández et al., J Clin Invest 2021: https://www.jci.org/articles/view/151331
- Dynaphos (eLife 2024): https://elifesciences.org/articles/85812



Top comments (0)