In the rapidly evolving landscape of Edge AI, privacy is no longer just a feature—it's a requirement. When dealing with sensitive medical data like skin photos, users are rightfully hesitant to upload their images to a remote server. Today, we are breaking boundaries by building a high-performance skin lesion segmentation tool that runs entirely on the client side. By leveraging Vision Transformers (ViT), Transformers.js, and ONNX Runtime, we can achieve real-time, "zero-server" inference that keeps data where it belongs: on the user's device.
This tutorial explores how to implement image segmentation using a miniaturized ViT architecture. We will weave together React for the UI and TailwindCSS for a sleek, medical-grade aesthetic, ensuring our Edge Computing solution is both powerful and user-friendly.
🏗️ The Architecture
The magic happens inside the browser's memory. Instead of sending a POST request to a heavy GPU cluster, we download a quantized ONNX model and execute it using WebAssembly (WASM) or WebGPU.
graph TD
A[User Selects Photo] --> B[React State Management]
B --> C[Canvas Pre-processing]
C --> D{Transformers.js Engine}
D --> E[ONNX Runtime / WASM]
E --> F[Miniaturized ViT Model]
F --> G[Segmentation Mask Output]
G --> H[Canvas Overlay Render]
H --> I[Privacy-Preserved Result]
🛠️ Prerequisites
To follow this advanced guide, you’ll need:
- React 18+ (Vite preferred)
- Transformers.js: The library that brings Hugging Face power to the browser.
- ONNX Runtime Web: For hardware-accelerated inference.
- TailwindCSS: For rapid UI development.
🚀 Step-by-Step Implementation
1. Project Setup & Model Selection
First, let's grab our dependencies. We are using transformers.js because it abstracts the complexity of tensor manipulation and model loading.
npm install @xenova/transformers react-dropzone lucide-react
For skin lesion segmentation, we typically use a model like SegFormer (a transformer-based framework for segmentation) quantized to q8 or fp16 to keep the bundle size under 30MB.
2. Creating the Inference Hook
We’ll create a custom hook to manage the model lifecycle. This prevents the browser from re-downloading the model on every render.
import { useEffect, useState, useRef } from 'react';
import { pipeline, env } from '@xenova/transformers';
// Allow worker-based multi-threading
env.allowLocalModels = false;
env.useBrowserCache = true;
export function useSegmentation() {
const [ready, setReady] = useState(false);
const segmenter = useRef(null);
useEffect(() => {
(async () => {
// Loading a lightweight SegFormer model
segmenter.current = await pipeline('image-segmentation', 'Xenova/segformer-b0-finetuned-ade-512-512', {
device: 'wasm', // or 'webgpu' for cutting-edge performance
});
setReady(true);
})();
}, []);
const runInference = async (imageSource) => {
if (!segmenter.current) return null;
const output = await segmenter.current(imageSource);
return output;
};
return { ready, runInference };
}
3. The UI Logic (React + Tailwind)
Now, let's build the component that handles the image upload and displays the segmentation mask as an overlay.
import React, { useState } from 'react';
import { useSegmentation } from './hooks/useSegmentation';
const SkinAnalyzer = () => {
const { ready, runInference } = useSegmentation();
const [image, setImage] = useState<string | null>(null);
const [isProcessing, setIsProcessing] = useState(false);
const handleFileUpload = async (e: React.ChangeEvent<HTMLInputElement>) => {
const file = e.target.files?.[0];
if (!file) return;
const url = URL.createObjectURL(file);
setImage(url);
setIsProcessing(true);
// Run the Edge AI Inference
const result = await runInference(url);
console.log("Segmentation Result:", result);
// Logic to draw mask on <canvas> would go here
setIsProcessing(false);
};
return (
<div className="max-w-4xl mx-auto p-8 bg-slate-50 rounded-2xl shadow-xl">
<h2 className="text-3xl font-bold text-slate-800 mb-4">🔬 On-Device Skin Screening</h2>
<p className="text-slate-600 mb-8">
Private. Secure. Fast. Your data never leaves this browser.
</p>
{!ready ? (
<div className="animate-pulse flex space-x-4">Loading AI Models...</div>
) : (
<div className="border-2 border-dashed border-indigo-200 rounded-xl p-12 text-center">
<input type="file" onChange={handleFileUpload} className="hidden" id="upload" />
<label htmlFor="upload" className="cursor-pointer bg-indigo-600 text-white px-6 py-3 rounded-lg hover:bg-indigo-700 transition">
{isProcessing ? "Analyzing..." : "Upload Lesion Photo"}
</label>
</div>
)}
{image && (
<div className="mt-8 relative flex justify-center">
<img src={image} alt="Target" className="max-w-full rounded-lg shadow-md" />
{/* Mask overlay canvas would be positioned absolute here */}
</div>
)}
</div>
);
};
💡 The "Official" Way to Scale
While running models in the browser is fantastic for privacy, deploying these solutions in a production healthcare environment requires robust pipelines, model versioning, and advanced optimization techniques like Knowledge Distillation.
For a deeper dive into production-ready AI patterns, enterprise-grade deployment strategies, and more advanced computer vision tutorials, I highly recommend exploring the technical resources at the WellAlly Blog. It’s a goldmine for developers looking to bridge the gap between "cool demo" and "scalable product." 🥑
🎨 Post-Processing the Results
Vision Transformers output a probability map. To make this useful for medical screening (as a preliminary tool), we need to apply a threshold and smooth the edges using a Gaussian filter or simple morphological operations on a hidden canvas.
// Example: Converting the model output to a visual mask
const processOutput = (masks) => {
const canvas = document.createElement('canvas');
const ctx = canvas.getContext('2d');
masks.forEach(mask => {
// Each mask object contains: score, label, and mask (RawImageData)
if (mask.label === 'skin_lesion_identifier') {
// Render mask pixels to canvas with 50% opacity
const imageData = new ImageData(mask.mask.data, mask.mask.width, mask.mask.height);
ctx.putImageData(imageData, 0, 0);
}
});
};
🏁 Conclusion
By bringing Vision Transformers to the edge, we are empowering users with sophisticated diagnostic tools without compromising their fundamental right to privacy. This "Learning in Public" project shows that with the right tech stack—Transformers.js, ONNX, and React—the browser is no longer just a document viewer; it's a high-performance AI execution environment.
What's next?
- Try implementing WebGPU support for 10x faster inference.
- Experiment with different backbones like MobileViT for even smaller footprints.
- Drop a comment below if you want a part 2 on "Fine-tuning ViT for specific dermatology datasets"!
Happy coding! 🚀🥑
Found this helpful? Check out more advanced guides at wellally.tech/blog.
Top comments (0)