This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built SignBridge for my close friend Minh, who was born deaf and communicates primarily through Vietnamese Sign Language (VSL). In university seminars and daily coffee chats, communication was slow because typing takes away personal eye contact.
SignBridge is a bidirectional, real-time sign language interpreter. It runs client-side hand-landmark tracking and local neural network inference to translate continuous hand gestures into natural Vietnamese speech and text with near-zero latency.
Demo
- Repository: github.com/minKasent/signbridge
- Architecture: Next.js 16 (in-browser MediaPipe) ➔ Spring Modulith 2.1 ➔ FastAPI (ONNX Runtime) ➔ SpeechSynthesis Audio Output.
Code
minKasent
/
signbridge
AI/Web sign language interpretation and translation bridge.
SignBridge 🖐️
Hệ thống dịch ngôn ngữ ký hiệu tiếng Việt (VSL) hai chiều, thời gian thực — đồ án tốt nghiệp 2026.
- 📋 Kế hoạch chi tiết 16 tuần: PLAN.md
- 🔑 Hướng dẫn đăng ký tài khoản & lấy API key: docs/SETUP-KEYS.md
- 🧠 Pipeline huấn luyện model: training/README.md
Kiến trúc
apps/web — Next.js 16: webcam + MediaPipe (landmark NGAY TRONG browser) → WebSocket
apps/core — Spring Boot 4.1 MODULAR MONOLITH (Spring Modulith 2.1):
identity | dictionary | translation | dataset | analytics
apps/ml — FastAPI: 1 service mỏng suy luận gloss từ chuỗi landmark (ONNX)
training/ — tiền xử lý + notebook huấn luyện (Colab/Kaggle, log lên W&B)
Luồng phiên dịch đầy đủ:
webcam → MediaPipe (browser) → WebSocket → cửa sổ trượt 32 frame
→ ML service (gloss + độ tin cậy) → buffer gloss theo phiên
→ nghỉ tay 2s (nhãn NO_SIGN) → LLM ghép câu tiếng…How I Built It
- MediaPipe (In-browser Computer Vision): 21 3D hand landmarks tracked per hand at 60 FPS right inside the browser. No raw video feed ever leaves the client machine.
- ONNX Runtime & Light CNN/LSTM: 32-frame sliding window model classifying continuous VSL gloss sequences.
- Spring Modulith 2.1: Modular monolith routing sessions, dictionary mapping, and LLM sentence reconstruction.
- Local Fallback: Works completely offline without external APIs when dictionary heuristics are used.
Why Does Open Innovation Matter?
- Zero Cloud Latency & Privacy: Deaf users deserve full camera privacy. Processing skeletal landmarks on-device means no video frames are transmitted to proprietary cloud vision servers.
- Cost & Sustainability: Minh uses this daily on a student laptop. Running an open-weight ONNX model costs $0 in monthly API subscriptions.
- Adaptability: Proprietary speech-to-text models neglect low-resource sign languages like VSL. Open-source models allowed us to collect localized datasets and train custom gloss classifiers tailored to Vietnamese signers.
Prize Categories
- Featured: Best Use of Gemma
- Partner: Best Use of GitHub Copilot
Top comments (0)