DEV Community

yug
yug

Posted on

Genetic Result Is Awful. I Built a Private Explainer That Runs on Your Laptop

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

MyVariant, Explained is a private tool that helps people understand a genetic test result for hypertrophic cardiomyopathy (HCM), an inherited heart condition.

Genetic reports often come back with a "variant of uncertain significance" (VUS): the lab found a change in a gene but can't say whether it matters. In our data, more than 4,500 of 8,137 missense variants across nine HCM genes are labeled uncertain in ClinVar. That's over twice the number with a confident label. For a family, that means no clear answer on whether relatives should be screened, plus a long wait for a counselor appointment.

I'll be upfront: I didn't have one specific friend's story to build around. I built this for the people in that waiting room, and for counselors who don't have time to explain every variant. You type the gene and protein change from your report (for example MYL3 E152K), and you get:

  • a research estimate of whether the change looks more like known disease-linked variants or harmless ones,
  • the features that drove that estimate,
  • the amino-acid sequence window around the change,
  • a calm plain-English summary with questions to bring to a genetic counselor.

It is a research prototype, not a medical device. It never diagnoses anything, and every result says so.

Demo

GITHUB

Code

YOUTUBE DEMO

How I Built It

The open-source pieces are the core of the project:

  • ESM-2 (facebook/esm2_t33_650M_UR50D, open weights): a protein language model. For each variant I embed the protein window before and after the change and take the difference, which captures how much the mutation disturbs the sequence.
  • A two-tower PyTorch network + a Random Forest + Platt calibration: these combine the ESM-2 shift with 51 biophysical features (chemical distance, charge, size, domain context). They're trained on 1,954 high-confidence variants (1,441 pathogenic, 513 benign) from nine genes: MYH7, MYBPC3, TNNT2, TNNI3, TPM1, MYL2, MYL3, ACTC1 and TNNC1.
  • Gemma via Ollama (local): turns the numbers into a plain-language explanation, with a prompt that forbids diagnosis and treatment advice. If no local LLM is running, a built-in template takes over, so the app never fails silently. <!-- DELETE the Gemma bullet and the Prize Categories line if you did not see "Generated by: local LLM (gemma3:4b via Ollama)" in the app -->
  • Flask serves the API and the web page.

I care most about how they were tested. We removed six features that come from the clinical curation process, and we evaluated with leave-one-gene-out cross-validation, so every score comes from a model that never saw that gene. Under that strict test, the Random Forest reached a mean AUPRC of 0.886, against 0.791 for EVE, 0.800 for AlphaMissense, 0.842 for REVEL, 0.849 for MetaRNN and 0.876 for CardioBoost. The two-tower model scored 0.857 and was best calibrated on thin-filament genes (Brier 0.145 on held-out TNNT2).

Honest limits: CardioBoost beats us on MYBPC3 and MYH7, the training set is small, we use nine genes, ESM-2 is frozen, there's no 3D structure, and nothing here is clinically validated.

Why Does Open Innovation Matter?

A genome is about the most personal data there is, and it doesn't belong in a hosted chatbot.

  • Privacy: ESM-2 and Gemma run on a laptop. The variant never leaves the machine, there's no account and no cloud API, and after the one-time model download it works without internet.
  • Control: I wrote the explainer's rules myself (never diagnose, never recommend treatment, always send people to a counselor) and I can swap the model with a single setting. With a closed API I couldn't guarantee where the data goes or that behavior stays the same next month.
  • Cost: a lookup costs nothing, so a clinic or patient group could run this for free.
  • Honest research: because the weights are open, the method can be inspected and reproduced, which matters a lot for a tool that touches health decisions.

Prize Categories

  • Best Use of Gemma: local Gemma writes the plain-language explanation.

Top comments (0)