This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I Built
I built ROGVEDA for my friend Rohan, a graduate student researching medicinal chemistry and natural product drug leads.
Rohan was facing a painful dilemma common to academic and early-stage drug discovery researchers:
The Cloud Privacy Trap: He develops novel, patent-pending small molecules. Under institutional IP rules and non-disclosure agreements, he cannot upload unpublished chemical structures to closed cloud AI services without risking intellectual property leakage and patent forfeiture.
Workflow Fragmentation: To evaluate a single candidate molecule, he was jumping between several disconnected tools: a 2D drawing utility, desktop software for logP and polar surface area calculation, online servers for toxicity prediction, and manual spreadsheets to record modifications.
I built ROGVEDA to give Rohan a unified, 100% offline-first computational drug discovery workbench that runs entirely on his local machine. With ROGVEDA, he can:
- Sketch 2D molecular structures or paste SMILES strings.
- Compute 17+ deterministic physicochemical properties such as molecular weight, logP, TPSA, Lipinski Rule of 5 violations, and synthetic accessibility scores.
- Run calibrated machine learning models across 6 bioactivity and safety targets including hERG cardiotoxicity, Tox21 mitochondrial toxicity, anti-inflammatory, anti-microbial, anti-cancer, and anti-oxidant potential.
- Perform high-throughput batch screening on CSV libraries of up to 500 molecules with hardware-aware battery pacing.
- Receive contextual biochemical trade-off insights powered by a locally hosted, open-weight language model without a single byte of chemical data leaving his device.
When I handed the working build over to Rohan, his feedback was direct:
"Being able to screen a series of newly designed derivatives right on my laptop in seconds—without worrying about confidentiality or cloud rate limits—cuts days of tedious manual lookups down to minutes."
Demo
Garg-Pankaj29
/
Rogveda
ROGVEDA is an offline, AI-assisted molecular research sandbox that unifies structure analysis, ML property predictions, and local AI interpretation into a single workspace. Its core innovation is an iterative "Modify → Recalculate → Compare" loop for rapid, private compound design.
ROGVEDA
In-Silico Molecular Intelligence & Drug Discovery Screening Platform
One molecule. One workspace. One reproducible research loop.
1. Project Overview
ROGVEDA is an offline-first, privacy-preserving computational chemistry and early-stage drug discovery workbench.
The Problem It Solves
Early-stage medicinal chemistry and drug discovery workflows are heavily fragmented. Researchers, students, and small academic laboratories frequently juggle disconnected utilities:
- Standalone drawing apps for 2D chemical sketching.
- Cloud web servers for calculating molecular descriptors.
- Cloud APIs for bioactivity and toxicity estimations.
- External databases for similarity matching.
- Manual spreadsheets and notes for tracking structural modifications.
This fragmentation creates severe context switching, eliminates reproducibility, and forces researchers to upload proprietary or unpublished molecular structures to third-party cloud servers.
Why ROGVEDA Exists
ROGVEDA unifies the complete molecular exploration cycle into a single, local-first desktop application
Application Walkthrough
Interactive 2D Sketching and 3D Conformer Modeling
Real-time chemical canvas powered by embedded Ketcher and client-side RDKit WebAssembly, paired with interactive 3D WebGL ball-and-stick rendering via 3Dmol.js.Calibrated Bioactivity and Toxicity Prediction
Instant multi-target inference with probability calibration bars explicitly communicating confidence distance from the decision boundary.High-Throughput Batch Screening
Screening multi-molecule CSV libraries with row-level error isolation, property filtering, and one-click PDF and CSV report exports.
Code
The entire codebase is open-source under the MIT license:
Garg-Pankaj29
/
Rogveda
ROGVEDA is an offline, AI-assisted molecular research sandbox that unifies structure analysis, ML property predictions, and local AI interpretation into a single workspace. Its core innovation is an iterative "Modify → Recalculate → Compare" loop for rapid, private compound design.
ROGVEDA
In-Silico Molecular Intelligence & Drug Discovery Screening Platform
One molecule. One workspace. One reproducible research loop.
1. Project Overview
ROGVEDA is an offline-first, privacy-preserving computational chemistry and early-stage drug discovery workbench.
The Problem It Solves
Early-stage medicinal chemistry and drug discovery workflows are heavily fragmented. Researchers, students, and small academic laboratories frequently juggle disconnected utilities:
- Standalone drawing apps for 2D chemical sketching.
- Cloud web servers for calculating molecular descriptors.
- Cloud APIs for bioactivity and toxicity estimations.
- External databases for similarity matching.
- Manual spreadsheets and notes for tracking structural modifications.
This fragmentation creates severe context switching, eliminates reproducibility, and forces researchers to upload proprietary or unpublished molecular structures to third-party cloud servers.
Why ROGVEDA Exists
ROGVEDA unifies the complete molecular exploration cycle into a single, local-first desktop application
Core Stack:
- Frontend: React 19, Vite, Tailwind CSS, WebAssembly (@rdkit/rdkit), 3Dmol.js
- Backend: Python 3.12, FastAPI, RDKit C++ binaries
- Machine Learning: Scikit-Learn calibrated Random Forest ensembles trained on 230,000+ compounds from ChEMBL 37 and Tox21
- Local AI: LM Studio / llama.cpp local REST connector
How I Built It
ROGVEDA is built entirely around an open-source AI and cheminformatics stack designed for strict local execution:
Local Open-Weight LLM Integration
Instead of sending chemical graphs to proprietary cloud APIs, ROGVEDA hooks into a locally running open-weight model (such as Llama-3.2-1B-Instruct or Qwen2.5-3B-Instruct served via LM Studio or llama.cpp). The backend constructs strict prompts embedding deterministic RDKit calculations and machine learning confidence intervals, asking the local model to generate structured summaries, highlight pharmacophore trade-offs, and suggest bioisosteric modifications. If the language model server is turned off, the core chemistry pipeline still functions without interruption.Calibrated Machine Learning Engine
We trained Random Forest ensembles across 6 pharmacology and toxicology targets using 230,000+ compounds from ChEMBL 37 and the NIH Tox21 dataset:Compounds are featurized using 2048-bit Morgan Fingerprints.
Datasets are partitioned using 80/20 Bemis-Murcko Scaffold Splits to reflect true out-of-distribution prospective generalization.
Predictions use CalibratedClassifierCV with sigmoid calibration, producing true confidence probabilities rather than raw heuristic scores.
Hybrid Client-Server Cheminformatics
Client-Side: RDKit compiled to WebAssembly runs inside the browser to validate chemical valences and render vector SVG thumbnails instantly on every change with zero network latency.
Server-Side: Python FastAPI calculates 3D energy-minimized conformers using the ETKDGv3 distance-geometry algorithm and MMFF94 force field optimization, returning standard SDF blocks to the frontend WebGL viewer.
Why Does Open Innovation Matter?
For chemistry and life sciences, open innovation is not merely a philosophical preference—it is a functional necessity:
Zero-Leakage Privacy for Scientific IP: Closed APIs require transmitting molecular representations to cloud servers. For researchers working toward provisional patents, that constitutes public disclosure or third-party risk. Running open-weight models locally ensures no proprietary chemical telemetry ever touches an external network.
100% Offline Capability: Field researchers, university laboratories with intermittent campus Wi-Fi, and air-gapped institutional environments can run the entire platform with zero internet connectivity.
Zero Operational Cost: Academic researchers and students cannot afford recurring token bills or enterprise subscriptions. Open-source models and libraries allow unlimited screening at zero API cost.
Deterministic Reproducibility: Closed cloud models update and change behavior without warning, altering conclusions and breaking academic provenance. Open-weight models and version-pinned machine learning models ensure experiments run today produce identical results years from now.
Prize Categories
- Hacktoberfest Weekend Challenge: Build for a Friend
Top comments (0)