This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass.
What I Built
EcoID is an offline-first plant identification tool for field observations.
The idea is simple: take a laptop or phone outside, photograph a real plant, and use a local AI model to get identification candidates. Then look at the plant yourself and verify the result.
EcoID is designed for:
- students learning about local plants
- field researchers and citizen scientists
- gardeners and small-scale farmers
- people who want to explore nature without sending their photos to a cloud service
The current model supports five plant classes:
- Mango —
Mangifera indica - Coconut —
Cocos nucifera - Banana —
Musa acuminata - Papaya —
Carica papaya - Cassava —
Manihot esculenta
The AI does not present its output as a fact. It returns ranked identification candidates with confidence scores. The user then chooses whether the result is:
- Verified
- Rejected
- Uncertain
This human-in-the-loop flow matters because a model prediction is only a suggestion until somebody looks at the actual plant.
EcoID stores field observations locally, including the original photo, model prediction, alternative predictions, confidence score, verification status, notes, optional GPS coordinates, capture time from EXIF metadata, and user corrections.
The project is designed to make the computer useful for a short moment, then get the person back to observing the real world.
Demo
EcoID runs locally on a laptop or a phone-accessible local network. There is no public hosted demo because local execution is part of the project's privacy and offline design.
The trained model and downloaded dataset are not committed to the repository. Model weights and image data are kept outside Git because of their size and individual dataset licensing terms.
After training or obtaining the ONNX model locally:
python -m app --model models/ecoid-global-20261010.onnx --open
The app opens at:
http://127.0.0.1:8765
The main flow is:
Capture or upload a photo
↓
Run local inference
↓
Review possible identifications
↓
Verify, reject, or mark uncertain
↓
Save the field observation
↓
Review observations in history and on the local map
Code
The source code is available on GitHub:
github.com/Bangkah/EcoID/tree/main
The source code is released under the MIT License. Dataset images retain their original licenses and attribution requirements.
The main parts of the project are:
-
app/ai/— preprocessing, inference, model contracts, and result post-processing -
app/observation/— drafts, verification, EXIF metadata, maps, statistics, and exports -
app/storage/— local SQLite persistence and image storage -
app/ui/— the local HTTP server and browser interface -
scripts/— dataset preparation, training, export, evaluation, and benchmarking -
tests/— unit, integration, offline, and browser tests
The architecture is:
How I Built It
Local image preprocessing
Every image follows the same preprocessing path:
- Decode the image
- Apply EXIF orientation
- Convert to RGB
- Resize to
224x224 - Normalize with ImageNet mean and standard deviation
- Convert to a
float32tensor with shape(1, 3, 224, 224)
Keeping preprocessing in one shared module helps avoid train/serve skew between the training pipeline and the production inference path.
Open-source AI and local inference
EcoID uses a MobileNetV3-Small image classifier exported to ONNX and executed with ONNX Runtime on the CPU.
The UI and storage layers depend only on a small model contract, so the inference backend can be replaced without rewriting the rest of the application.
The output is converted into a Top-3 list of candidates. If the confidence score is below the configured threshold, the result is marked as LOW_CONFIDENCE.
The initial threshold was calibrated using validation data. The current configuration is:
threshold = 0.73
On the current validation set, this produced:
- 53.7% coverage
- 97.5% selective accuracy
- 80.0% negative rejection
These numbers are provisional and should be recalibrated when the field dataset grows.
Training results
The first real dataset download produced 2,447 usable licensed images, approximately 370 MB in total. The dataset was collected from iNaturalist using CC0, CC BY, and CC BY-SA images, with attribution metadata preserved.
The validated dataset contains approximately 2,000 training images across the five target classes, 300 validation images, and 149 negative samples across lookalike, non-plant, and low-quality groups.
The dataset pipeline includes:
- license metadata
- author attribution
- photographer-based split separation
- duplicate and near-duplicate checks
- validation and evaluation splits
- negative examples
- dataset structure validation
The initial MobileNetV3-Small training run achieved:
- best validation accuracy: 80.3%
- five target plant classes
- local CPU inference after ONNX export
On a separate 50-image evaluation set (10 images per class), the model achieved:
- Top-1 accuracy: 86.0% (43/50)
- Top-3 accuracy: 94.0% (47/50)
- negative rejection: 77.1% (84/109)
- selective accuracy at threshold
0.73: 94.1% - coverage: 68.0%
This evaluation set was collected from licensed iNaturalist images and is useful as a reproducible benchmark. It is not a substitute for a field evaluation with locally captured phone photos.
The ONNX export was verified against the PyTorch model:
- maximum logit difference: approximately
1.9e-6 - Top-1 agreement:
20/20samples
The repository also generates contact sheets for mandatory manual label review. Automated dataset validation passed with zero errors, but a complete human review of all contact sheets remains a dataset-quality task.
Human verification
A prediction is first stored as a temporary draft. It becomes a permanent observation only after the user makes a verification decision.
This prevents an abandoned or uncertain prediction from silently becoming part of the observation history.
The flow is:
photo → identification draft → human verification → saved observation
Draft photos that are abandoned are automatically removed after 24 hours.
Offline storage
The application stores everything locally:
~/.ecoid/
├── ecoid.db
├── images/
├── thumbs/
└── drafts/
The application does not require a cloud API for inference. The browser UI also uses a strict Content Security Policy and does not load external scripts, fonts, or images.
GPS is opt-in. A user can attach a photo's EXIF location, use device geolocation, or enter coordinates manually.
Performance
The model was benchmarked on a Windows laptop using CPU inference:
| Stage | Average |
|---|---|
| Decode and resize | 6.3 ms |
| ONNX inference | 5.8 ms |
| Softmax and Top-K | 0.4 ms |
| Total | 12.4 ms |
The measured p95 total latency was 33.8 ms per image, below the project's target of 10 seconds per image.
Testing
The project includes:
- Python 3.11 and 3.12 test coverage
- Ruff linting
- byte-compilation checks
- tests with non-UTF-8 default encoding
- browser tests with real headless Chromium
- offline tests that block non-loopback sockets and DNS lookups
- UI tests for capture, verification, history, map, statistics, export, and deletion
The current local test result is:
222 tests passed
25 browser tests skipped unless browser testing is enabled
The CI workflow is defined in .github/workflows/ci.yml.
Why Does Open Innovation Matter?
Plant identification is a good example of why open innovation matters.
A closed cloud API could return a label quickly, but it would also introduce several limitations:
- photos would leave the device
- the application would depend on an internet connection
- the model's behavior would be difficult to inspect
- users would have less control over their field data
- the tool might not work in remote areas
Using open-source tools and local inference makes a different design possible. The complete pipeline can be inspected and adapted:
- image preprocessing is visible
- model inputs and outputs are explicit
- confidence thresholds can be calibrated
- observations remain on the user's device
- the backend can be replaced
- the application can run without a reliable internet connection
Open innovation also makes the project easier to learn from. Someone can inspect the training scripts, run the synthetic smoke test, replace the model, add new plant classes, or adapt the observation workflow for another field research problem.
The goal is not to pretend that a small classifier can replace botanical expertise. The goal is to build a transparent tool that helps people notice, record, and learn from the plants around them.
Limitations and Next Steps
This is an initial working model, not a finished botanical identification system.
The most important remaining limitation is domain coverage. The 50-image independent evaluation set is sourced from licensed iNaturalist images, while a dedicated evaluation set of locally captured Indonesian field photos is still needed.
The next steps are:
- Add Indonesian field photos and compare them with the current 50-image iNaturalist benchmark.
- Perform the manual visual quality-control pass on all contact sheets.
- Remove or document near-duplicate training images.
- Re-run calibration using the expanded validation set.
- Re-run final evaluation on a fixed, independent test set.
- Record a physical field demonstration with the trained model.
My Agent Session
This project was developed with an AI coding assistant using Copilot SDK in VS Code.
The agent helped inspect the repository, understand the architecture, connect the AI pipeline to the local observation workflow, run the dataset pipeline, train the initial model, verify ONNX parity, calibrate the confidence threshold, and validate the test suite.
The agent session is optional for this submission and is not linked here.
Prize Categories
I am entering the overall Hacktoberfest Open-Source AI Challenge.
I am not entering a partner category because EcoID does not currently use a partner-specific technology.

Top comments (0)