This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
built DataPilot, an open-source autonomous data scientist for tabular datasets.
I built DataPilot, an open-source autonomous data scientist for tabular datasets.
I built it for a friend who works with data and often starts with the same frustrating workflow:
Upload CSV → inspect columns → check missing values → perform EDA → train a baseline → try another model → compare results → figure out what to investigate next.
A lot of this work is repetitive, and the bigger problem isn't running a model — it's knowing what to investigate next and why.
DataPilot tries to solve that.
You give it a CSV and a target column, and it combines deterministic Python data-science tools with a local open-weight LLM to investigate the dataset.
The goal isn't to build another "chat with your CSV" application.
Instead, DataPilot follows an investigation loop:
Dataset
↓
Profile
↓
EDA
↓
Hypothesis
↓
Experiment
↓
Evidence
↓
Next Investigation
↓
Final Report
The AI scientist decides what it should investigate next, while Python tools perform the actual data analysis and calculate the metrics.
This separation is important:
The LLM reasons about the investigation.
Python produces the evidence.
For example, on my loan-default dataset, DataPilot found:
credit_score → correlation: -0.1472, p-value: 0.00096
income → correlation: -0.1112, p-value: 0.01286
debt_ratio → correlation: 0.0619, p-value: 0.167
age → correlation: -0.0132, p-value: 0.768
It then compared a Logistic Regression baseline against TabPFN.
LogisticRegression
Accuracy: 0.770
Weighted F1: 0.6786
TabPFN
Accuracy: 0.775
Weighted F1: 0.6768
Rather than simply saying "TabPFN is better", DataPilot's evidence evaluator determines that there is no clear overall improvement because the models disagree slightly on accuracy vs. F1.
That evidence-first behavior is the foundation I'm building around.
Demo
The project currently runs locally with an open-weight model through Ollama.
Example:
git clone https://github.com/Manawariqbal/datapilot.git
cd datapilot
python test_agent.py
The agent runs the investigation and produces a scientific report from the generated evidence.
I'm also working toward a Streamlit interface where the workflow becomes:
Upload CSV
↓
Select target
↓
Run DataPilot
↓
Watch investigation
↓
Read evidence-backed report
Code
https://github.com/Manawariqbal/datapilot
The project is intentionally built as a modular open-source system rather than one large agent prompt.
Current structure:
datapilot/
├── app.py
├── requirements.txt
├── README.md
├── data/
│ └── loan_default_sample.csv
├── tests/
│ ├── test_profiling.py
│ └── test_eda.py
└── datapilot/
├── profiling.py
├── problem_detector.py
├── models.py
├── tabpfn_tool.py
├── eda.py
├── experiments.py
├── evidence.py
└── agent/
├── state.py
├── llm.py
├── nodes.py
├── graph.py
└── report.py -->
How I Built It
The core of DataPilot is built around open-source AI and local inference.
Open-weight AI
I'm using:
Gemma 3 4B
running locally through Ollama.
The LLM acts as the scientific reasoning layer.
It decides things such as:
What should I investigate next?
What hypothesis should I test?
Is the current evidence sufficient?
Should I continue or produce a report?
But the LLM does not calculate the statistical metrics itself.
Agent orchestration
I'm using LangGraph to orchestrate the investigation.
The architecture is roughly:
┌───────────────┐
│ Gemma 3 4B │
│ Scientist │
└───────┬───────┘
│
Next investigation
│
┌─────────────────┼─────────────────┐
↓ ↓ ↓
EDA Tool Baseline Tool TabPFN Tool
│ │ │
└─────────────────┼─────────────────┘
↓
Evidence Evaluator
│
↓
Scientific Report
The Python tools are deterministic.
For example:
Pandas performs data analysis.
SciPy calculates statistical relationships.
scikit-learn trains the baseline.
TabPFN performs the foundation-model experiment.
Python calculates the evaluation metrics.
The LLM interprets the resulting evidence.
This prevents the agent from simply inventing numbers.
Machine Learning
The current stack includes:
Python
Pandas
NumPy
SciPy
scikit-learn
TabPFN
LangGraph
LangChain
Ollama
Gemma 3
Streamlit
Pydantic
One of the experiments compares:
Logistic Regression
vs.
TabPFN
The project is designed so additional experiments can become new tools.
Evidence-driven reasoning
This is probably the part of the project I'm most interested in.
Instead of:
LLM:
"Credit score seems important."
I want DataPilot to produce:
Hypothesis:
Credit score is associated with loan default.
Experiment:
Calculate Pearson correlation and significance.
Evidence:
r = -0.1472
p = 0.00096
Interpretation:
A weak negative association is present.
Next investigation:
Determine whether the relationship
has predictive value beyond linear association.
That creates a much more useful agent:
Hypothesis
↓
Experiment
↓
Evidence
↓
Reasoning
↓
Next hypothesis
Why Does Open Innovation Matter?
This project is specifically designed around the idea that the AI reasoning layer should be replaceable and locally runnable.
Using an open-weight model makes it possible to run the scientist locally instead of sending the dataset to a proprietary API.
That matters for data-science workflows because datasets can contain sensitive information.
With the current architecture, the workflow can look like:
Private Dataset
↓
Local Python Analysis
↓
Local Gemma
↓
Local Investigation
↓
Local Report
No proprietary LLM API is required for the core reasoning loop.
Open models also make experimentation much easier.
I can change the reasoning model without redesigning the data-science tools:
Gemma
↓
Qwen
↓
Another open-weight model
while keeping the same:
EDA
Models
TabPFN
Evidence
LangGraph
architecture.
That's the part of open innovation I wanted to explore: AI agents should not have to be tied to one closed model provider.
What I Want DataPilot to Become
The current version is the foundation.
The next major step is making the investigation genuinely autonomous.
Instead of the agent choosing between a fixed set of tools, I want it to discover increasingly specific investigations:
Dataset
↓
EDA
↓
"Credit score has a significant relationship"
↓
Hypothesis
↓
Feature analysis
↓
Evidence
↓
"Linear relationship is weak"
↓
New hypothesis
↓
Non-linear analysis
↓
Evidence
↓
Model experiment
↓
Final scientific report
Eventually, every run should leave behind reproducible artifacts:
runs/
└── run_20261005/
├── profile.json
├── hypotheses.json
├── experiments.json
├── model_results.json
├── charts/
└── report.md
So DataPilot isn't just an AI that talks about your dataset.
It's an AI that investigates the dataset and leaves behind evidence of what it actually did.
Prize Categories
Gemma
DataPilot uses Gemma 3 4B as its local scientific reasoning model through Ollama.
TabPFN
DataPilot uses TabPFN as one of its model experimentation tools for tabular datasets.
Final Thought
I started this project because I wanted to build something more meaningful than another chatbot wrapper.
A useful data scientist doesn't just answer:
"What does this dataset contain?"
They ask:
"What should I investigate next?"
That's what I'm trying to build with DataPilot.
An open-source AI data scientist that doesn't just answer questions about your data — it investigates them.
Top comments (0)