DEV Community

Cover image for What I learned building a study platform for Nigerian universities, as a student founder with a team of AI agents
Ngbaronye Nmesirionye
Ngbaronye Nmesirionye

Posted on

What I learned building a study platform for Nigerian universities, as a student founder with a team of AI agents

I'm a student at the Federal University of Technology, Owerri (FUTO), and for the last five months I've been building UniUI: a platform where a Nigerian university student uploads their whole semester (courses, notes, past questions, deadlines, groups) and gets help running it.
It's live at FUTO with 152 students. Public launch is in 50 days. I'm doing most of the engineering with a set of AI agents working in parallel, and I want to write down what I've learned, because most of it wasn't what I expected.
The short version: the hard part wasn't the AI. It was the data, the money, and keeping a fast-growing codebase from eating itself.
The stack
Here's what runs it, layer by layer.
Layer
What I use
Web
Next.js 15 (App Router), TypeScript everywhere, Tailwind, shadcn/ui, TanStack Query, Zustand, react-hook-form with zod
Mobile
React Native with Expo and Expo Router, SQLCipher for encrypted on-device storage
API
Python 3.12, FastAPI (2,000+ routes), Pydantic v2, asyncpg, Alembic, Celery
Data
PostgreSQL 18 with pgvector, pg_cron, pg_trgm and pgcrypto; Supabase for identity; Redis; MinIO; MeiliSearch
AI
Voyage AI embeddings, OpenAI models in three tiers, Whisper, Tesseract, FFmpeg
Payments
A primary provider with Paystack and Monnify as fallbacks, NIBSS for bank transfers
Ops
Docker Compose, Caddy, Cloudflare, Sentry, Prometheus, Grafana, Loki, UptimeRobot
All of it, more than 20 containers, runs on a single Oracle Cloud VPS with 7.8GB of RAM. The database has 237 tables across five schemas. The public marketing site is separate and lives on Vercel.
The real problem is a data problem
The product idea is simple to say. A student asks a question and gets an answer built from their course materials, with sources. Not a generic answer from the internet.
But before any of that works, you need course materials in a shape a machine can use. And Nigerian university material is the opposite of that. Notes live in WhatsApp PDFs, phone photos and handwritten scans. Past question papers get passed down through seniors, often without a course code, a year, or a lecturer's name. Nobody knows which version is current.
So most of my time went into a content pipeline. Here's what's in the system today:
• 4,936 courses catalogued
• 11,071 structured study notes, generated from 11,098 enriched sections
• 21,607 past questions in the question bank
• 3,708 formulas, structured and indexed
• 64,747 content chunks
Two pipelines feed that corpus.
The upload pipeline handles anything a student submits, in ten stages: client-side validation, chunked upload, virus scan, text extraction (Tesseract for images, Whisper for audio, FFmpeg for media, pdftotext), classification, deduplication (SHA256, MinHash and TF-IDF), quality scoring on a six-factor 0 to 100 rubric, human review by a course rep when needed, enrichment, and attribution through the wallet ledger so the uploader gets credit.
The content pipeline turns raw material into study notes in six stages: chunking, processing into sections, enrichment, rich notes, derived content (practice questions, formulas, definitions) and indexing.
Two things surprised me. Quality gates matter more than throughput: 1,951 chunks are quarantined as low quality, and I'd rather lose them than serve them. And embedding everything wasn't the right goal. Only about 15% of the 64,747 raw chunks are embedded, but all 11,128 retrieval sections are, and those are what the AI actually searches.
If I could redo one thing, I'd normalise on write. Section enrichment and rich notes are still stored as JSONB, which is fast to create and painful to query, so I'm writing ETL scripts to pull that data out into normalised tables.
If I had to give one piece of advice from this part, it's to attach provenance to everything from day one. Every chunk carries its course code, level and source. That sounds like bookkeeping, but it's what makes everything downstream possible, because the whole product promise is "this answer comes from your course". If you can't prove where a chunk came from, you can't make that promise.
Answering from a student's own materials
A generic answer can be correct about thermodynamics and still be wrong about what a particular lecturer taught and will test. So retrieval is scoped by course first and only then ranked, and every answer shows its sources.
Here's how it works. Sections and chunks are embedded with Voyage AI at 1,024 dimensions and stored in pgvector with HNSW indexes. When a student asks a question, the orchestrator first loads their context from the Student Graph (a graph of 25 node types and 30 edge types, with edges like enrolled_in, mastered and struggles_with). Then it searches within their course, takes the top five sections, and has the model write an answer from those. A confidence engine scores every answer, and the sources come back with it. [Add what happens when confidence is below your threshold.]
The graph is what makes "study this, here's why" possible. Plain retrieval can find a relevant paragraph. Knowing that this particular student keeps missing questions on one topic is what lets the system point at the right one.
For cost, I split models into three tiers: a nano model for cheap classification and noise filtering, a mini model for daily work, and a larger one reserved for escalation. Repeated queries go through an LLM cache. The lesson I'd pass on is to test models on your own data, not on benchmarks. On my workload the cheapest model beat the expensive one, and I only found out by running both on real questions.
One decision I'm happy with is selling usage by quota instead of by the month. The plans look like this:
• Quick Ask: ₦500 for 24 hours, 10 questions
• Sprint: ₦2,500 for 14 days, 100 questions
• Scholar: ₦2,500 for 30 days, 120 questions
• Scholar Plus: ₦3,500 for 30 days, 250 questions
• Deep Study: ₦10,000 for 30 days, unlimited, with offline access
Quotas make the cost of serving each student predictable. My model projects AI inference at roughly 3% of revenue. That's a projection from my own assumptions, not a measured number, so treat it as one.
Moving money between students
The part that made me most nervous was the wallet. UniUI has a marketplace, a notes store, a task board and tutoring, so real naira moves between students who don't know each other. That means escrow, fees and disputes.
Take the task board. A student posts a task and a reward. The poster pays the reward plus 5%, the worker receives the reward minus 5%, and the platform keeps the difference, which is 10% of the reward. The money sits in escrow until the work is done.
// Simplified and illustrative: not the production code
function settleTask(reward: number) {
const posterPays = reward * 1.05; // 5% fee on top
const workerGets = reward * 0.95; // 5% fee off the top
const platformKeeps = posterPays - workerGets; // 10% of the reward
return { posterPays, workerGets, platformKeeps };
}

// settleTask(2000) -> { posterPays: 2100, workerGets: 1900, platformKeeps: 200 }
Tutoring has a longer flow because the stakes are higher and the work is harder to verify:
PAID (funds held in escrow)
-> SESSION HAPPENS
-> tutor submits a session note
-> both sides confirm attendance
-> 7-day dispute window
-> RELEASED (tutor keeps 80-90%)

any step -> DISPUTED -> RESOLVED
A few other rules came out of thinking about abuse before launch:
• Withdrawals start at ₦10,000 and require a verified bank account.
• Payouts go out in weekly batches, so there's a review step before cash leaves the platform. [Add your reasoning for batching and how you review them.]
• Withdrawal is limited by role. At launch, creators, ambassadors, tutors and a few other approved roles can cash out. Regular users earn tokens, which don't convert to cash on their own.
• Referral commissions (creators, ambassadors, student unions) share the same attribution and fraud rules: last-click with a 30-day window, plus checks for self-referral, duplicate signups and click farming. That leaves one system to harden instead of three.
Underneath, the wallet is an append-only ledger with 31 entry types. Balances are computed from the ledger and never edited in place, so every kobo is traceable. Payments go through a primary provider, with Paystack and then Monnify as fallbacks if one fails, and NIBSS handles direct bank cashouts. Tokens behave differently from naira: the token price has a floor of 40 kobo, floats up to ₦1.20, and reprices every two hours through a pg_cron job.
If you've built escrow or wallets in Nigeria, I'd love to hear what bit you. [Mention anything you've already hit in testing, like webhook retries, idempotency, or reconciliation.]
Building with AI agents
I run a multi-agent setup for development, content generation and operations. Fixed costs are close to zero because of it. It's also how one student built this much in five months.
The honest downsides are real, though. Agents are very good at producing things, and nobody tells you that producing things is its own problem. I ended up with 908 frontend pages and had to consolidate them down to about 150 that are critical for launch. Volume is cheap. Coherence is not.
There was also a lesson that no agent could help with. The mobile app, with 152 screens built, was blocked for weeks on an Apple Developer account. Some blockers aren't code at all.
To keep agent-written code coherent, I lean on constraints. Every frontend page has to use components from one design reference file, and if a pattern isn't in it, I add it there first. TypeScript is strict on the frontend, and Pydantic models sit on every request and state object on the backend. Agents are much better at working inside tight rules than inside vague ones.
I also run a second set of agents that watch the platform instead of building it. I call the layer Jarvis. It has six guardian agents (infrastructure, security, users, growth, content and finance) plus a persona named Treasure, which is the only one that talks to humans. Each agent is a container defined by a YAML profile and a tools module, built with miragen and Pydantic AI for typed tool calls.
A guardian that spots something creates a signal. The core orchestrator scores it, then Open Policy Agent decides whether the action is allowed. Risky actions go to Treasure, which asks me for approval on Telegram, email or in-app. Every decision goes into an audit ledger in Postgres, and resolved incidents are stored in Qdrant so similar ones can be matched later. The policy gate sits outside the agents, so an agent can't talk its way past it.
I try to keep one distinction honest here. A cron job that restarts a container is automation. I only call something autonomy when the system decides what to do, does it, learns from the result and adjusts.
Building for Nigerian constraints
A few product choices exist because of where the users are:
• Naira pricing and small amounts. Quick Ask is ₦500 for a day, because many students will try something before they commit to a month.
• Offline. The top plan includes offline study, because data is expensive and connections are unreliable.
• Phone first. Students live on their phones, which is why the mobile app matters even though the web app is live.
• Curriculum-specific content. Nigerian course codes, Nigerian syllabi and Nigerian exam patterns. A global tool can't fake that.
Offline is the least finished part of the platform, and I'd rather say so. The design is device-first. The app writes to an encrypted local database (SQLCipher) first and adds the change to an outbox. A sync engine then pushes it to the server, which fans it out to the student's other devices and resolves conflicts. For offline AI, the plan is quantised models running on-device through llama.cpp, fetched from Hugging Face at download time instead of stored on my VPS, plus a limited WebLLM path in the browser.
The boring parts took most of the time
Auth, billing, notifications and sync took about 60% of the engineering time. Nobody writes blog posts about them, but everything else stands on them.
One person plus agents isn't the same as one person. The agents built as much as I did, and at this pace it's the only way the work gets done.
What's still hard
Legal is the unglamorous one. The platform stores student data, so NDPA compliance, CAC registration and a proper privacy policy aren't optional, and any future use of aggregated, anonymised learning data depends on getting them right.
The other open question is whether students will pay for a study tool at all. I think the answer is yes if the product proves its value fast, but that's still a hypothesis.
What's next
Public launch is in 50 days, with a planned expansion to six more Nigerian universities in November 2026. I'm also recruiting 500 student creators to help with distribution.
If you've built retrieval over messy educational content, payments in Nigeria, or a product with a lot of agent-written code, I'd like to hear how you approached it. Comment below or find me on X.
• App: app.uniui.com.ng
• Waitlist and creator program: waitlist.uniui.com.ng
• More: uniui.com.ng

Top comments (0)