This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
hackreport is a pipeline that reads public pages about past UK hackathons and reports on the student-led events season by season. For each event it records the organiser, whether it was student-led, and where available the size and number of projects.
I built it for a colleague of mine who was interested in seeing statistics of events across the year.
The problem: there is no single place showing which UK hackathons are student-run, who ran them, and how big they were, and organiser knowledge gets lost as students graduate.
Demo
Check out the results here!
Code
Check out the folder 0 Weekend here:
Run it with
./run.sh. It sets up the environment, starts Ollama, pulls the model if missing, and writes data/report.md and data/events.csv.
How I Built It
- Model: Gemma 3 (4B), run locally through Ollama.
-
Collect: Python collectors read the Hackathons UK season pages and MLH season pages. They then follow each event's own website, plus its About or Team page.
- Requests respect
robots.txt, are rate limited, and are cached. - If a site is dead, the pipeline falls back to the Wayback Machine copy from within the same season.
- Devpost blocks scripted access, so it reads pages you save from a browser.
- Requests respect
-
Extract: Gemma gets each page and returns JSON matching a schema. It is told to use only what the text states and to leave unknowns as
null, with temperature 0 and a verbatim evidence quote for each classification. -
Guardrails: a 4B model guesses student-led from a university name, so code checks the evidence quote.
-
student_led: trueneeds society or student wording in the quote. -
falseneeds company or charity wording. - Anything else is demoted to unclear.
-
- Seasons: seasons run September to August and are named by the ending year (2024 = Sep 2023 to Aug 2024).
- Merge and report: events seen in several sources are merged, with the organiser's own site taking priority. The report is a per-season summary plus a CSV.
Why Does Open Innovation Matter?
- Cost and repeatability: an open-weight model running locally means a pipeline that reads hundreds of pages costs nothing per call and can be re-run whenever the sources change. This is great as it means there is no per-token cost for a volunteer or student organiser who may not have access to LLMs.
- Constrained output: local inference with a JSON schema, temperature 0 and a fixed model version makes extraction reproducible, so anyone can rerun it and check the results.
- Honest limits: because the model is small and open, I had to build the evidence checks. That makes every classification traceable to a quote on a page.
Prize Categories
- Best Use of Gemma
- Best Use of GitHub Copilot
Top comments (0)