This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
Gully cricket is street cricket in India. A tennis ball, a lane between two buildings, a brick or a stack of slippers for stumps, and rules that change from one street to the next. In one lane hitting it over the wall is a six. In the next lane it's out, and you go and get the ball.
Somebody always has to keep score, and that somebody spends the match looking at a phone instead of playing.
So I built Gully Umpire. Whoever's scoring holds one big button and shouts what happened: "four!", "wide!", "out, caught by Rohan!", or just "chauka!". It keeps the scorecard, says the score back out loud, and after wickets, boundaries and the end of an over a commentator chimes in, in English or Hinglish. Friends who couldn't make it follow along on a live link, and when it's over you get a match report for the group chat.
The part I like most: you describe your lane's rules the way you'd explain them to a new kid. "Wall ke upar gaya toh out. One tip one hand. Last man akela khelega." The umpire reads that back as switches you can check before the match starts. (One tip one hand: caught one-handed after a single bounce is out. Last man stands: the last batter can bat alone.)
The screen is only really needed for the first minute. After that it stays on by itself, the scorer talks, and the phone talks back. If it's too noisy to talk, there's a pad of big buttons, and every call is read back, so a wrong one is a single "undo" away.
Demo
Try it: gully-umpire-fu6e.onrender.com
It's on a free plan, so the first load after a quiet spell takes about half a minute. Tap Start a match, hold the yellow button and say "chauka", or try something messier like "wide hai aur do run bhaag liye". If you're at a desk, typing the call works too.
Code
abjt01
/
gully-umpire
hacktoberfest'26 w1
Gully Umpire
Shout the score, keep the phone in your pocket. Gully Umpire scores street cricket by voice: someone yells "four!", "wide!" or "out, caught by Rohan!", and it keeps the scorecard, says the score back out loud, and plays commentator after the big moments. It plays by your lane's rules too: over the wall is out, one tip one hand, last man stands.
Try it: gully-umpire-fu6e.onrender.com. It's on a free plan, so the first load after a quiet spell takes about half a minute.
Built for the Hacktoberfest 2026 DEV challenge, week 1: Touch Grass.
How a match goes
- Set up in a minute. Team names, overs, and your lane's rules in plain words ("wall ke upar gaya toh out"). The umpire reads them back as switches you can check. Player names are optional.
- Put the phone down. Whoever's scoring holds the big button and shouts the…
How I Built It
It's a Next.js 16 app with plain CSS, hosted on Render. Every model is open-weight and served free on Groq:
- Ears: Whisper large v3 turbo turns the shout into text.
- Umpire: GPT-OSS 20B turns messy calls into ball events.
- Commentator: Qwen 3.8 27B, one line after the big moments.
- Writer: GPT-OSS 120B writes the match report.
The models don't keep score
That was the first decision, and everything else hangs off it. A match is a setup plus a list of ball events, and the scorecard is whatever you get by replaying them in plain code: runs, strike changes, overs, extras, wickets, all out, targets, results. Undo just drops the last event. The engine has its own unit tests and never talks to a model.
The lane rules live in the engine too. If your lane plays over the wall is out, a six becomes a wicket however it was called. If there's no LBW, an LBW call is ignored. A model can't argue its way around that.
The models only turn words into events. Simple calls ("four", "chauka", "do run", "wide", "bowled") are matched in code and scored instantly. Only the messy ones go to GPT-OSS, and whatever comes back is checked before the engine sees it. A typical call takes well under a second from letting go of the button to hearing the score.
Everything that broke
I tested with real models and synthetic voices: macOS's Indian English and Hindi voices, saying calls into the actual pipeline. Things broke in ways I wouldn't have guessed:
- Whisper hears silence as "Thank you." A quick tap on the button would have been scored as something. Now known silence phrases count as nothing heard, and presses under half a second are ignored.
- "Chauka" came back as "Пока!", in Cyrillic. With a vocabulary hint it became "чauka", half Cyrillic. So the parser now reads Cyrillic letters as Latin before anything else.
- "No ball pe chauka" came back as "No ball, pichauka". If a call has "no ball" and a number word anywhere in it, that's enough.
- "Clean bowled" came back as "Clean ball" and got scored as a dot ball. That one's dangerous, because a wicket just disappears. Now "clean ball" means bowled, and anything that doesn't sound like cricket gets asked again instead of guessed.
- "Wall pe laga, do run" was given out. It means the ball hit the wall and they ran two. Going over the wall is out; hitting it isn't. One example in the prompt fixed it.
- The commentator called a four a "chakka". Now it gets told exactly what the ball was: a four, a six or a wicket.
- The match report crowned the wrong top scorer. Now the code works out the top scorer, the best bowler and the real moments from the ball-by-ball log, and the writer can only use those.
Testing it like a phone would
Chrome's fake microphone recorded pure silence in headless mode, so I couldn't test voice by playing a file into a fake mic. Instead, the test plays the audio through Web Audio into a MediaRecorder inside the page, which is the same encoder a phone uses, and uploads it to the real voice endpoint. That's how I found the Cyrillic chauka.
There's also a proper test suite: 34 unit tests for the engine and the call parser, 15 API tests, and 17 tests in real Chrome. The Chrome tests load every screen at 320 to 1280px wide, in light and dark mode, with the longest names the app allows, Devanagari and emoji, then click through a whole match. GitHub Actions runs all of it on every push, and Render only deploys a commit after it passes.
Why Does Open Innovation Matter?
I could pick a model per job. The app has four jobs: hearing, umpiring, commentary and writing reports. Each one is a different open model behind one setting. When I ran the same moments through GPT-OSS 20B, GPT-OSS 120B and Qwen, Qwen wrote the most natural Hinglish ("Arre bhai! Neel ki delivery, Raj ne pakda! Aman gayab hai") and was the fastest, so it got commentary. Swapping a model is changing one line.
Free changed what I built. Groq serves these models on a free tier, which also comes with quirks you only find by using it. Qwen has an output limit of 1,000 tokens a minute that I couldn't find in the docs. That's far too little for match reports and plenty for one line of commentary, so that's the job it got, with GPT-OSS as a backup. The whole thing runs on Groq's free tier and Render's, which feels right for a game played with a tennis ball and a brick.
Nothing's locked to one provider. The app only speaks the OpenAI-compatible API, including for Whisper. To move it to self-hosted models you'd change a base URL and model names, not the code.
Prize Categories
-
Best Use of Render: the app runs on Render's free tier from a Blueprint (
render.yaml), set to deploy only commits that pass CI. - Best Use of GitHub Copilot, for the GitHub Actions part: every push runs lint, unit tests, API tests and the real-Chrome tests, uploads screenshots when a layout breaks, and gates the Render deploy.
The real test is still ahead: wind, traffic, and ten people shouting at once. That's what the undo button is for.







Top comments (0)