DEV Community

Wellington Filipe
Wellington Filipe

Posted on

trilha: identifying birds by ear on the trail, with no signal

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is my submission for the Hacktoberfest Open-Source AI Challenge: Week 1 — theme: Touch Grass.

What I built

You're on a trail. Something calls from the brush, twice, and then stops. You don't know what it was.

The normal move is to hold your phone at the canopy and wait for a bar of signal. The places worth birding are exactly the places that don't have one. So you get home six hours later with a recording you can no longer place, and it sits in your phone forever.

trilha is a command-line tool that does the whole thing on your laptop, offline. It listens to a recording, names what it heard, and writes the outing into a field journal in prose.

$ trilha listen manha.wav --lat -23.55 --lon -46.63 --place "Parque do Carmo"

  manha.wav
  596 species in range · 5.2s on CPU · week 40

   0.97 ###################  Rufous-bellied Thrush
        Turdus rufiventris · first at 30s · 22 window(s)
   0.94 ###################  Rufous-collared Sparrow
        Zonotrichia capensis · first at 42s · 26 window(s)
   0.92 ##################   Southern Lapwing
        Vanellus chilensis · first at 63s · 4 window(s)
   0.59 ############         Southern Yellowthroat
        Geothlypis velata · first at 78s · 2 window(s)

  Field note
  The presence of the Rufous-bellied Thrush, Rufous-collared Sparrow, and
  Southern Lapwing strongly suggests a mixed habitat of open woodlands and
  grasslands within Parque do Carmo. Next, while on this trail, pay particular
  attention to the calls of the Southern Yellowthroat — they often favour dense
  undergrowth and shrubbery, and their distinctive song is relatively
  high-pitched.
Enter fullscreen mode Exit fullscreen mode

No account, no API key, no upload. 5.2 seconds on CPU.

How it works

Three open pieces, each doing the one thing it's good at.

The ear — BirdNET 3.0, through ONNX on CPU. It knows 11,560 species. It slices the audio into 3-second windows and scores each window independently, which is why the output can tell you a bird was first heard at 78s across 2 windows rather than just handing you a label.

The filter — BirdNET's range model, 14,082 species. This is the part I find most interesting and it's the part nobody demos. Give it latitude, longitude and the week of the year, and it narrows the candidates to what could plausibly be there right now. At -23.55, -46.63 in week 40 that's 596 species instead of 11,560.

That reduction is the difference between a useful answer and a list of everything that has ever had feathers. A classifier asked to choose from 11,560 classes will cheerfully hand you a Himalayan bird in São Paulo. Asked to choose from the 596 that are actually in range in October, it won't.

The voice — Gemma 3, through Ollama, on my machine. It turns the detection list into a few sentences for the journal. It never touches the identification.

That last point is a hard architectural boundary, not a disclaimer: if Ollama isn't running, the note is skipped and the detections still print. The language model is the part with taste in it, so it's the part that's allowed to be absent.

The two bugs that mattered

I had the thing working and almost wrote it up at that point. Then I ran it once more, properly, and read the output instead of just checking that it ran.

The field note came back like this:

  Field note
  Aqui está a nota de campo em português:

  **Nota de Campo – Reserva do Pampa**

  A combinação de corruíras, sabiás-laranjeira e sabiá-poca sugere um habitat
  de matagal aberto...
Enter fullscreen mode Exit fullscreen mode

Two separate problems in one paragraph.

The small one: the model announced what it was about to do, then added a Markdown heading. In a terminal pretending to be a paper journal, **Nota de Campo –** is noise.

The one that actually mattered: Reserva do Pampa. There is no such place in this story. I passed no --place, so the prompt sent bare coordinates — -23.55, -46.63, which is São Paulo. The model read the numbers, decided it knew where that was, invented a reserve name, and put the wrong biome on it. The Pampa is 1,000 km south.

This is the failure mode that would have killed the project. The whole premise of a field journal is that it's a record of where you actually were. One invented place name and the entire artifact becomes untrustworthy — not just that line, all of it, retroactively. A tool that is right about the birds and wrong about the ground is worse than no tool, because you'd believe it.

The root cause was in the prompt contract, not the code path. My original SYSTEM already said "never invent a species that is not listed, and never invent a count" — I had thought about hallucination, but only about the taxonomy. I never constrained the geography. So I added the rule I'd been missing:

Only ever refer to the location exactly as it is given to you. When you are given
coordinates and no place name, leave the place unnamed — never guess a park,
reserve, region, biome or city from them.

Return only the note itself: no preamble, no sign-off, no heading, no title, no
Markdown formatting, no restating of this instruction. Start with the first
sentence of the note.
Enter fullscreen mode Exit fullscreen mode

Same recording, after:

A forte presença de corruíras e sabiás indica um habitat de matagal aberto, possivelmente com árvores frutíferas na região. A identificação do sabiá-poca com menor confiança sugere a possível existência de áreas mais sombreadas ou de vegetação densa próximas. Próximo, procure por cantos altos nas árvores, pois o sabiá-laranjeira tende a se exibir em posições elevadas para anunciar sua presença.

No preamble, no invented place, and it still flags the low-confidence detection as low-confidence.

Then I checked the inverse, because a fix that only works in one direction isn't a fix: with --place "Parque do Carmo" it uses the real name. And one more paranoid check — the English note told me to listen for Southern Yellowthroat, a species I hadn't noticed in the list. Was that a third hallucination? No: it's there at 0.59 confidence, 2 windows. The species rule had been holding the whole time. Only the location was unguarded.

The lesson I'm taking: when you write an anti-hallucination rule, you're enumerating the categories of thing the model must not invent. I enumerated species and counts and felt covered. Place was a category I hadn't thought of, and it was sitting right there in the prompt as a bare pair of floats. Worth asking, of any prompt: what else is in this context window that the model could decide it recognizes?

That session, if you want to watch the diagnosis happen:

Catching two prompt bugs in an offline bird identifier
You

Check how we're doing for Hacktoberfest — we were already building a new project.

Agent

The Week 1 theme is "Touch Grass", and the first example in the prompt is literally "a bird call identifier that works on the trail with no signal". That is exactly what trilha-offline is. Deadline Oct 11, 11:59 PM PDT.

The code is 677 lines across 6 modules and the CLI runs, but there is no README (and pyproject points at one, so the build breaks), no git repo, and no DEV post.

You

What's left to finish?

Agent

Running it end to end first, because a checklist written without running the thing is a guess.

It works — the identification is solid. But the field note has two defects.

  1. Preamble leak: the note opens with "Aqui esta a nota de campo em portugues:" and a Markdown heading, instead of starting with the note.

  2. Hallucinated location — the worse of the two. The note names "Reserva do Pampa". I passed no --place, so the prompt sent bare coordinates (-23.55, -46.63, which is Sao Paulo). The model invented a place name and got the biome wrong too. A field journal that invents where you were destroys its own credibility.

You

Fix it!

Agent

Both defects are in the prompt contract, not the code path — SYSTEM never forbade a preamble, and never constrained how the location may be referred to. Two rules added to SYSTEM in trilha/narrate.py:

  • "Only ever refer to the location exactly as it is given to you. When you are given coordinates and no place name, leave the place unnamed — never guess a park, reserve, region, biome or city from them."
  • "Return only the note itself: no preamble, no sign-off, no heading, no title, no Markdown formatting."

Plus a reinforcement at the end of the user prompt: "Output the note text and nothing else."

Preamble gone, no invented place. Now the inverse case — does it still honour a place when one is given?

Uses the real place. One more check: the note tells the reader to listen for Southern Yellowthroat — is that species actually in the detections, or is it a third hallucination?

It is there at 0.59. No species hallucination — the "never invent a species that is not listed" rule in the original SYSTEM was already holding. Only the location was unguarded.

You

Yes, let's finish this challenge.

Agent

Wrote the README around what the tool actually prints (captured from real runs, not invented), added the MIT LICENSE file that pyproject declared but did not ship, and made the first commit.

A blind test across six countries

I want to be straight about something: I have not yet taken this out on a real trail. The challenge offers bonus points for doing that and I'd rather tell you what I actually did than dress up a demo as a morning in the woods.

What I did instead is arguably harder on the tool. iNaturalist has an open API full of bird observations that carry an audio recording, the exact coordinates, the exact date, and a species identification verified by the community (research grade). That is a labelled test set with real geography attached — and nothing about it was chosen by me to flatter the tool. I pulled eight at random across six countries and ran each one with its own true coordinates and date.

Where Truth trilha's top guess
🇧🇷 Parque Nacional das Emas Melanopareia torquata Collared Crescentchest 0.98 ✅
🇺🇦 Lviv Oblast Corvus corax Northern Raven 0.95 ✅
🇮🇹 Reggio Calabria Certhia brachydactyla Short-toed Treecreeper 0.94 ✅
🇺🇸 Bannock County, Idaho Corvus brachyrhynchos American Crow 0.93 ✅
🇩🇪 Bad Wurzach Rallus aquaticus Water Rail 0.92 ✅
🇩🇪 Berg im Gau Tadorna ferruginea Ruddy Shelduck 0.91 ✅
🇩🇪 Duisburg Larus canus Common Whitethroat 0.52 → Common Gull 0.32 at #2 ⚠️
🇹🇭 Chanthaburi Arborophila cambodiana nothing above 0.10 ❌

Six of eight correct on the first guess, seven of eight in the top three, 3.2–4.9 seconds each on CPU. The range filter shrank the candidate pool differently at every site — 592 species in the Brazilian cerrado in January, 232 in western Ukraine in late November — which is the whole mechanism working as intended.

The two that didn't land are the interesting ones.

Duisburg is a city gull recording, and the tool put a warbler on top of the real answer. The gull is right there at #2. This is the honest shape of the thing: on a messy urban recording with several birds and traffic, you get a ranked list, not an oracle. The confidence bars exist so you can see when it isn't sure, and 0.52 versus 0.32 is visibly not sure.

Chanthaburi is the one that taught me something. Complete miss. So I asked the obvious question — does the model even know this bird? I grepped the acoustic model's label files:

$ grep -i arborophila ~/.local/share/birdnet/.../labels/no.txt
Arborophila cambodiana_khmerhøne
Enter fullscreen mode Exit fullscreen mode

It knows it. All 11,560 labels include the Chestnut-headed Partridge. Then I checked what the range filter allowed at that coordinate:

$ trilha here --lat 12.9257 --lon 102.1796 --date 2024-03-16 --limit 0 | grep -i partridge
  Green-legged Partridge  (Tropicoperdix chloropus)
Enter fullscreen mode Exit fullscreen mode

429 species in range, and Arborophila cambodiana is not one of them. The ear could have identified this bird. The filter — the component I'd just finished praising as the thing that makes the tool trustworthy — had made the correct answer unreachable. It's a Cardamom Mountains endemic with a tight range, exactly the kind of bird a range map is least confident about, and exactly the kind of bird someone pointing a recorder at a Thai forest most wants named.

The bug the error message confessed

While tracing that, I found this in detect.py:

raise RuntimeError(
    "The range model expects no species at this coordinate. "
    "Check --lat/--lon, or pass --no-range to listen worldwide."
)
Enter fullscreen mode Exit fullscreen mode

There was no --no-range flag. I had written the error message for a feature I never built, and nothing caught it because that code path is hard to reach — you need a coordinate where the range model expects nothing at all.

The Thai partridge turned it from a cosmetic lie into a missing feature with a concrete use case, so I built it. Same recording, filter off:

$ trilha listen th-partridge.wav --lat 12.9257 --lon 102.1796 \
    --date 2024-03-16 --no-range --min-conf 0.1

  11560 species worldwide · 3.6s on CPU · week 11

   0.35 #######              Chestnut-headed Partridge
   0.34 #######              Taiwan Barbet
   0.21 ####                 Taiwan Scimitar Babbler
   0.17 ###                  Malayan Partridge
   0.10 ##                   Blue-throated Barbet
Enter fullscreen mode Exit fullscreen mode

Correct answer, first place. And look at what came with it: Taiwan Barbet at 0.34, a hair behind, from 4,000 km away. Taiwan Scimitar Babbler at 0.21.

That is the trade-off made legible in one screen. With the filter you get 0.98 on a Brazilian crescentchest and a wall in front of a Cambodian partridge. Without it you can reach the partridge, but Taiwan is suddenly a live hypothesis in Thailand. Neither mode is correct in general, so the tool now does what it should have done from the start: default to the filter, and give you a documented way out when you have reason to think it's the thing standing in your way.

The blind test and the Chanthaburi diagnosis, end to end:

A blind test across six countries, and the range filter that hid the answer
You

Get it from the internet!

Agent

I'll do it, but not as a "Taking it outside" section — writing a field report from a downloaded recording would be claiming, publicly and in your name, something that did not happen. That is literally the bug we just fixed in the code.

The honest version is stronger. iNaturalist's open API has bird observations carrying audio, exact coordinates, exact date, and a community-verified species ID (research grade). That is a labelled test set with real geography attached. Run trilha with each recording's true coordinates and date and check whether it gets the known species. A real blind test, not a demo.

xeno-canto v3 needs an API key and the site is behind anti-bot protection. Two closed doors — switching to iNaturalist rather than retrying.

Eight research-grade observations across six countries. Downloading each and running trilha with its own real coordinates and date.

6/8 on the first guess, 7/8 in the top three, 3.2-4.9s each on CPU. The candidate pool varied by site exactly as designed: 592 species in the Brazilian cerrado in January, 232 in western Ukraine in late November.

The Chanthaburi miss is worth diagnosing rather than excusing. Two hypotheses: either the acoustic model does not know this species, or the range filter excluded it. Those are completely different bugs.

Diagnosed. Arborophila cambodiana IS in the acoustic model's 11,560 labels — the ear knows the bird. It is not among the 429 species the range filter allowed at that coordinate. The filter, the component that makes the tool trustworthy, had made the correct answer unreachable. A Cardamom Mountains endemic is exactly the kind of bird a range map is least sure about.

Agent

Found a third bug while tracing it, and this one is in code, not in a prompt.

The error message tells the user to pass --no-range. There is no --no-range flag in the CLI. The message was written for a feature never built, and nothing caught it because the code path is hard to reach — you need a coordinate where the range model expects nothing at all.

The Thai partridge turns it from a cosmetic lie into a missing feature with a concrete use case, so: implement it. use_range threaded through Listener.listen, custom_species_list set to None when off, and a ranged field on Listening so the output says "worldwide" instead of "in range" when the filter is off.

Correct answer in first place at 0.35 — and the trade-off visible in the same screen: Taiwan Barbet a hair behind at 0.34, from 4,000 km away. Neither mode is right in general. Default stays filtered; --no-range is the documented way out.

No regression on the ranged path. Committed and pushed.

Why open matters here

Not as a philosophical preference. This tool does not function without it.

No signal is the normal case, not the edge case. Every model here is a file on disk. You run trilha doctor once at home on wifi, it confirms the weights are cached, and then it tells you: "Models are cached on disk. After this check you can go offline." A hosted API is a hard dependency on a network that, by the nature of the activity, isn't there.

Coordinates are a movement profile. Where you go birding, at what hour, how often, how it changes over a year — that's not a bird fact, it's a map of your habits and your free time. Sending that to someone else's server to get a bird's name back is a bad trade, and it's a trade you'd be making weekly for years. Here it never leaves the machine. There is no server to leak it.

I could read the filter's mind. This is the one I didn't expect to matter and it turned out to be the whole story of the Thai recording. I could grep the label files to confirm the species existed, query the range model directly with trilha here to see what it allowed, find the discrepancy, and ship a flag to work around it — in an afternoon, without filing a ticket with anyone. Against a closed endpoint the entire investigation would have been "it returned nothing, I guess the bird isn't supported." Open weights let you distinguish the model doesn't know this bird from my pipeline threw the answer away. Those are completely different bugs with completely different fixes, and from the outside they look identical.

Swappable voice. --model takes anything Ollama runs. The note is the subjective part, so it should be the replaceable part — gemma3:4b on a laptop in the field, gemma3:12b at a desk with RAM to spare. Same tool, no code change, no vendor to ask permission from.

It costs nothing per use. That blind test was 9 inference runs over 8 recordings, plus every re-run while I was chasing the partridge. On per-call pricing I would have tested three and called it good — and the Chanthaburi miss, the only genuinely instructive result of the day, was the eighth.

Try it

git clone https://github.com/Tinhomagri/trilha-offline
cd trilha-offline
python3 -m venv .venv
.venv/bin/pip install -e .
ollama pull gemma3:4b
trilha doctor
Enter fullscreen mode Exit fullscreen mode

Two xeno-canto recordings are included so you can try it before going outside:

trilha listen samples/corruira-XC436932.mp3 --lat -23.55 --lon -46.63 --lang pt
Enter fullscreen mode Exit fullscreen mode

There's also trilha here --lat --lon, which lists what's in range at a coordinate with no audio at all — useful the night before, to know what you might be about to hear. And trilha journal, which prints every outing you've recorded, or exports it to Markdown.

Code: https://github.com/Tinhomagri/trilha-offline — MIT.

Prize categories

Best Use of Gemma — Gemma 3 through Ollama writes every field note, locally, and it's the component that makes the journal readable rather than a CSV of confidence scores. It's also deliberately scoped: constrained by prompt to the detections it was handed, forbidden from naming the ground, and fully optional at runtime. The identification never depends on it.


Built with BirdNET (birdnet-team/birdnet), Gemma 3 and Ollama. Test recordings from iNaturalist observers under their respective licenses; samples from xeno-canto.

Top comments (0)