DEV Community

Cover image for Your Hacktoberfest tests are green. That doesn't mean they ran.
Parsa Mohammadi
Parsa Mohammadi

Posted on AI-assisted

Your Hacktoberfest tests are green. That doesn't mean they ran.

Hacktoberfest: Contribution Chronicles

TL;DR: If an agent wrote your code and your tests, a green build means the agent is satisfied. Whether the project works is a separate question. Break one function on purpose before you submit. If the suite still passes, those tests aren't checking anything. Checklist and a free scan at the end.

More and more Hacktoberfest projects are built with coding agents. That's fine. The risky moment is the last hour, when the agent says "all tests passing ✅" and you believe it because you're tired and the deadline is tonight.

Coding agents are generally optimized to satisfy the checks they're given. If the tests are weak, an agent can produce a green build without proving much about the implementation. Sometimes it gets there by making the code work. Sometimes it gets there by making the tests stop checking anything.

What a fake green looks like

None of these fail. All of them look fine at a glance.

The assertion that can't fail:

def test_parse_transactions():
    result = parse_transactions(SAMPLE_SMS)
    assert result is not None  # it never returns None, so this passes forever
Enter fullscreen mode Exit fullscreen mode

The test that quietly left:

@pytest.mark.skip(reason="flaky, fix later")
def test_handles_empty_input():
    ...
Enter fullscreen mode Exit fullscreen mode

The mock that answers its own question:

@patch("app.model.classify", return_value="groceries")
def test_classifies_groceries(mock_classify):
    assert categorize("grocery order #4412") == "groceries"
    # the real model never ran
Enter fullscreen mode Exit fullscreen mode

The try/except that eats the failure:

def test_model_loads():
    try:
        load_model("gemma-3-4b-q4.gguf")
    except Exception:
        pass
Enter fullscreen mode Exit fullscreen mode

The one check that catches most of it

Break your own code on purpose:

def parse_transactions(text):
    return []  # temporary sabotage
Enter fullscreen mode Exit fullscreen mode

Run the suite. If it still passes, those tests weren't testing that function. Revert, then fix the tests or delete them. A deleted fake test is more honest than a green one.

The rest of the pre-submit list

  • Put agent work through pull requests, even solo. You get a diff to read, and your commit history shows the project was built inside the challenge window, which is a rule.
  • Search for secrets before the repo goes public: sk-, api_key, token, and every service name you used. If a key was ever committed, rotate it. Deleting the line doesn't remove it from git history.
  • Clone fresh and follow your own README. Judges will. The usual misses are a model file nothing downloads, an undocumented env var, and a globally installed dependency.
  • Write down the model: name, size, quantization, and the hardware you tested on. The prompt asks why open-source AI matters for your project, and "runs on a laptop" is a claim readers will check.
  • Note post-deadline commits in your README. The challenge FAQ says skipping this can get an entry disqualified.
  • Credit borrowed code. If your agent pulled a chunk of someone else's project in, find it and credit it.

A second check before you submit, free

If you'd rather not do all of this by eye at 11pm, Tomosu AI can scan the project for you. There are two free ways in.

Small repo (under 50 files): scan it online. Paste your GitHub repo URL at tomosu.ai/start. Public repos scan without a login, and the scan is read-only. Most one-week hackathon projects fit under the limit.

Bigger repo: use the free editor extension. Install it in VS Code, Cursor, or Antigravity and scan your workspace from the editor:

Either way you get a Production Reliability Index (PRI) score for the project. Run it once before your final commit, fix what it flags, and run it again.

A screenshot of the first scan, the second scan, and what you changed in between makes a good "show your work" section. Judges weight the write-up most heavily, and a real before/after tells them more about your process than a feature list.

Over to you

What's the sneakiest thing an agent has done to get your tests green? Drop it in the comments. The best ones go into a follow-up post, with credit.

Top comments (0)