DEV Community

Ashen
Ashen

Posted on

No JSON Schema, No Postman, No Effort: A Local Gemma 4 API Tester for My Lazy Friend

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

I built CRUD API Tester, a tool that points at a REST API, reads its OpenAPI spec, and uses a local Gemma 4 model to generate and run test cases for one endpoint at a time. You describe the data you want in plain English, like "Sri Lankan names" or "edge cases: empty strings, age 0 and 200", and it does the rest.

I built it for a friend I met about a year ago on a uni team project. They took the Q&A part of the project, and they're also, with love, frustratingly lazy. They don't want to set up a standard test tool, and they really don't want to hand-write JSON schemas just to check that POST /customers rejects a bad email. They want to point at an API, say what kind of data they'd like, and get test cases back.

So that's the whole tool: pick an endpoint, type a sentence, get classified test cases. If you want, it also fires them at the live API and tells you what passed.

It has two modes:

Generate only: accepts a URL or an OpenAPI file and stops after saving a JSON template.
Run tests: needs a live base URL, generates the cases, sends them with httpx, and writes a pass/fail report next to the template.

It has two interfaces over the same service function: a CLI (python -m tester) and a single-page web UI on port 8080.

Demo

Code

CRUD API Tester

A local tester that reads an OpenAPI document, builds a Pydantic model for one endpoint, and asks Ollama (gemma4:e4b by default) to fill classified test cases. Each case is marked happy, edge, or invalid. You can call the endpoint with that data, or only save it.

Requirements

  • Python 3.11+
  • Ollama running locally
  • The model: ollama pull gemma4:e4b

OLLAMA_MODEL overrides the model name. OLLAMA_HOST overrides the Ollama address (default http://127.0.0.1:11434).

Install

python -m venv .venv
.venv\Scripts\Activate.ps1
pip install -r requirements.txt
Enter fullscreen mode Exit fullscreen mode

Sample API

The Customer CRUD server uses SQLite (sample_api/customers.db) and publishes /openapi.json.

uvicorn sample_api.main:app --port 8000
Enter fullscreen mode Exit fullscreen mode

Endpoints:

  • POST /customers
  • GET /customers with optional city and min_age
  • GET /customers/{customer_id}
  • PUT /customers/{customer_id}
  • DELETE /customers/{customer_id}

Interactive docs: http://localhost:8000/docs

Tester

Two modes, in the terminal or in the browser:

  1. Run tests. Enter a live address such as localhost:8000. The tester…

How I Built It

Model: Gemma 4, specifically gemma4:e4b, running locally through Ollama and the ollama Python client. OLLAMA_MODEL overrides it if you want to swap models.

The core idea: the schema is the contract. My friend hates writing JSON schema, so the tool writes it from the API's own spec and hands it to the model.

Load the spec. openapi_loader.py accepts localhost:8000, a full URL, a direct openapi.json link, or a local file. A bare host gets normalized to http://{host}/openapi.json. It resolves $ref against components.schemas and returns one operation per method and path, with its summary, path params, query params, JSON body schema, and documented status codes.
Build a dynamic Pydantic model. schema_builder.py turns the endpoint's body, path, and query schemas into models using pydantic.create_model, then wraps them:
'''python
class TestCase(BaseModel):
name: str
case_kind: Literal["happy", "edge", "invalid"]
description: str
path_params: PathModel
query_params: QueryModel
body: BodyModel | None
expected_status: int

class TestBatch(BaseModel):
cases: list[TestCase]
'''

case_kind is the classification. Happy and edge cases expect the success codes from the spec, and invalid cases expect 422 or 404. Only common FastAPI shapes are mapped: string, integer, number, boolean, enum, array, nested object, and nullable.

Constrain the model with that schema. TestBatch's JSON schema goes to Ollama as the format parameter. Pydantic is both the constraint and the classifier: every case the model produces is forced into the shape and labeled. The user's prompt says what the data should look like, and the schema says what shape it must take.
Keep it small for a small model. generator.py makes one ollama.chat call per chunk of up to 5 cases, so a 4B-class model stays inside a small schema. Temperature stays moderate so rows actually differ. Each response is validated with TestBatch.model_validate_json, and if validation fails, that chunk is retried once with the error text appended.
Run them. runner.py sends the requests with httpx. For happy and edge cases on GET/PUT/DELETE /customers/{customer_id}, it first POSTs a seed customer and injects the real id, so those cases test real behavior. Invalid cases keep the model's id or body so a 404 or 422 actually means something. A case passes when the response status equals expected_status. The report records status, a short body snippet, latency, and pass/fail.

Every run saves a template to runs/{timestamp}.json with the source, endpoint, prompt, count, and cases, so my friend can reuse it or download it from the web UI.

Checks: parsing, schema building, the sample API, and the web form can all be checked without the model. The generate and run paths need gemma4:e4b pulled locally. The tool never pulls the model automatically, since it's several GB.

Why Does Open Innovation Matter?

Open weights are what made this tool possible for this friend. They don't want accounts, API keys, or a billing page between them and a quick check of an endpoint. With Gemma 4 running through Ollama, there is nothing to sign up for: pull the model once, and everything runs on their own machine.

It also keeps their data local. API specs and test payloads from a real project can be sensitive, and with a local model none of it leaves the laptop.

Finally, structured output is the heart of this project. Passing a Pydantic-derived schema as the format constraint to a small local model, with chunking and a retry, is what makes the output dependable enough to run against a real API. Being able to pick, run, and tune a small open model for that is what let me design the tool around it.

Prize Categories

Gemma 4

Top comments (0)