DEV Community

Cover image for Teaching small models to read job postings: our arena and the JOA Small, Advance and Deep fine-tunes
Loukas Tzekos
Loukas Tzekos

Posted on

Teaching small models to read job postings: our arena and the JOA Small, Advance and Deep fine-tunes

Disclosure: Job Opportunities API (JOA, jobopportunitiesapi.org) is an independent data business that sells API access to job-posting data. AI helped run the experiments, check the numbers and draft this text; Loukas (Luca) Tzekos is editorially responsible. Contact: hello@jobopportunitiesapi.org.

The problem

Job Opportunities API (JOA) collects job postings from employers' own recruiting systems and careers pages and re-checks them every day. A posting arrives as a page of text written by the employer. Our customers want fields: the title, the company, where the job is, the pay and its period, the seniority, whether it is remote, when it was posted and whether the poster is the employer or a staffing agency.

Large models do this well, but running one over every posting, every day, is slow and expensive, and the best of them reason for thousands of tokens per posting. We wanted to know whether a small model trained for this one task could do as well. To answer that honestly we first needed a fair test.

The arena

The arena is 250 real job listings from JOA, 11 fields each: title, company, location, salary currency, minimum, maximum and period, seniority, remote type, posted date, and agency or not.

  • Same input for every model: the same prefetched page text and structured data, and up to 24,000 tokens to answer. A reply cut off mid-thought counts as no answer.
  • Reference labels were built with Claude Opus agents that read our stored text and the live page, with a confidence level per field.
  • Two judges, Kimi-Linear-48B and Seed-OSS-36B, graded every answer field by field; a cell scores the mean of the two.
  • What counts: the main rule leaves out values that appear nowhere in the model's input; a missing answer where the value could be found scores 0.
  • Uncertainty: every score has a 95% range from a paired bootstrap over listings. Re-running one model moved its score by up to 1.5 points, so smaller gaps are noise.

By 25 September 2026, 68 open-weight models and their quantised variants had run it. The leader, Qwen3.5-27B, scored 89.2, with ten others statistically tied, up to a 235-billion-parameter model. Reasoning models scored higher but spent thousands of tokens per listing, and size was not destiny: granite-3.3-8b-instruct scored 86.4, close to models many times its size.

The fine-tunes

Qwen3.5-27B became the trainer. It extracted the same fields from a large set of JOA listings, and a refinement step kept only answers grounded in the text (a title or a salary that actually appears in the posting) while keeping the original share of agency postings and empty fields. None of the 250 arena listings was in the training data, and the Claude-built reference labels were used only to evaluate, never to train.

We trained LoRA adapters on three open base models:

Name Base model (licence)
JOA Small ibm-granite/granite-4.2-3b (Apache 2.0)
JOA Advance ibm-granite/granite-3.3-8b-instruct (Apache 2.0)
JOA Deep Qwen/Qwen3.8-27B (Apache 2.0)

Results

Arena, 84 models, graded 29 September (Advance) and 2 October 2026 (Small, Deep):

Rank Model Score 95% range Median tokens per answer
1 JOA Small 93.3 92.2–94.4 117
2 JOA Advance 92.7 91.6–93.9 123
3 JOA Deep 92.2 91.0–93.4 111
4 Trainer, Qwen3.5-27B 89.2 86.9–91.4 2,946
71 Stock granite-4.2-3b 64.5 59.7–69.0 4,411

Three things we did not expect:

  1. The smallest model ranks first. Fine-tuning lifted granite-4.2-3b from rank 71 to rank 1. Small and Advance are statistically tied; Deep is just behind. Once the task was learned, size stopped mattering on this test.
  2. The students beat their teacher. Partly because the trainer sometimes reasoned until it ran out of tokens, and partly because the training answers were tidied into one consistent form. On the listings the trainer did answer, Advance was still 1.9 points ahead (95% range +0.5 to +3.5).
  3. They answer in about 110–120 tokens. The trainer needs about 2,900, almost all of it reasoning.

A second, stricter test used 250 different, recent listings, each verified twice by Claude Opus against the live page and the application flow. On exact match, all three fine-tunes scored about 85%, the trainer about 86% and the stock base models 71–82% (2 October 2026). Here the teacher is still slightly ahead.

Limits

  • They read what we store, not the live page. Our stored text is usually the description body; the header, where the title, location and posted date often sit, can be missing.
  • They cannot tell that a job has closed. From stored text alone, all three called a few dead listings open. The fix is better input, not a bigger model.
  • JOA Deep is weak on posted date (72, against 89 for JOA Small).
  • JOA Small cannot drive a tool-using agent. It works when given the input, not when asked to fetch it.
  • The references and the judges are models. Two judges, checked references and confidence ranges reduce that risk; they do not remove it.

Status and what is next

No weights are published yet; the three model pages are status pages. JOA Small runs on a single consumer graphics card, which makes it the candidate for everyday use. Next we will feed the models a richer input (the rendered page, structured data and the recruiting system's own data), fix the posted-date weakness, re-run both tests and then decide what to release.

This work used the EuroHPC supercomputer Discoverer+ in Bulgaria, made available by the EuroHPC Joint Undertaking through an AI Factories Playground access allocation (project EHPC-AIF-2026PG01-1124).

Contact. General: hello@jobopportunitiesapi.org · Legal matters: luca@tzekos.eu · Phone: +30 2311 113 603. Job Opportunities API (JOA) is a sole proprietorship of Loukas Tzekos, Didaskalisis Papathanasiou Vas. 79, 54629 Thessaloniki, Greece. VAT EL117613696.

Top comments (0)