DEV Community

Brandon Favor
Brandon Favor

Posted on

I Tested Whether Frontier AI Can Do a Texas Real Estate Agent's Desk Work — Here's What 3 Models Taught Me (and What the 4th Taught Me About Benchmarks)

Kaggle Benchmarking Challenge Submission

I'm a Texas real estate agent (eXp Realty). I spend my mornings doing desk work that has to be exactly right: the math on investor packages, filtering listings against a buy box, Texas compliance questions where a wrong answer isn't trivia — it's liability. Public leaderboards rank models on math olympiads and coding. They say nothing about whether a model knows when a Texas agent must hand over the IABS form.

So I built DeskBench-RE: 36 tasks covering the actual desk work, and ran four frontier models against them. The results surprised me — just not the way I expected.

The benchmark (public, Apache 2.0, all 36 tasks attached): https://www.kaggle.com/benchmarks/brandonfavor/hunter

What I ran

DeskBench-RE asks one question: can a frontier AI model perform the desk work of a Texas real estate agent? Five categories, each with its own scoring:

Category Tasks What it tests How it's scored
Package math A1–A10 (10) Investor-package arithmetic (commissions, pricing, splits) Exact numeric answer
Buy-box filtering B1–B8 (8) Apply multi-rule listing filters (price, new construction, county, property type) Exact set match
Texas compliance C1–C8 (8) TREC/IABS licensing and brokerage rules Multiple-choice letter
Listing-data extraction D1–D6 (6) Pull structured fields out of listing blurbs Field-by-field JSON match
Outreach drafting E1–E4 (4) Cold emails to builders/investors (word cap, CTA, honest tone) LLM-judge rubric

One task asks whether each listing in a small set meets a four-rule buy box ($400K cap, new construction only, single-family, specific Texas counties) — the filtering I do every morning. Another asks the TREC question every Texas agent has heard: when must you provide the Information About Brokerage Services form? (At the first substantive communication about a specific property — get this wrong and it's a compliance problem, not a trivia miss.)

Which models I tested

All runs executed on Kaggle between October 6–8, 2026, one run per task per model (144 total attempts):

  • gpt-6-astra (OpenAI) — 36/36 runs completed
  • claude-opus-5-5-default (Anthropic) — 36/36 runs completed
  • gemini-3.8-flash (Google) — 36/36 runs completed
  • deepseek-r1-0528 (DeepSeek) — 30/36 runs completed; 6 errored (see below)
  • gemini-3.7-flash — public baseline from the benchmark leaderboard (24/36, scored September 2026)

The scores

Model Overall Math (10) Buy-box (8) Compliance (8) Extraction (6) Outreach (4) Cost, all tasks
Claude Opus 5.5 25/36 (69%) 10 2* 8 1* 4 $0.17
GPT-6 Astra 24/36 (67%) 10 2* 7 1* 4 $0.14
Gemini 3.8 Flash 24/36 (67%) 10 2* 7 1* 4 $0.13
Gemini 3.7 Flash (baseline) 24/36 (67%) 10 2* 7 0* 4 —
DeepSeek R1 0528 4/30† 0† 0† 0† —† 4 $0.24 (30 runs)

* Buy-box and extraction scores are polluted by bugs in the benchmark's own expected answers — verified and detailed below. They don't measure model capability on those tasks.
† DeepSeek's <think> reasoning trace broke the harness's exact-format extraction — verified below. Not a capability ranking.

What I found

1. The benchmark was wrong more often than the models were.

This is the headline. Eleven of the 36 tasks are effectively unpassable — not because the models failed the desk work, but because the grader contradicts its own instructions. I verified each one by reading the task files and the models' actual outputs:

  • Buy-box (6 tasks): the prompt says "reply with ONLY the qualifying addresses, copied exactly as written above" — the listings read "123 Main St, Dallas, Dallas County." GPT-6 did exactly that. The expected answer was 123 main st dallas — street and city only, county stripped. Every model that followed the instructions failed. The two buy-box tasks that "pass" (B3, B8) are the ones where no listing qualifies and the answer is NONE.
  • Extraction (5 tasks): the listing blurbs write "88 Canyon Creek Dr, Forney TX 75126" — no comma between city and state. The expected JSON demands "Forney, TX 75126" — with a comma the source never contained. Every model that faithfully extracted the address failed. (D1 passes because its blurb happens to include the comma.)

The models did the desk work right. My grader marked it wrong. I'm fixing the expected answers and re-running — the benchmark gets better when working agents argue with it, and this is me arguing.

2. On the tasks that actually measure something, the frontier is a commodity.

Package math: 10/10 across all three complete models. Outreach drafting: 4/4 across all four models. Texas compliance: 8/8 for Claude, 7/8 for the rest. The cheapest model I tested (Gemini 3.8 Flash, $0.13 for all 36 tasks) matched GPT-6 Astra ($0.14) and trailed Claude ($0.17) by a single task. For a working agent, that's the money insight: you don't need the most expensive model to do desk work — you need any of them, wired up correctly. Total inference cost for the entire 36-task evaluation ran under twenty cents per model. The model is not the expensive part of this pipeline. I am — my time verifying the answers.

3. DeepSeek knew the answers and still "failed" — blame the parser, not the model.

DeepSeek R1's 4/30 looks catastrophic until you read its actual output. On task A1 (7 homes × $385,000 × 3% commission), its response begins with a <think> chain-of-thought trace — "There's a package of 7 homes…" — and the harness's regex grabs the first number in the response: 7. The correct answer, 80850, appears three times later in the same response, computed correctly, with the work shown. The model did the math right; the extractor read its thinking.

The six errored D-tasks are the same story at higher volume: the model can't suppress its reasoning trace to emit bare JSON, so the output never parses. Meanwhile it went 4/4 on the outreach tasks — the only category scored by an LLM judge that reads content instead of demanding exact format.

If you're building agent pipelines: your parser matters as much as your model. A reasoning model that "thinks out loud" will fail every exact-format assertion you write, no matter how smart it is. That's an integration lesson, not a model ranking — so I'm not ranking DeepSeek here.

4. The failures that would cost real money are the scattered compliance misses.

Two models missed one Texas compliance question each: GPT-6 on earnest-money timing under the TREC 1-4 contract (C3), Gemini 3.8 and 3.7 on the TREC-promulgated-forms rule (C5). In production, that's a wrong answer to a client or a counterparty on something with legal teeth. The math and drafting categories are safe to delegate with a spot-check; the compliance category is where I'd keep a human in the loop the longest. One wrong TREC answer costs more than every cent of inference in this entire evaluation.

5. What surprised me most.

I expected the story to be "frontier models ranked on real-estate desk work." Instead the story is: the models are interchangeable on the work, the cheap one ties the expensive ones, and the most valuable thing the benchmark produced was its own bug report. I set out to grade the machines and ended up grading my test. That's not a failure of the exercise — it's the exercise working. A benchmark you can't argue with is a press release.

What's next

  • Fix the 11 expected answers (buy-box address format, extraction phantom commas), re-run all models, republish scores.
  • Add a trace-tolerant extraction mode so reasoning models aren't punished for thinking.
  • Next task packs: multi-turn negotiation, longer-context listings, agent-with-tools variants (county appraisal district lookups, TREC rule retrieval).

Reproduce it

  • Benchmark: https://www.kaggle.com/benchmarks/brandonfavor/hunter — public, Apache 2.0, all 36 tasks and scoring logic inspectable.
  • Every number in the table above comes from a Kaggle benchmark run (October 6–8, 2026); per-task pass/fail is on the task pages. DeepSeek's D1–D6 runs show as errored after repeated retries — no valid result, not a zero.
  • Total spend: under $1 of Kaggle inference across all models. Zero paid top-ups.

Built for the DEV × Kaggle Benchmarking Challenge.

Top comments (0)