<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Brandon Favor</title>
    <description>The latest articles on DEV Community by Brandon Favor (@brandon_favor_cf98fad9596).</description>
    <link>https://dev.to/brandon_favor_cf98fad9596</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4171473%2F681235f7-3b37-4830-8452-54b24c39e465.png</url>
      <title>DEV Community: Brandon Favor</title>
      <link>https://dev.to/brandon_favor_cf98fad9596</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/brandon_favor_cf98fad9596"/>
    <language>en</language>
    <item>
      <title>I Tested Whether Frontier AI Can Do a Texas Real Estate Agent's Desk Work — Here's What 3 Models Taught Me (and What the 4th Taught Me About Benchmarks)</title>
      <dc:creator>Brandon Favor</dc:creator>
      <pubDate>Thu, 08 Oct 2026 13:53:23 +0000</pubDate>
      <link>https://dev.to/brandon_favor_cf98fad9596/i-tested-whether-frontier-ai-can-do-a-texas-real-estate-agents-desk-work-heres-what-3-models-71n</link>
      <guid>https://dev.to/brandon_favor_cf98fad9596/i-tested-whether-frontier-ai-can-do-a-texas-real-estate-agents-desk-work-heres-what-3-models-71n</guid>
      <description>&lt;p&gt;I'm a Texas real estate agent (eXp Realty). I spend my mornings doing desk work that has to be exactly right: the math on investor packages, filtering listings against a buy box, Texas compliance questions where a wrong answer isn't trivia — it's liability. Public leaderboards rank models on math olympiads and coding. They say nothing about whether a model knows when a Texas agent must hand over the IABS form.&lt;/p&gt;

&lt;p&gt;So I built &lt;strong&gt;DeskBench-RE&lt;/strong&gt;: 36 tasks covering the actual desk work, and ran four frontier models against them. The results surprised me — just not the way I expected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The benchmark (public, Apache 2.0, all 36 tasks attached):&lt;/strong&gt; &lt;a href="https://www.kaggle.com/benchmarks/brandonfavor/hunter" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/brandonfavor/hunter&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I ran
&lt;/h2&gt;

&lt;p&gt;DeskBench-RE asks one question: &lt;em&gt;can a frontier AI model perform the desk work of a Texas real estate agent?&lt;/em&gt; Five categories, each with its own scoring:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Tasks&lt;/th&gt;
&lt;th&gt;What it tests&lt;/th&gt;
&lt;th&gt;How it's scored&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Package math&lt;/td&gt;
&lt;td&gt;A1–A10 (10)&lt;/td&gt;
&lt;td&gt;Investor-package arithmetic (commissions, pricing, splits)&lt;/td&gt;
&lt;td&gt;Exact numeric answer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Buy-box filtering&lt;/td&gt;
&lt;td&gt;B1–B8 (8)&lt;/td&gt;
&lt;td&gt;Apply multi-rule listing filters (price, new construction, county, property type)&lt;/td&gt;
&lt;td&gt;Exact set match&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Texas compliance&lt;/td&gt;
&lt;td&gt;C1–C8 (8)&lt;/td&gt;
&lt;td&gt;TREC/IABS licensing and brokerage rules&lt;/td&gt;
&lt;td&gt;Multiple-choice letter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Listing-data extraction&lt;/td&gt;
&lt;td&gt;D1–D6 (6)&lt;/td&gt;
&lt;td&gt;Pull structured fields out of listing blurbs&lt;/td&gt;
&lt;td&gt;Field-by-field JSON match&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Outreach drafting&lt;/td&gt;
&lt;td&gt;E1–E4 (4)&lt;/td&gt;
&lt;td&gt;Cold emails to builders/investors (word cap, CTA, honest tone)&lt;/td&gt;
&lt;td&gt;LLM-judge rubric&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One task asks whether each listing in a small set meets a four-rule buy box ($400K cap, new construction only, single-family, specific Texas counties) — the filtering I do every morning. Another asks the TREC question every Texas agent has heard: when must you provide the Information About Brokerage Services form? (At the first substantive communication about a specific property — get this wrong and it's a compliance problem, not a trivia miss.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Which models I tested
&lt;/h2&gt;

&lt;p&gt;All runs executed on Kaggle between October 6–8, 2026, one run per task per model (144 total attempts):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;gpt-6-astra&lt;/strong&gt; (OpenAI) — 36/36 runs completed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;claude-opus-5-5-default&lt;/strong&gt; (Anthropic) — 36/36 runs completed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;gemini-3.8-flash&lt;/strong&gt; (Google) — 36/36 runs completed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;deepseek-r1-0528&lt;/strong&gt; (DeepSeek) — 30/36 runs completed; 6 errored (see below)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;gemini-3.7-flash&lt;/strong&gt; — public baseline from the benchmark leaderboard (24/36, scored September 2026)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The scores
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Overall&lt;/th&gt;
&lt;th&gt;Math (10)&lt;/th&gt;
&lt;th&gt;Buy-box (8)&lt;/th&gt;
&lt;th&gt;Compliance (8)&lt;/th&gt;
&lt;th&gt;Extraction (6)&lt;/th&gt;
&lt;th&gt;Outreach (4)&lt;/th&gt;
&lt;th&gt;Cost, all tasks&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5.5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;25/36 (69%)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;2*&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;1*&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;$0.17&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;24/36 (67%)&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;2*&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;1*&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;$0.14&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;24/36 (67%)&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;2*&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;1*&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;$0.13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash (baseline)&lt;/td&gt;
&lt;td&gt;24/36 (67%)&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;2*&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;0*&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek R1 0528&lt;/td&gt;
&lt;td&gt;4/30†&lt;/td&gt;
&lt;td&gt;0†&lt;/td&gt;
&lt;td&gt;0†&lt;/td&gt;
&lt;td&gt;0†&lt;/td&gt;
&lt;td&gt;—†&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;$0.24 (30 runs)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;* &lt;em&gt;Buy-box and extraction scores are polluted by bugs in the benchmark's own expected answers — verified and detailed below. They don't measure model capability on those tasks.&lt;/em&gt;&lt;br&gt;
† &lt;em&gt;DeepSeek's &lt;code&gt;&amp;lt;think&amp;gt;&lt;/code&gt; reasoning trace broke the harness's exact-format extraction — verified below. Not a capability ranking.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I found
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. The benchmark was wrong more often than the models were.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is the headline. Eleven of the 36 tasks are effectively unpassable — not because the models failed the desk work, but because the grader contradicts its own instructions. I verified each one by reading the task files and the models' actual outputs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Buy-box (6 tasks):&lt;/em&gt; the prompt says "reply with ONLY the qualifying addresses, &lt;strong&gt;copied exactly as written above&lt;/strong&gt;" — the listings read "123 Main St, Dallas, Dallas County." GPT-6 did exactly that. The expected answer was &lt;code&gt;123 main st dallas&lt;/code&gt; — street and city only, county stripped. Every model that followed the instructions failed. The two buy-box tasks that "pass" (B3, B8) are the ones where no listing qualifies and the answer is NONE.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Extraction (5 tasks):&lt;/em&gt; the listing blurbs write "88 Canyon Creek Dr, Forney TX 75126" — no comma between city and state. The expected JSON demands "Forney, TX 75126" — with a comma the source never contained. Every model that faithfully extracted the address failed. (D1 passes because its blurb happens to include the comma.)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The models did the desk work right. My grader marked it wrong. I'm fixing the expected answers and re-running — the benchmark gets better when working agents argue with it, and this is me arguing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. On the tasks that actually measure something, the frontier is a commodity.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Package math: 10/10 across all three complete models. Outreach drafting: 4/4 across all four models. Texas compliance: 8/8 for Claude, 7/8 for the rest. The cheapest model I tested (Gemini 3.8 Flash, $0.13 for all 36 tasks) matched GPT-6 Astra ($0.14) and trailed Claude ($0.17) by a single task. For a working agent, that's the money insight: &lt;strong&gt;you don't need the most expensive model to do desk work — you need any of them, wired up correctly.&lt;/strong&gt; Total inference cost for the entire 36-task evaluation ran under twenty cents per model. The model is not the expensive part of this pipeline. I am — my time verifying the answers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. DeepSeek knew the answers and still "failed" — blame the parser, not the model.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;DeepSeek R1's 4/30 looks catastrophic until you read its actual output. On task A1 (7 homes × $385,000 × 3% commission), its response begins with a &lt;code&gt;&amp;lt;think&amp;gt;&lt;/code&gt; chain-of-thought trace — "There's a package of 7 homes…" — and the harness's regex grabs the &lt;em&gt;first number in the response&lt;/em&gt;: &lt;strong&gt;7&lt;/strong&gt;. The correct answer, &lt;strong&gt;80850&lt;/strong&gt;, appears three times later in the same response, computed correctly, with the work shown. The model did the math right; the extractor read its thinking.&lt;/p&gt;

&lt;p&gt;The six errored D-tasks are the same story at higher volume: the model can't suppress its reasoning trace to emit bare JSON, so the output never parses. Meanwhile it went 4/4 on the outreach tasks — the only category scored by an LLM judge that reads &lt;em&gt;content&lt;/em&gt; instead of demanding exact format.&lt;/p&gt;

&lt;p&gt;If you're building agent pipelines: your parser matters as much as your model. A reasoning model that "thinks out loud" will fail every exact-format assertion you write, no matter how smart it is. That's an integration lesson, not a model ranking — so I'm not ranking DeepSeek here.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The failures that would cost real money are the scattered compliance misses.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Two models missed one Texas compliance question each: GPT-6 on earnest-money timing under the TREC 1-4 contract (C3), Gemini 3.8 and 3.7 on the TREC-promulgated-forms rule (C5). In production, that's a wrong answer to a client or a counterparty on something with legal teeth. The math and drafting categories are safe to delegate with a spot-check; the compliance category is where I'd keep a human in the loop the longest. One wrong TREC answer costs more than every cent of inference in this entire evaluation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. What surprised me most.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I expected the story to be "frontier models ranked on real-estate desk work." Instead the story is: the models are interchangeable on the work, the cheap one ties the expensive ones, and the most valuable thing the benchmark produced was &lt;em&gt;its own bug report&lt;/em&gt;. I set out to grade the machines and ended up grading my test. That's not a failure of the exercise — it's the exercise working. A benchmark you can't argue with is a press release.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Fix the 11 expected answers (buy-box address format, extraction phantom commas), re-run all models, republish scores.&lt;/li&gt;
&lt;li&gt;Add a trace-tolerant extraction mode so reasoning models aren't punished for thinking.&lt;/li&gt;
&lt;li&gt;Next task packs: multi-turn negotiation, longer-context listings, agent-with-tools variants (county appraisal district lookups, TREC rule retrieval).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Benchmark:&lt;/strong&gt; &lt;a href="https://www.kaggle.com/benchmarks/brandonfavor/hunter" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/brandonfavor/hunter&lt;/a&gt; — public, Apache 2.0, all 36 tasks and scoring logic inspectable.&lt;/li&gt;
&lt;li&gt;Every number in the table above comes from a Kaggle benchmark run (October 6–8, 2026); per-task pass/fail is on the task pages. DeepSeek's D1–D6 runs show as errored after repeated retries — no valid result, not a zero.&lt;/li&gt;
&lt;li&gt;Total spend: under $1 of Kaggle inference across all models. Zero paid top-ups.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Built for the DEV × Kaggle Benchmarking Challenge.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>llm</category>
      <category>benchmarking</category>
    </item>
  </channel>
</rss>
