DEV Community

Daniel Nwaneri
Daniel Nwaneri Subscriber

Posted on AI-assisted

10 Internal Inconsistencies in 3 Published Groundwater Surveys

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

A resistivity reading of 2,950 Ω·m sounds like an answer. It is not. Resistivity measures how strongly the ground resists electric current, and what counts as normal depends on the geology of that exact site. The same number can mean clean groundwater in one place and a warning sign in another.

I studied geophysics at university, and this is a problem I was trained to work on. In the Niger Delta, some surveyors use Vertical Electrical Sounding (VES) to decide where to sink a borehole. They send current into the ground, measure the resistance of each layer, and compare the results with published surveys of the same area. I have also seen the alternative: boreholes sited by artisans with no geology training, based on local folklore or wherever the last well worked. In a region with decades of oil pollution, a bad siting decision can put a community on water that should never have been approved.

Ask my agent about a reading at a named site, and it answers in plain English: normal, anomalous, or uncertain. It also tells you which paper, and which table or figure, each number came from.

While I built it, I found 10 internal inconsistencies inside the 3 published papers it uses. None of them were staged. I checked each one against the source PDFs. The full list, with pages, is on the findings page.

  • A swapped table: the Bori paper's summary table on page 7 swaps the names of two stations. Its own coordinate table on page 3 and its Figure 2 captions show the correct order.
  • Curve labels that contradict their own data: three stations are labeled "A-type." An A-type curve requires resistivity to increase with every layer. The papers' own layer values drop and then rise again. The Choba paper even states the pattern ℓ1<ℓ2>ℓ3<ℓ4<ℓ5 one sentence before it calls the curve A-type.
  • Depth arithmetic that does not add up: at six stations in the Etche paper, a layer's printed depth does not equal the depth above it plus its printed thickness. For example, at Egwi, 6.29 m + 37.95 m = 44.24 m, but the paper prints 37.25 m. Both numbers also appear in the paper's text for Figure 2. So the table and the text are consistent with each other, but not with arithmetic. The paper does not give enough information to decide which number is wrong.

I have contacted the corresponding authors of all three papers.

The agent does not repeat these errors. It reports them, shows both values, and explains which one it trusts and why, or says clearly that it cannot decide.

How a surveyor uses it

A surveyor in Bori finishes a sounding and inverts it. The aquifer layer comes out at 2,950 Ω·m. Before recommending a borehole, they ask: is this normal here? The agent returns what the published surveys of Bori show, flags any published value that is unreliable, and gives the sources to cite in the siting report. At a site with no signal, they run the same check with --offline.

A second user is the consultant who writes the report. If the report cites the Etche paper, the agent shows which of its numbers do not add up, before those numbers go into a recommendation.

Demo

Try it live: ves-interpretation-agent.vercel.app. No login, no setup. The findings page lists all 10 internal inconsistencies; the ask page runs the same rule-based --offline mode as the CLI, live, with example questions ready to click.

The video shows:

  1. The Sanity Dashboard: the Knowledge Base, the 14 issues Context raised, and the standing instructions.
  2. The table-swap question: "I got a 2,950 Ω·m reading at BMGS Bori Field. Is that normal or anomalous, and can I trust it?"
  3. The curve-label question for Choba (full transcript: examples/choba-demo.md).
  4. The same question with --offline: no network and no LLM. It reports "label mismatch," the same conclusion in rule-based words.

Here is a real answer, unedited:

[NORMAL] Yes, 2950 ohm-m is normal for BMGS Bori Field — it matches the trusted figure caption and coordinate table in the paper.

Resistivity measures how strongly a material (like underground rock or water) resists electric current; higher numbers often mean drier or more solid layers, lower numbers suggest wetter or clay-rich layers. For this site, multiple independent parts of the same paper agree on 2950 ohm-m.

SOURCE DISAGREEMENT FOUND
The paper gives two different aquifer resistivity values for BMGS Bori Field: 2950 ohm-m from the figure caption and 3706 ohm-m from the summary table.

  • Claim A: Figure 2 panel caption (IPI2Win-inverted layer table) says 2950 ohm-m
  • Claim B: p.7 Table 1 (summary table), row printed 'BMGS bori field' says 3706 ohm-m
  • Trusted: the figure caption and coordinate table, which say 2950 ohm-m (the paper's own p.3 coordinate table and Figure 2 captions are internally consistent and contradict the swapped labels in the summary table — this is a documented label-swap error)

Document IDs cited: reading-bori-bmgs-bori-field, reading-bori-bmgs-bori-field-p7table. Full run with every tool call and the GROQ query: examples/bori-demo.md.

Code

GitHub logo dannwaneri / ves-interpretation-agent

An agent that answers whether a groundwater VES resistivity reading is normal or anomalous, grounded in real published survey papers via Sanity Context over MCP.

VES Interpretation Agent

An agent that answers whether a groundwater resistivity reading is normal or anomalous for a specific Niger Delta drinking-water survey site, grounded in real, published Vertical Electrical Sounding (VES) papers, queried through Sanity Context over MCP.

Built for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content.

Sanity project ID: c78nb8ch (dataset production, public) Knowledge Base ID: kbIp9dbX1pcY Model: Qwen (qwen-plus, via DashScope)

Try it live: ves-interpretation-agent.vercel.app, no login or setup needed. A static findings page lists all 10 internal inconsistencies; an ask page runs the same rule-based --offline mode as the CLI below, live, with example questions ready to click. Source: web/ and api/ask.js, which calls agent/index.js --offline as an unmodified child process, so the deployed page and the CLI can never give different answers.

The problem

The same resistivity reading means something different depending on the…

How I Used Sanity

Structured content: papers, sites, readings

I transcribed the three papers by hand into a Sanity dataset with three document types:

  • surveyPaper: one per paper, with authors, year, and DOI.
  • surveySite: one per survey area, with a reference to its paper.
  • vesReading: one per station, with the layer resistivities in depth order, the thicknesses, the depths, and the exact source location (table or figure, and page).
defineType({
  name: 'vesReading',
  type: 'document',
  fields: [
    defineField({name: 'station', type: 'string'}),
    defineField({name: 'site', type: 'reference', to: [{type: 'surveySite'}]}),
    defineField({name: 'paper', type: 'reference', to: [{type: 'surveyPaper'}]}),
    defineField({name: 'sourceLocation', type: 'string'}), // e.g. "p.7 summary table, row 'Kenpoly sec school field'"
    defineField({name: 'curveTypePublished', type: 'string'}), // the paper's label, checked by curveType.js
    // also: rmsPercent, reportedAquiferResistivityOhmM, reportedAquiferDepthM, reportedAquiferThicknessM
    defineField({name: 'layers', type: 'array', of: [{
      type: 'object',
      fields: [
        defineField({name: 'layerIndex', type: 'number'}),
        defineField({name: 'resistivityOhmM', type: 'number'}),
        defineField({name: 'thicknessM', type: 'number'}),
        defineField({name: 'cumulativeDepthM', type: 'number'}),
      ],
    }]}),
  ],
})
Enter fullscreen mode Exit fullscreen mode

The dataset holds 3 papers, 3 sites, and 24 station readings. The Knowledge Base built from it has 8 entries as of the latest rebuild; this number has changed with every rebuild so far, since Context restructures entries each time, not just their content. The dataset is public, so you can query it yourself. This link runs a live GROQ query against the real production dataset, no login needed, and returns the swap itself: two documents for the same station, 2950 and 3706 Ω·m.

Two endpoints, two jobs

A Sanity Context endpoint serves one mode. So my agent connects to two:

  • GROQ mode (c78nb8ch.production): groq_query returns exact numbers from the structured documents.
  • Knowledge Base mode (kbIp9dbX1pcY): knowledge_base_read returns the explanations and the conflict history.

The rule is simple: numbers come only from groq_query. Explanations come only from knowledge_base_read. Both endpoints serve a tool named initial_context, so the agent prefixes each tool name to keep them apart.

The Knowledge Base holds what the raw data cannot: the decisions. When I resolved the table swap, that decision became part of the Knowledge Base, and every later answer uses it. The offline mode gives the verdict from numbers alone. Only the online agent can also say which conflicts a person already checked and resolved.

flowchart TD
    Q[Surveyor question] --> G[groq_query: exact layer values]
    Q --> K[knowledge_base_read: explanations and conflicts]
    G --> C[Code checks: curve type, depth arithmetic, table swap]
    C --> S[Qwen writes the answer]
    K --> S
    S --> V[Grounding check: drop any number not in the GROQ data]
    V --> A[Verdict, sources, raw numbers]

Here is a real query the agent ran, answering "is 2,950 Ω·m normal for BMGS Bori Field?" (full run):

*[_type == "vesReading" && site->name == "Bori Metropolis" && station == "BMGS-bori-field"]{
  _id, station, sourceLocation, curveTypePublished, rmsPercent, layers,
  reportedAquiferResistivityOhmM, reportedAquiferDepthM, reportedAquiferThicknessM,
  "siteName": site->name, "paperId": paper._ref, "paperTitle": paper->title,
  "paperCitation": paper->citation
}
Enter fullscreen mode Exit fullscreen mode

This returns two documents for the same station. First one only, the rest of the fields are the same shape:

{
  "_id": "reading-bori-bmgs-bori-field",
  "station": "BMGS-bori-field",
  "sourceLocation": "Figure 2 panel caption",
  "reportedAquiferResistivityOhmM": 2950,
  "siteName": "Bori Metropolis"
}
Enter fullscreen mode Exit fullscreen mode

The second document, reading-bori-bmgs-bori-field-p7table, reports 3706 Ω·m from the page-7 summary table instead. That disagreement between two real documents for the same station is the table swap.

Code checks decide the facts. The model explains them.

A paper's curve label is a claim. The layer values are evidence. So the agent never asks the model whether a label is correct. Code decides:

  • curveType.js derives the curve type (A, Q, H, K, and multi-layer forms) from the raw layer values and compares it with the paper's label.
  • depthArithmetic.js checks each printed depth against the depth above plus the printed thickness.
  • grounding.js scans the final answer and removes any conflict that cites a number not present in the GROQ data.

When a table is inconsistent, a fixed code template writes the statement, not the model. Early versions let the model word it, and it sometimes chose which number was "wrong." Now the statement is identical on every run and in every mode.

These checks have 22 unit tests, including tests on the real layer values of the stations above.

What Context caught, and what it could not

Context raised 14 issues while it built the Knowledge Base, across several rebuilds. I checked every one against the PDFs. 6 pointed to real inconsistencies in the papers, and together they cover 3 of the 10: the Bori table swap, the Choba curve label, and, first caught on a later rebuild, the Etche Odufor curve label.

Bori and Choba are contradictions stated in the paper's own text: two documents giving different numbers, or a label contradicting a pattern the paper itself describes. Odufor is different, and worth being precise about. The dip for Odufor is not stated in the paper's own words. It is stated in my own transcription note, written while I entered the data, recording the dip I found in Table 9. I checked the raw dataset: Odufor and Opiro both carry an identical note. Context compared my note against the paper's A-type label and caught the contradiction, but only for Odufor; I checked the current Knowledge Base entry and it correctly lists both stations now, so the gap was specific to that one rebuild's generated entry, not a limit of same-fact detection itself.

The papers themselves never state the other 7 in words. For Opiro, my note exists just like Odufor's, but Context's entry missed it on every rebuild so far. The six depth inconsistencies are never written anywhere, not even in my own notes; they only appear when you add the numbers. No text comparison can find those, so my agent computes them in code. The two methods complement each other.

One issue deserves a name. Context paired "Kenpoly Convocation Arena" and "Kenpoly sec school field" as competing claims about one fact. They are two different stations, both correctly labeled. The only link is the word "Kenpoly." My own agent made the same name-similarity mistake early in this build. I verified the pairing against the paper, deleted the standing instruction it had created, rebuilt the Knowledge Base, and reran the affected questions to confirm nothing changed.

I also used two Knowledge Base features on purpose:

  • Purpose: it names the real users and questions, so the outline tags surveyor topics as [core].
  • Instruction: a rule anchored to the Choba and Etche papers. It says a curve label is a claim, and an A-type label requires rising resistivity at every layer.

Issues screen, Resolved (14)
Instructions screen

Works in the field, without a signal

Survey sites in the Niger Delta often have weak or no network. So the agent has three modes:

Mode Network Knowledge Base LLM
Online (default) Yes, with a token Yes Yes
--public Yes, no token No Yes
--offline None No No, rule-based

--offline runs the same GROQ queries on a local snapshot with groq-js. The code checks run locally, and the verdict comes from rules, not a model. On the 9 evaluation questions, the offline verdict matches the online agent on 8. The one difference is a question that needs language understanding, which the rule-based mode does not have.

The offline mode is the proof: remove the model, and the structured data still gives the verdict.

Judges can run --offline with no keys at all. --public needs no Sanity token, only a Qwen API key.

git clone https://github.com/dannwaneri/ves-interpretation-agent
cd ves-interpretation-agent/agent
npm install
node index.js --offline "Is the Choba site's curve really the type the paper says it is?"
Enter fullscreen mode Exit fullscreen mode

Evaluation, and what I learned from it

I compared the agent with a plain keyword-search baseline: BM25 over the raw text of the three PDFs, with the same model and a similar prompt. 9 questions, 3 runs each:

Structured agent Keyword baseline
Correct verdict 27/27 15/27
Correct conflict detection 27/27 23/27

This is not a held-out test. Claude Code drafted these questions, I approved them, and I fixed the system against them during development. They show that the agent handles the cases it was built for.

So I also wrote 4 held-out questions myself, with the expected answers written first. Two are in Nigerian Pidgin, the way a field worker might ask. On the first run, the agent answered 3 of 4 correctly. The comparison question ("Between Maakoro-street and Akpoku, which has higher resistivity?") failed: the agent answered about the wrong station. I fixed the question routing, and the same 4 questions now pass. That rerun is not a held-out result, because I fixed the agent after I saw the question.

Limitations

  • Small dataset: 3 papers and 24 station readings, transcribed by hand and checked against the PDFs. Small and verified, not large.
  • The depth tolerance was chosen after inspection. Clean stations show at most 0.300 m of rounding noise. The smallest real inconsistency is 0.795 m. I set the tolerance to 0.5 m after I saw all stations. My first value, 1.0 m, missed a real inconsistency at Akpoku.
  • The table-swap rule is general but tested on one case. The code has no station names in it, but the Bori swap is the only real swap in the data.
  • Earlier failures I fixed: the --public mode once reversed the Bori conclusion. The model once reported a conflict that did not exist. The online agent once had no depth check, so the model guessed which number was wrong. A clean-clone test, run specifically because the post tells judges to clone and run these exact commands, found that every mode crashed with no .env file present, including --offline, which needs no keys at all. Each fix has a commit and a note in the README.
  • Qwen's API is intermittently slow to connect, not down, just unreliable. While preparing this submission, node agent/index.js failed with fetch failed several times in a row. Traced it to UND_ERR_CONNECT_TIMEOUT reaching Qwen's endpoint specifically; both Sanity endpoints succeeded every time in the same test. A second attempt usually goes through. If a judge hits this, it is the network path to Qwen, not the agent or the Knowledge Base.
  • A reference check, not a full field tool, though it no longer needs a terminal. The agent takes one reading and a named survey area. It does not match a new location by coordinates, and it does not read a full sounding curve. The full online agent (Sanity Context plus an LLM, for free-form questions) still runs only from the command line, but the rule-based offline checks now also run on a phone, no install: ves-interpretation-agent.vercel.app/ask.html.
  • A Knowledge Base rebuild is not a stable operation. Resolving one issue's standing instruction, then rebuilding, surfaced a fresh issue, three times in a row, before a rebuild finally came back with 0 pending. The entry count changed too: 8 entries as of the latest rebuild, after being 9 and then 11 at earlier points in this project. None of it was a paper error; each new issue was checked against the PDFs the same way as every one before it. But it means "the Knowledge Base is done" is only ever true as of the last rebuild you actually looked at, not a fact that holds on its own.
  • The held-out set is small: 4 questions. It is a check, not a full test.
  • The model is Qwen (qwen-plus). Another model can behave differently.

Sanity Project Details

  • Project ID: c78nb8ch
  • Dataset: production (public)
  • Knowledge Base ID: kbIp9dbX1pcY

Sources

Top comments (0)