DEV Community

Morgan Li
Morgan Li

Posted on

Model Critique or Server Evidence: A Debate for Agent SQL Acceptance

A Thursday review queue held one agent draft that joined orders, payments, and refunds for a weekly finance extract. The author had requested a seven-day window, yet the draft projected every column from all three source tables. Two reviewers debated the draft for several minutes before either would approve it for the scheduled job. One reviewer wanted a second model critique, while the other demanded execution evidence from a disposable server.

Both positions can be defended with operational evidence, yet they answer different questions about failure and exposure. A written critique can name a missing predicate, but only a server run shows the plan, the row count, and the locks. This article compares those gates, supplies a reproducible evidence worksheet, and ends with a decision rule for promotion.

The question that should frame the debate

The useful question is not which tool sounds more modern during a crowded afternoon review meeting. The useful question is which failure you fear, and which artifact can actually observe that failure. A missing date predicate is visible in the statement text, while a bad join order appears only after planning. Lock waits, temp file growth, and underestimated row counts appear only when a server executes the statement.

If your acceptance rule collects the wrong artifact, you will feel confident while the risky behavior remains untested. That mismatch, rather than model brand or server price, is the failure mode this debate should reduce. Headlines about model quality do not change the observation problem, because a fluent explanation is still not a plan. Treat news about newer models as a reason to retest your gate, not as a substitute for the artifact.

Position A: a structured critique as the acceptance gate

Position A says a second model pass should accept or reject the draft before any database connection opens. Supporters point to speed, repeatability, and the absence of lock risk while the statement remains plain text. A critique checklist can require named checks for predicates, join keys, selected columns, and explicit limits. The output can be stored beside the draft, diffed across revisions, and reviewed without copying table rows.

The evidence for this position is real, but it is evidence about the text rather than evidence about execution. A careful critique often catches a cartesian join, an unbounded scan, or a write hidden inside a function call. It cannot tell you whether the planner chooses a nested loop, a hash join, or a sequential scan. It also cannot show whether a selective predicate matches zero rows because of a type or timezone mistake.

Position A becomes weak when the same model family drafts the statement and then reviews its own work. Shared blind spots travel from the draft into the critique, especially around business calendars and late-arriving facts. A second pass remains useful as a prefilter, yet it should not be renamed into proof that the query is safe. Keep the critique artifact, and label it as a text review rather than as an execution result.

Position B: server evidence as the acceptance gate

Position B says no agent statement is accepted until a disposable server returns a plan and a bounded result. Supporters argue that production-shaped statistics reveal mistakes that a fluent critique will describe and then miss. A rehearsal session can set a statement timeout, a lock timeout, and a read-only transaction before the draft runs. Those settings do not make the draft correct, but they make a bad draft fail where you can afford it.

The evidence for this position is the plan, the estimates, the bounded output, and the error text from the session. A sequential scan on a fixture of known size is a fact, while a model warning about scans is only a hypothesis. Server evidence also exposes search path surprises, missing indexes, and casts that prevent index use. Those facts are difficult to invent from the SQL text alone, even when the critique sounds specific and confident.

Position B becomes weak when the fixture is tiny, stale, or shaped differently from the production tables. A clean plan on a small synthetic sample can hide a hash spill that appears on a much larger month. A hosted rehearsal server also does not grant permission to upload regulated rows, secrets, or unmasked customer attributes. If the fixture cannot legally leave your network, this position must be satisfied on infrastructure you control.

Where a free model pass and a free server participate

MonkeyCode offers free model access for the critique pass and a free server option for the rehearsal database.

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

This draft states neither model names, token quotas, hardware size, nor how long either option remains available. Confirm the current terms in the product documentation before any promotion gate depends on them. The same worksheet still works if you replace that model access with another API and that server with local Postgres. Removing the product name should not remove the debate, the session limits, or the decision rule below.

Use the free options only for masked fixtures and for statement text, not for raw extracts from production. Record which option produced each artifact so a later review can see the boundary you actually tested. A hosted target is a collector of evidence, not a quieter copy of the production role. If the documented terms are narrower than this worksheet, follow the terms and shrink the fixture.

Numbered workflow: collect both artifacts, then decide

The following commands are a proposal for a local rehearsal, and they were not executed for this article. If you use a hosted free server, keep the SQL and replace only the connection target after you confirm its terms. Do not point this workflow at production, even with a read-only role, until a separate change process allows it. Start from an empty database so leftover objects from an earlier experiment cannot silently change the plan.

1. Start a disposable database

# Local throwaway password. Do not reuse it. Pin a supported image tag before relying on this.
docker run --rm --name agent-sql-rehearsal \
  -e POSTGRES_PASSWORD=rehearsal \
  -e POSTGRES_DB=rehearsal \
  -p 55432:5432 \
  postgres:16
Enter fullscreen mode Exit fullscreen mode

The image tag above is an example for a local container, not a description of any hosted free server. Wait for the port to accept connections, then load a masked fixture you are allowed to copy. Keep the fixture row counts in a small file so later steps can detect drift before you trust a plan. Stop the container when the review ends so the data does not linger on a laptop.

2. Record fixture volume before planning

SELECT 'orders' AS relname, COUNT(*) AS fixture_rows FROM orders
UNION ALL
SELECT 'payments', COUNT(*) FROM payments
UNION ALL
SELECT 'refunds', COUNT(*) FROM refunds;
Enter fullscreen mode Exit fullscreen mode

Compare those three counts with the ranges published by the team that owns the weekly finance extract. A fixture that is empty, duplicated, or missing a month will make both the plan and the critique look calmer than production. Write the counts into the evidence record before you open a model prompt or run EXPLAIN. If the drift exceeds the agreed threshold, rebuild the fixture and do not promote the draft.

3. Run the draft inside a read-only, time-bounded transaction

BEGIN READ ONLY;
SET LOCAL statement_timeout = '1500ms';
SET LOCAL lock_timeout = '400ms';
EXPLAIN (FORMAT JSON, COSTS, VERBOSE)
SELECT o.id, p.amount_cents
FROM orders AS o
JOIN payments AS p ON p.order_id = o.id
WHERE o.created_at >= DATE '2026-10-02'
  AND o.created_at < DATE '2026-10-09'
LIMIT 100;
ROLLBACK;
Enter fullscreen mode Exit fullscreen mode

BEGIN READ ONLY rejects writes in that transaction, even if a later statement tries to insert rows. The timeouts bound waiting and execution, but they do not certify that the selected rows are the rows the author wanted. Save the JSON plan, the SQLSTATE if the statement fails, and the wall time reported by psql. If you need actual row counts, run a second statement with EXPLAIN ANALYZE only after the read-only plan is understood.

4. Ask for a critique without sending row samples

Review this SQL as text only. Do not request sample rows.
Checks: date predicate present, join keys named, selected columns minimal,
no DDL, no DML, no unbounded scan language, limit present.
Return a JSON object with pass, findings, and the fear each finding addresses.
Enter fullscreen mode Exit fullscreen mode

Send the statement text and the checklist, not a CSV extract and not the JSON plan's private filter values if those values are sensitive. A critique that asks for example rows is a data-handling event, and it should fail the review even if the SQL looks sound. Store the model output next to the plan, using separate fields so neither artifact overwrites the other. If the critique and the plan disagree, keep both and let the decision rule below pick the stricter gate.

5. Score the record with an explicit function

# Proposal only: not executed against a live database or model for this article.
def promotion_allowed(record: dict) -> bool:
    writes = record["statement_class"] in {"dml", "ddl"}
    if writes and not record["server_evidence"]:
        return False
    if record["fear"] in {"plan", "locks", "row_estimate"} and not record["server_evidence"]:
        return False
    if record.get("regulated_columns") and record.get("rows_sent_to_model"):
        return False
    if record.get("fixture_drift_ratio", 0) > record.get("drift_limit", 0.2):
        return False
    if not record.get("human_reviewed"):
        return False
    return True
Enter fullscreen mode Exit fullscreen mode

The function encodes the debate instead of hiding it inside a prompt that a later editor cannot diff. A 0.2 drift limit is a placeholder threshold for the example, not a measured standard from a production incident. Change that placeholder only when the owning team publishes a numeric range for the masked rehearsal fixture. Run the function in your test suite against the cases in the next section before you trust it in review.

Evidence map and proposed cases

Feared failure Critique can observe it Server can observe it Acceptance rule
Missing date predicate Yes, from the text Yes, as an unexpectedly large scan Either artifact may block; do not accept on silence
Bad join method Only as a guess Yes, in the plan nodes Require server evidence
Hidden write Sometimes, if the syntax is visible Yes, read-only mode errors Require server evidence
Lock wait No Yes, when lock_timeout fires Require server evidence
Rows copied into a prompt No No Fail closed; neither gate excuses the copy
Fixture drift No Only if you counted rows first Discard the plan and rebuild

Use the table as a test oracle for promotion_allowed, not as a slogan to paste into a model prompt. The following cases are proposed inputs, and they illustrate the rule without claiming a measured pass rate. A read with fear plan, server evidence present, no regulated columns, drift 0.05, and human review returns true. A write without server evidence returns false, even when the critique field says the statement looks safe.

A regulated extract that sent sample rows to the model returns false, regardless of how clean the plan looks. An empty fixture with drift above the limit returns false, because the plan would describe the wrong data set. These cases are small on purpose, so a reviewer can implement them in minutes and argue about the rule rather than about a framework. Add a case from your own incident archive only when that archive can be shared inside the team.

Decision rule

  1. Name the feared failure before you choose a tool, and write that name into the evidence record.
  2. If the statement writes, changes objects, or calls a function you have not classified, require server evidence on non-production data.
  3. If the fear is a plan shape, a lock, or a row estimate, a critique may prefilter the draft but cannot accept it.
  4. If any column is regulated, do not upload those rows to a third-party server and do not paste samples into a model prompt.
  5. If fixture drift exceeds the published threshold, discard both the plan and the critique's confidence language.
  6. Promote only when the collected artifact matches the fear, the session limits match the script, and a human has read the record.

Apply the six checks in order, and stop at the first failed check instead of averaging them into a soft score. A soft score hides a regulated-data failure behind a fluent critique and a fast rehearsal plan. The rule is intentionally strict about copies and writes, because those failures are expensive relative to a delayed review. Textual doubts can still be resolved by editing the draft and collecting a new pair of artifacts.

Limitations, and who should skip this approach

This approach does not estimate production latency, and it should not be cited as a benchmark of any model or server. A masked fixture can still leak rare identifiers, so legal review remains necessary even when the SQL session is read-only. Timeouts can abort the statement before the needed plan is printed, so a timeout is not itself a safe plan. Shared rehearsal servers can also be noisy, so a slow plan there is not automatic proof of a slow plan in production.

Skip the hosted free server when the fixture cannot leave a private network, or when your policy forbids third-party compute. Skip the model critique path when the statement contains secrets, customer keys, or unpublished financial figures in literal predicates. Skip the whole worksheet if the only available credentials can also reach production, because a pasted connection string defeats the rehearsal boundary. Teams that need a change-advisory approval for every query should keep this method in a lab until that process explicitly allows it.

Closing

Before the next promotion meeting, run one unsettled draft through the six checks and keep the artifact that matches the feared failure. If those free collectors are currently available to you, use them only as described above after you confirm the live terms. That comparison is useful when it changes a decision, and it is noise when it merely adds another confident paragraph. Leave the production role untouched until the written record shows which gate actually observed the risk.

Top comments (0)