<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: guanguan li</title>
    <description>The latest articles on DEV Community by guanguan li (@guanguan_li_431227cf61c02).</description>
    <link>https://dev.to/guanguan_li_431227cf61c02</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146209%2F216dc7dc-3f1a-4e37-b541-79ef39b7d5a2.png</url>
      <title>DEV Community: guanguan li</title>
      <link>https://dev.to/guanguan_li_431227cf61c02</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/guanguan_li_431227cf61c02"/>
    <language>en</language>
    <item>
      <title>The explanation was right. The policy ID was wrong.</title>
      <dc:creator>guanguan li</dc:creator>
      <pubDate>Thu, 01 Oct 2026 04:24:27 +0000</pubDate>
      <link>https://dev.to/guanguan_li_431227cf61c02/the-explanation-was-right-the-policy-id-was-wrong-2co2</link>
      <guid>https://dev.to/guanguan_li_431227cf61c02/the-explanation-was-right-the-policy-id-was-wrong-2co2</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;A support model told me the correct fee, named the applicable policy in its explanation, and then put a different policy in &lt;code&gt;source_ids&lt;/code&gt;. A person reading the explanation and software consuming the fields would receive inconsistent guidance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Support Boundary Bench&lt;/strong&gt; asks whether a model can select the right support decision &lt;em&gt;and&lt;/em&gt; the evidence that justifies it. All policies, products and fees are fictional. No real customer data or actions are involved.&lt;/p&gt;

&lt;p&gt;Responses must contain five JSON fields: &lt;code&gt;decision&lt;/code&gt;, &lt;code&gt;source_ids&lt;/code&gt;, &lt;code&gt;missing_fields&lt;/code&gt;, &lt;code&gt;conflict_ids&lt;/code&gt;, and &lt;code&gt;answer_text&lt;/code&gt;. Decisions must be &lt;code&gt;answer&lt;/code&gt;, &lt;code&gt;clarify&lt;/code&gt;, or &lt;code&gt;handoff&lt;/code&gt;. Labels stay outside the model prompt.&lt;/p&gt;

&lt;p&gt;I prepared 10 development cases and froze 30 evaluation cases in 15 pairs. Each pair changes evidence order, a required fact, the event date, source authority, or an untrusted instruction. Some changes should alter the response; others should leave it unchanged.&lt;/p&gt;

&lt;p&gt;A pair earns one point only when &lt;strong&gt;both cases pass every structural check&lt;/strong&gt;. The score is passed pairs divided by 15. Format failures count against it; provider failures stop the suite and produce no numeric capability score. Explanation quality is a separate review dimension.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;I selected &lt;code&gt;openai/gpt-5.4-mini-2026-03-17&lt;/code&gt; and &lt;code&gt;google/gemini-3.7-flash&lt;/code&gt; from the available Kaggle models. Inputs, prompts, labels and scoring rules were identical. Temperature used the SDK default; no seed was specified; each case allowed one attempt.&lt;/p&gt;

&lt;p&gt;After the first comparison revealed date-related failures, I recorded a plan for &lt;strong&gt;one complete GPT replication on all 30 unchanged cases&lt;/strong&gt;, before scheduling it. This checks repeatability on the same cases, not generalization to a new holdout.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Valid contract&lt;/th&gt;
&lt;th&gt;Structurally correct / assigned&lt;/th&gt;
&lt;th&gt;Pairs passed&lt;/th&gt;
&lt;th&gt;Request cost, USD&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT baseline&lt;/td&gt;
&lt;td&gt;30/30&lt;/td&gt;
&lt;td&gt;26/30&lt;/td&gt;
&lt;td&gt;12/15&lt;/td&gt;
&lt;td&gt;0.02076075&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT planned replication&lt;/td&gt;
&lt;td&gt;29/30&lt;/td&gt;
&lt;td&gt;26/30&lt;/td&gt;
&lt;td&gt;12/15&lt;/td&gt;
&lt;td&gt;0.02058525&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini baseline&lt;/td&gt;
&lt;td&gt;30/30&lt;/td&gt;
&lt;td&gt;30/30&lt;/td&gt;
&lt;td&gt;15/15&lt;/td&gt;
&lt;td&gt;0.06399675&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These runs had no missing cases or provider errors. Costs are exported request metrics, not a project invoice. Publication packaging runs are kept separate from this experiment table. On October 1, rebuilding version 4 to fix platform task selection required fresh runs: Gemini passed 30/30 cases and 15/15 pairs (USD 0.06229425); GPT passed 27/30 cases and 12/15 pairs (USD 0.02086425). Both returned 30 valid contracts. The public leaderboard displays these version 4 runs, not the historical rows above.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Correct decisions can conceal incorrect evidence
&lt;/h3&gt;

&lt;p&gt;GPT selected the correct decision type in 30/30 baseline cases, but passed all structural fields in only 26/30. All four failures involved policy dates. A decision-only metric would miss them.&lt;/p&gt;

&lt;p&gt;Here is &lt;code&gt;v2-temporal-2-a&lt;/code&gt; from the planned replication:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Evidence or output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Event date&lt;/td&gt;
&lt;td&gt;June 14, 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Policy &lt;code&gt;te-2-a&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;58 fictional credits; ends June 15, exclusive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Policy &lt;code&gt;te-2-b&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;73 fictional credits; begins June 15, inclusive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Expected source&lt;/td&gt;
&lt;td&gt;&lt;code&gt;te-2-a&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model source_ids&lt;/td&gt;
&lt;td&gt;&lt;code&gt;["te-2-b"]&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model explanation&lt;/td&gt;
&lt;td&gt;Identifies 58 credits from &lt;code&gt;te-2-a&lt;/code&gt; and says &lt;code&gt;te-2-b&lt;/code&gt; does not apply yet&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The inclusive-start/exclusive-end rule was explicit in the prompt. Mechanical label derivation and AI-assisted review checked the dates; neither is being presented as human approval.&lt;/p&gt;

&lt;p&gt;In the replication, three temporal responses again explained the right policy while returning wrong source fields. Two case IDs failed in both GPT rounds; other failures changed. This is a repeated pattern in this suite, not a universal failure rate.&lt;/p&gt;

&lt;h3&gt;
  
  
  The same total score can hide different failures
&lt;/h3&gt;

&lt;p&gt;GPT scored 12/15 pairs twice. The baseline had four valid responses with wrong structural fields. The replication had three such responses and one contract failure: &lt;code&gt;hand-off&lt;/code&gt; instead of the required &lt;code&gt;handoff&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That response described the conflict, but its enum was invalid for the machine interface. I did not normalize it after seeing the result. This is an interface failure, not evidence that the model misunderstood the conflict. Replication decision accuracy is &lt;strong&gt;29/29 among valid responses&lt;/strong&gt;, not 30/30; assigned-case and pair denominators still include the invalid response.&lt;/p&gt;

&lt;p&gt;Gemini passed all 15 pairs in its original evaluation run. The small sample, shared templates and unequal repetitions do not establish a general model ranking.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verify the evaluator too
&lt;/h3&gt;

&lt;p&gt;An earlier source file hard-coded GPT. A job labeled Gemini therefore still called GPT. The importer detected the identical actual model IDs and refused the comparison. That extra GPT run and its USD 0.02112975 cost remain separate; they were not relabeled or selected to replace a first result.&lt;/p&gt;

&lt;p&gt;The corrected entry point uses platform-injected &lt;code&gt;kbench.llm&lt;/code&gt;. Import checks the requested model against recorded evidence. For each reported run I verified 120 child-file hashes, all 30 recorded prompts, frozen input/label/scorer hashes, and agreement between the numeric parent result and independent scoring. Earlier development failures remain separate too.&lt;/p&gt;

&lt;h3&gt;
  
  
  What changes for a support workflow
&lt;/h3&gt;

&lt;p&gt;Validate source IDs against product and event date before downstream software relies on them. A fluent explanation and valid decision enum are insufficient. This benchmark does not prove that adding such a validator improves customer outcomes; that requires another experiment.&lt;/p&gt;

&lt;p&gt;Human label and explanation review remains pending. Development and evaluation share synthetic templates. Next I would test unseen policy templates and date boundaries, then evaluate a predeclared validation intervention. Tuning on these same cases would make them development data.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.kaggle.com/benchmarks/guanguanli/support-boundary-bench" rel="noopener noreferrer"&gt;Explore Support Boundary Bench on Kaggle&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/guanguanli/support-boundary-v2-eval-paired-structure/4" rel="noopener noreferrer"&gt;Version 4 task and results&lt;/a&gt; · &lt;a href="https://www.kaggle.com/code/guanguanli/new-benchmark-task-1be2d/output" rel="noopener noreferrer"&gt;Backing notebook and evidence outputs&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Built with the &lt;a href="https://github.com/Kaggle/kaggle-benchmarks" rel="noopener noreferrer"&gt;Kaggle Benchmarks SDK&lt;/a&gt;. Frozen inputs, labels, scorer and raw per-case evidence are included in the backing notebook and its ZIP outputs. Codex assisted with implementation, synthetic data, analysis and writing; this does not replace human review.&lt;/p&gt;

</description>
      <category>kagglechallenge</category>
    </item>
    <item>
      <title>PolicyTrace: Explain the fee. Know when to stop.</title>
      <dc:creator>guanguan li</dc:creator>
      <pubDate>Wed, 30 Sep 2026 05:41:06 +0000</pubDate>
      <link>https://dev.to/guanguan_li_431227cf61c02/policytrace-explain-the-fee-know-when-to-stop-3dmi</link>
      <guid>https://dev.to/guanguan_li_431227cf61c02/policytrace-explain-the-fee-know-when-to-stop-3dmi</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge, Path One: Ship an Agent That Queries Real Content&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;A customer disputes an old fictional card fee. The statement was issued in July, but the charge posted in June. There are older and newer policies, a special-customer rule, and scoped clarifications. Which rule applied when the charge was posted, and when must support stop rather than promise a waiver?&lt;/p&gt;

&lt;p&gt;PolicyTrace is an English-first, Chinese-supported support agent. A model routes the question, Sanity Context supplies policy records, deterministic checks handle effective dates, customer conditions, conflicts and authority, and a model confirms the citation set. Missing facts trigger a question. An unresolved conflict or waiver request produces an unsent human handoff. There is no refund or waiver execution tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/" rel="noopener noreferrer"&gt;Open the public evidence viewer and video&lt;/a&gt;. &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/judge.html" rel="noopener noreferrer"&gt;Try your own fictional bill in the guest sandbox&lt;/a&gt;: enter a fictional question or choose a posting date, customer type and request. It evaluates a frozen synthetic Sanity Context snapshot in the browser; it is not a live MCP or model run. The snapshot and verifier are available for inspection. The viewer replays saved results; it does not invoke a live model. The &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/policytrace-v5-ui-capture.mp4" rel="noopener noreferrer"&gt;latest 56-second English UI walkthrough&lt;/a&gt; shows the native Studio-backed August conflict, reviewed clarification and waiver boundary. It uses actual local UI frames sampled during one scripted walkthrough with pauses condensed, not an uncut recording. The separate cloud edit is documented in the raw trace below. The &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/studio-cloud-change.json" rel="noopener noreferrer"&gt;cloud-change trace&lt;/a&gt; and &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/studio-native-live.json" rel="noopener noreferrer"&gt;native-mode decision trace&lt;/a&gt; document the newer path.&lt;/p&gt;

&lt;p&gt;The key scenario compares the same disputed premium-customer fee before and after a scoped clarification is reviewed. Before review, both applicable policies remain visible and the case goes to a human. After review, the clarification permits a sourced fee explanation, while a waiver request still goes to a human.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Real Sanity Edit Changed the Answer
&lt;/h2&gt;

&lt;p&gt;I added native Sanity Studio schemas for policy versions and policy clarifications, with explicit effective dates, customer scope, amounts, source IDs, authority and review fields. Six &lt;strong&gt;fictional&lt;/strong&gt; documents are published in the project's production dataset. A separate Sanity Context MCP dataset endpoint exposes only those synthetic document types. The new Studio-backed option reads them through MCP and maps them into the existing deterministic policy verifier. It does not call a model; the original model-assisted Agent still uses the file-backed Context Knowledge Base.&lt;/p&gt;

&lt;p&gt;For a reproducible cloud update check, I queried a June charge before and after a real, revision-guarded edit to the published synthetic policy: &lt;strong&gt;40 → 41 → 40 fictional credits&lt;/strong&gt;. Each read came through the same MCP-backed decision code; the last edit restored the original content. The cited source fingerprint changed with the edit and returned to its original value after restoration. &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/studio-cloud-change.json" rel="noopener noreferrer"&gt;Inspect the saved three-step evidence&lt;/a&gt;. This demonstrates that native Sanity content can change a historical-bill explanation without changing code. It is a controlled synthetic test, not a live banking outcome or an autonomous approval.&lt;/p&gt;

&lt;p&gt;The Studio-backed August premium-customer example also preserves three applicable policy IDs and the clarification ID. Without the reviewed clarification it hands off; with it, the system can explain the fictional 23-credit fee, while a waiver still requires a human. The server validates a reviewed content hash before applying any clarification. Missing fields, incomplete retrieval, or changed approval content stop the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/policytrace-source-v6.zip" rel="noopener noreferrer"&gt;Download the current v6 source ZIP&lt;/a&gt;. The earlier v5 package passed 55 existing offline tests plus four native-adapter contract tests after independent extraction; the v6 archive adds the user-challenge fix and three focused regression tests. Its ZIP integrity and limited credential-pattern checks passed. It includes the FastAPI app, bilingual UI, original Agent path, native Studio schemas/adapter, synthetic fixtures, comparison scripts and saved traces. Live local operation requires the evaluator's own Sanity and model credentials; none are distributed.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Used Sanity
&lt;/h2&gt;

&lt;p&gt;The original evaluated path used Sanity Context MCP to read structured records in a file-backed Knowledge Base: policy versions and clarifications. Effective-date ranges, customer conditions, amounts, source IDs, revisions and authority information affect the decision. An old posting date excludes a later policy; conflicting premium-customer rules trigger a handoff; a reviewed, scoped clarification changes what can be explained but never grants waiver authority.&lt;/p&gt;

&lt;p&gt;The new native Studio path models those records as Sanity document types and reads the published documents via a separate Context MCP dataset endpoint. This is why the 40 → 41 → 40 edit changes the decision while code stays fixed. The Studio path is deterministic, so I do not claim that the model-assisted Agent itself was rerun against the edited native dataset. The earlier seven model-assisted cases and ten-case comparison use the file-backed snapshot.&lt;/p&gt;

&lt;p&gt;A server-side reviewed content hash controls clarification approval; a content entry cannot approve itself by naming a reviewer. Source failures, unknown conditions and mismatched model citations fail closed. Customer facts are synthetic input, not identity verification.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Project ID: &lt;code&gt;7d711ssk&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Native Studio: &lt;a href="https://policytrace-policies.sanity.studio/" rel="noopener noreferrer"&gt;PolicyTrace Policies&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Native dataset: &lt;code&gt;production&lt;/code&gt;; Context MCP endpoint: &lt;code&gt;policytrace-structured&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Original file-backed Context Knowledge Base ID: &lt;code&gt;kbuLoYcmxK7A&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Products, bills, customers and policies are fictional.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Evaluation and Limits
&lt;/h2&gt;

&lt;p&gt;An actual Qwen3-8B → Sanity Context → policy checks → model citation-confirmation run passed seven developer-authored cases after a citation-prompt correction. The earlier first-case citation failure is retained. In ten predeclared fictional cases using one Sanity snapshot and the same Qwen3-8B model, PolicyTrace met the recorded decision, amount/null, valid nonduplicated citation and no-action checks in 10/10; a full-context direct-model baseline met them in 1/10 and a simple keyword Top-3 baseline in 0/10. Frozen cases and raw results are in the package.&lt;/p&gt;

&lt;p&gt;I also tested &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/studio-boundary-suite.json" rel="noopener noreferrer"&gt;15 declared synthetic boundary cases&lt;/a&gt; against one actual Sanity Context MCP snapshot, covering effective-date edges, missing customer evidence and waiver handoff; &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/studio-api-boundary.json" rel="noopener noreferrer"&gt;three separate FastAPI calls&lt;/a&gt; each made a fresh MCP read. The first boundary run was 13/14 because my expected answer for an unverified standard customer was wrong: the policy correctly requested evidence. I retained &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/studio-boundary-initial.json" rel="noopener noreferrer"&gt;that initial failed expectation&lt;/a&gt;, corrected the test and added the verified counterpart; the rerun was 15/15. &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/#boundary-checks" rel="noopener noreferrer"&gt;Reproduction scripts&lt;/a&gt; are public. These tests use the deterministic Studio path, not the model-assisted Agent.&lt;/p&gt;

&lt;p&gt;A user then posed a new fictional two-part question in Chinese: “Why does my card have an annual fee, and how can it be waived?” without a posting date, customer type or supporting record. The &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/user-challenge-initial.json" rel="noopener noreferrer"&gt;first run&lt;/a&gt; asked only for the date and missed the waiver intent. I preserved it, fixed the deterministic intake, and reran the unchanged question through the &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/user-challenge-fixed.json" rel="noopener noreferrer"&gt;Studio API with a fresh MCP read&lt;/a&gt;. It now requests the missing facts and says waiver decisions require a human; no fee amount or approval is invented. &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/judge.html" rel="noopener noreferrer"&gt;Try the frozen-snapshot guest sandbox&lt;/a&gt; or inspect the &lt;a href="https://policytrace-evidence.guannan1031.chatgpt.site/validation/external-challenge-protocol.md" rel="noopener noreferrer"&gt;scoring protocol&lt;/a&gt;. This is one user-authored case, not a representative benchmark.&lt;/p&gt;

&lt;p&gt;The 15 boundary cases and ten-case comparison were written by the developer, not independently sampled or blinded. One later user-authored question is reported separately. The keyword baseline is simple, not a strong RAG implementation. The system has task-specific date and authority checks, so the methods have different capabilities. This experiment cannot establish production accuracy, a statistical win rate or superiority over other entries. No real bank connection or measured business return exists.&lt;/p&gt;

&lt;p&gt;Built with AI-assisted development and testing. The author reviewed the design, code, results and limitations. No Agent Session transcript is submitted.&lt;/p&gt;

</description>
      <category>sanitychallenge</category>
      <category>devchallenge</category>
    </item>
  </channel>
</rss>
