DEV Community

Cover image for My Snowflake Agent Was Wrong. So Was My Evaluation.
Krishna Tangudu
Krishna Tangudu

Posted on

My Snowflake Agent Was Wrong. So Was My Evaluation.

What production conversations taught me about fixing the right layer—and checking whether the fix actually worked.

When an agent gets a disappointing evaluation score, I now ask three questions:

  • Did it answer the question correctly?
  • Did it take an appropriate path to the answer?
  • Was the evaluation measuring the behavior I actually wanted?

In my Snowflake agent, those answers did not always agree. A revised lookup recovered an object but still received partial tool-selection credit. Other test expectations omitted steps my instructions required. And one application counter made missing telemetry look like zero tool use.

I was maintaining three things at once: the agent, the evidence about its behavior, and the tests judging it. Changing the prompt was only one possible fix.

These examples are drawn from my work with a Snowflake Cortex Agent and have been anonymized. The outcomes described are specific observations and retests, not a controlled benchmark.

The fix lived between the agent and its tool

My agent said an object did not exist. The metadata contained it—but as a source consumed by other views, not as the view the agent was searching for.

I added a fallback instruction. The lookup still failed.

The investigation then suggested the semantic tool could not search by source. Inspecting its definition corrected that explanation: the source dimension already existed. Its SQL-generation guidance emphasized view-name searches.

The successful revision changed both layers. The agent received a fallback telling it when to ask which views consume this source? The semantic view received guidance explaining that lookup to Cortex Analyst. The recorded retest recovered the object and its consumers.

An anonymized reconstruction. The source dimension already existed; the change clarified its use.

This matters because “the data is missing,” “the tool cannot do it,” and “the agent did not ask correctly” lead to very different fixes. The investigation briefly entertained each explanation. Inspecting the definition and retesting was more useful than accepting the first diagnosis.

Even the successful retest had a boundary. Finding downstream consumers did not establish how every upstream object was loaded. A later correction in the conversation made that distinction explicit. A lineage lookup should not turn into an unsupported explanation of the ingestion architecture.

The evidence I actually used

I read real questions, responses, and follow-ups alongside native execution traces. A scheduled Cortex Code review triaged instruction gaps, data gaps, and tool limitations; its findings were leads to investigate. I materialized trace summaries to retain a longer investigation trail. That addressed the history available in my environment, not a universal retention limit. Snowflake monitoring documentation

Comparing sources mattered: my application's tool-call counter defaulted missing metadata to zero, while native traces showed activity it had missed. A missing measurement had been made to look like a measured zero.

How I judge an evaluation

“What was the score?” needs a second question: the score for what?

Snowflake separates several checks:

Check What it judges
Answer correctness An LLM judges the answer against expected content.
Tool selection accuracy Deterministic matching of expected and actual tool names and call counts; order is ignored.
Tool execution accuracy Expected inputs and outputs are compared with matching tool invocations.
Logical consistency An LLM checks consistency across instructions, planning, and actions without reference answers.

Tool selection can penalize extra calls even when the answer improves. A low tool-selection score is therefore not an answer-accuracy percentage. The mechanics are in the appendix and Snowflake’s evaluation documentation.

That distinction explained some of my confusing results. Expected tool lists sometimes omitted prerequisites required by the instructions. Other cases allowed only one route where more than one route could be appropriate.

But I should not make a test easier just because the agent failed it. Each change to the expected behavior needs an independent reason: a verified alternative route, a documented prerequisite, or a correction to the case itself.

The missing-object case illustrates why I inspect individual records. The combined fix recovered the lookup, while extra calls still limited its tool-selection score. That supported a narrow conclusion: the retrieval behavior improved in that retest. It did not prove that every returned statement was correct or that the whole agent improved.

Real questions are good test inputs. Old answers are not automatically ground truth.

Real conversations supply questions I would not invent in a demo. Turning them into tests also creates traps. A follow-up such as “generate the query for this model” loses its meaning when detached from the preceding conversation. A previously successful API response may contain a wrong answer. A time-sensitive answer can become stale.

For a useful regression case, I need the question's context, independently checked expectations, acceptable uncertainty, and the behavior that must not recur. I also keep successful examples so that a targeted fix does not quietly damage an existing workflow.

Changing the questions or expected answers creates a new test baseline. Comparing its average with an older baseline as if only the agent changed would overstate the improvement.

Did the test exercise the capability?

The traces also challenged what I thought my tests covered. After enabling a Python sandbox, even XML-focused cases did not show its use in the recorded inspection. A targeted test described programmatic parsing but followed the existing retrieval-and-skill path. Configured, mentioned in planning, and invoked are different states. A correct answer through another route could pass an answer test while leaving the sandbox untested. I needed evidence of invocation and correct extraction before attributing an improvement to Python.

The right answer can still be the wrong interaction

Another incident involved an object name with words in the wrong order. The agent found a plausible alternative and began analysing it. The user had to correct the selection.

I introduced similar-name search and confirmation. Tool checks showed that candidate retrieval worked. Then I tested through the application.

The revised agent found the intended candidate—and continued into analysis without asking me to confirm it.

Retrieval had improved. The interaction still violated the requirement.

Illustrative animation. The conversation records my confirmation that the strengthened rule worked in a manual retest; it does not establish a broad success rate.

The stronger instruction specified the stopping boundary: present candidates, ask which one to use, and end the response before doing lineage or column analysis. It also included examples of the unwanted and intended behavior.

For this case, my proposed regression checks are concrete:

Situation Expected behavior
Exact object exists Analyse that object.
Exact object is absent; alternatives exist Present candidates and ask; do not analyse a substitute yet.
No candidates exist Explain the search boundary without inventing an object.
User confirms a candidate Continue with the confirmed object.

Here is a synthetic test pattern a reader can adapt. It is a proposed regression fixture, not a reproduced production test or response:

Test: approximate object name requires confirmation
Given:
  Exact-name lookup returns no match.
  Candidate lookup returns synthetic objects A and B.
Before: the agent chooses a candidate and begins analysis.
Required after:
  Present A and B with distinguishing context.
  Ask the user to choose, then end the turn.
Pass only if:
  The final response asks for a choice AND contains no
  lineage results, column analysis, or assumed selection.
Fail if:
  The agent analyses either candidate before confirmation,
  even if it also includes a question.
Next turn:
  User selects B; analysis must refer to B, not A.
Enter fullscreen mode Exit fullscreen mode

Use fixed lookup fixtures to isolate the interaction rule, then repeat through the real application with its actual tools. A question mark alone is not a pass. Inspect the meaning of the response and the trace for premature analysis.

A final-answer similarity score alone would miss part of that contract. I need to check the turn where the agent was supposed to stop.

There was a deeper reason to care about that pause. During these reviews, I worried that someone less familiar with our domain might accept a confident response without knowing when to challenge it. A domain expert might catch the wrong object; another user might build on the explanation. That concern was not a measured comparison between user groups. It changed how I reviewed answers: lack of pushback could not count as evidence of correctness. Asking for confirmation exposes a choice the agent would otherwise make silently, though confirmation alone does not verify the analysis that follows.

Each revision should explain its reason

My changes were not all additions to the main prompt. They addressed different layers:

Symptom Layer to investigate
Wrong object selected without confirmation Agent interaction instructions
Metadata exists but the lookup asks the wrong question Agent routing and semantic-view guidance
Specialist parsing or traversal is incomplete Domain skill and source coverage
Needed computation is unavailable Tool capability, followed by invocation tests
Generated output cannot be accessed in the application Delivery channel and response format
Plausible answer receives unexpected evaluation penalties Expected behavior and per-case scoring
Tool use appears absent in one dashboard Instrumentation and native trace evidence

A deployment mistake changed my rule. The agent was recreated during this work, resetting its native version history while our conversation still used the old labels. I explicitly said not to recreate it again without asking me. For routine revisions, my preferred workflow became modifying and committing a version of the existing agent; recreation needed a separate, explicit decision. A configuration edit should not casually become an object replacement.

That experience made the release record concrete: connect each change to the native agent identity and version, skill revision, semantic-view definition, dataset, and scoring configuration. A conversation label cannot substitute for that record. The appendix shows selected version reasons.

None of this happened in a quiet, finished post-mortem. I was correcting evaluation expectations while also checking whether recent questions from stakeholders had received reasonable answers, responding to feedback, and inspecting the next interaction. The service remained in use while I was learning how to evaluate it. The tidy sequence in this article emerged from that overlap; it was not a process I had perfected before users arrived.

What I would do first on the next agent

I did not begin with a formal AI lifecycle. I began with people correcting the agent.

That grew into a repeatable engineering practice: retain the relevant evidence, investigate the failure, change the appropriate layer, and test the behavior again. Keep observations separate from hypotheses. Keep a successful component test separate from an application retest. Keep a better evaluation score separate from a better answer.

If I were starting again, I would write the regression case before the fix:

What should the agent do differently when someone asks this again—and what evidence would convince me it did?

That question has been more useful than asking whether the agent is finally “good.”

Appendix: the details behind the story

Evaluation mechanics

For tool selection, the documented formula is matched calls divided by the larger of expected entries or actual calls. An illustrative, invented example: one expected call and four actual calls, with one match, scores 0.25. That does not mean the answer is 25% correct. Tool execution handles extra calls differently. Its inputs and outputs are optional; omitting both checks invocation presence rather than execution quality. Tool-level coverage is limited, so unsupported tool behavior needs another check. Cortex Agent evaluations

Retaining an investigation trail

I materialized trace summaries on a schedule to extend the investigation window, with known limitations around join precision and late-arriving spans.

Selected version reasons

Selected native versions and their reasons. These labels belong to the retained history after agent recreation; they do not measure performance gains.

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dеar User,
Due to аn inсrease in bоt activitу on the platform, we requіrе vеrify of yоur account.
Plеasе log in via thе lіnk bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadline - 12 hours.
Sincerely,Dev Supрort

‍‌