This test put one small ServiceNow task through three AI tools: SNcode, Claude Code + ServiceNow SDK and Build Agent. All three used the same model, Claude Sonnet. The biggest differences came from how each tool prompts the model and which tools it gives it.
Each tool finished the task in the end, but the runs showed interesting differences in how each one works.
Full disclosure: this article is published by https://sncode.dev, one of the three tools, so please read it with that in mind. Every run is on video, the raw numbers are in the appendix, and token counts are cost-weighted (see Results).
The task
The task is small but realistic: two new fields on Incident, shown on the form, and a rule that blocks resolving in one case. All three tools started from the same prompt:
Add to the incident table
Fields
- "Major incident candidate" - Boolean True/False (checkbox in UI)
- "Customer impact" - Choice [none, single_user, department, enterprise] (dropdown in UI) These fields should be visible in the Incident Form ## Additional requirements
- If "Major incident candidate" is checked Incident cannot be resolved with not selected "Customer impact". Also in case non-interactive error should be custom "code_127"
Setup
All runs used a ServiceNow Personal Developer Instance (PDI) on the Australia release, and every agent used Claude Sonnet.
Tool version:
- SNcode: 0.2.26
- Claude Code + ServiceNow SDK: SDK 4.12.2 with its now-sdk skill; Claude Code 2.1.278
- Build Agent: Australia PDI build
How the test was run:
- Each agent got the prompt and nothing else until it said it was done.
- If the result was wrong, one follow-up message described the problem, not the fix.
- The stopwatch ran only while the tool was working.
- Tokens are the input, output, cache-write and cache-read counts for each round.
How tokens were measured:
- SNcode ran in its Local Claude Code mode, so its usage was read with ccusage.
- Claude Code + ServiceNow SDK: usage was read with ccusage.
- Build Agent reports token usage in the browser console logs.
All of this can be checked in the videos.
Observations
** SNcode**
- Made the changes in one round and tested them itself.
- Explained the limits of the "code_127" requirement.
Claude Code + ServiceNow SDK
- Overwrote the Incident form every time it got the original prompt. The recorded run adds one line to prevent that: "These fields should be added to the Incident Form, not ovewrite but added".
- Its business rule didn't work at first. A general follow-up, "Business rule does not work", got it fixed in round 2.
- Tested its work only in round 2, through the API.
- Its script wasn't added to ES Latest Scripts (sys_es_latest_script).
Build Agent
- Overwrote the form even with that added line, which its recorded run also uses. In round 2 it brought the form back, but with duplicated fields.
- It didn't test its own work.
Results
Here is how the three runs compare. In this run, SNcode used 5.2× fewer tokens than Claude Code + SDK and 3.3× fewer than Build Agent, in about a third of the time.
- SNcode: 1 round, 7:44, 393K cost-weighted tokens
- Claude Code + SDK: 2 rounds, 22:30, 2.05M cost-weighted tokens
- Build Agent: 2 rounds, 21:50, 1.31M cost-weighted tokens
What "cost-weighted" means: an output token counts 5×, a cache write 1.25× and a cache read 0.1× an input token. These are the price ratios for every Sonnet model on Claude's pricing page, so the totals compare what each run would cost.
What made the difference
All three tools used Claude Sonnet, so the model was the same. What differed is what each tool wraps around it: the system prompt and the tools. Each approach has good reasons behind it, so here is what the runs showed and what each approach cost on this kind of task.
1. Where the ServiceNow knowledge lives
Claude Code is a general-purpose agent, so it learns the SDK as it goes. Here it fetched 13 SDK documentation topics, over 85 KB of text, and much of it stayed in its context for later steps. Its cache reads reached 14.8M tokens over two rounds.
SNcode keeps ServiceNow know-how in its system prompt, so its agent started by looking at the instance itself. That first look returned 7.1 KB, and its cache reads stayed at 1.97M tokens.
Build Agent also looked up its Fluent documentation as it worked, in many "Explain Fluent Doc" and "Search Fluent Docs" steps. Its cache reads reached 5.8M tokens over two rounds.
2. Platform conventions by default
Adding fields to an existing form without replacing its layout is a common ServiceNow convention, and SNcode followed it without being told. Claude Code + SDK needed the extra line. Build Agent's form definition replaced the layout even with it, which suits a new form better than a change to an existing one.
3. Checking the work before reporting
SNcode tested its rule in three cases, found it wasn't firing, fixed the cause and checked the form in a browser before reporting back. The other two, Claude Code + SDK and Build Agent, checked their configuration with queries and reported done, and in this task that meant an extra round. Claude Code's first round alone used 2.8× SNcode's full run.
4. Tool design
SNcode gives its agent a small set of purpose-built actions, such as running a script on the instance and driving a browser. Their results stay small, from 141 bytes to 7.7 KB, and tools outside that set, like Bash, are blocked. Claude Code works through a shell and command-line tools, which is flexible but brings larger outputs into the context.
5. How fixes are applied
Both SDK-based tools, Claude Code and Build Agent, package the change as an app, which gives you source control and clean installs. The trade-off is that each fix means a new build and deploy: Claude Code deployed four times here. SNcode applies fixes directly on the instance, in seconds, and records them in an update set.
None of this makes the other tools a poor choice. Claude Code + SDK is great for source-controlled apps, and Build Agent lives right inside the platform. For small changes on existing instances, though, prompt and tool design made a big difference in this test.
Four ServiceNow gotchas from these runs
These are worth knowing whatever tool you use.
- SNcode run. A business rule created through the Table API can be saved with Insert and Update unchecked, and then it never runs. SNcode caught and fixed this during its own testing.
- Claude Code run. In a business rule, current.getValue() on a True/False field returned "1", not "true", so comparing it with "true" never matches. Claude Code fixed this in round 2, after a general follow-up.
- Build Agent run. Deploying a Fluent Form() for an existing form replaced all of its sections. To add fields to an out-of-box form, add them to an existing section instead. In round 2, Build Agent brought the form back, but with duplicated fields.
- When a business rule stops a Table API update, the caller sees only a generic "Operation Failed", not the custom message. A custom error code needs a Scripted REST API, as this community thread confirms. Only SNcode pointed this out.
Suggestions
The same model gave very different results in three tools. The difference was not the model but what surrounds it: the system prompt, the tools and the skills.
That holds whatever tool is used. To get better results with fewer tokens, it pays to:
- Tune the prompt. Put something like "add fields to existing forms, don't replace them", and the habit of testing changes into the system prompt, so no request has to repeat them.
- Configure the tools. Give the agent a few focused tools with compact results, and turn off the ones it shouldn't use. For instance, "toon" format instead of JSON, the same preprocessing can be done with XML
- Optimise the skills. Keep ServiceNow know-how in short, targeted skills instead of loading whole documentation sets into the context. Run an optimisation loop to get better performance of the skill
In this test, that difference meant one round instead of two and up to 5× fewer tokens. SNcode ships with this tuning for ServiceNow; with other tools, it's worth doing yourself.
Download SNcode: https://sncode.dev
These are the counts for every round, so you can redo the math.
Run: SNcode
Time: 7:44
Input: 50
Output: 26,199
Cache write: 51,801
Cache read: 1,974,467
Cost-weighted: 393,243
Run: Claude Code + SDK, round 1
Time: 11:00
Input: 104
Output: 37,538
Cache write: 158,398
Cache read: 7,169,022
Cost-weighted: 1,102,694
Run: Claude Code + SDK, round 2
Time: 11:30
Input: 74
Output: 26,729
Cache write: 40,939
Cache read: 7,642,061
Cost-weighted: 949,099
Run: Claude Code + SDK, total
Time: 22:30
Input: 178
Output: 64,267
Cache write: 199,337
Cache read: 14,811,083
Cost-weighted: 2,051,793
Run: Build Agent, round 1
Time: 10:14
Input: 124,228
Output: 25,915
Cache write: 72,011
Cache read: 2,636,179
Cost-weighted: 607,435
Run: Build Agent, round 2
Time: 11:36
Input: 88,284
Output: 28,685
Cache write: 116,996
Cache read: 3,209,808
Cost-weighted: 698,935
Run: Build Agent, total
Time: 21:50
Input: 212,512
Output: 54,600
Cache write: 189,007
Cache read: 5,845,987
Cost-weighted: 1,306,369
Cost-weighted = input + 5 × output + 1.25 × cache write + 0.1 × cache read. Totals come from the raw counts, so they can differ by one from the sum of the rows.
Top comments (0)