<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Thomas Kalnik</title>
    <description>The latest articles on DEV Community by Thomas Kalnik (@skyblueballykid).</description>
    <link>https://dev.to/skyblueballykid</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F142892%2Fb157b61f-e4a1-49f4-ba20-a3dbc40d358f.jpeg</url>
      <title>DEV Community: Thomas Kalnik</title>
      <link>https://dev.to/skyblueballykid</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/skyblueballykid"/>
    <language>en</language>
    <item>
      <title>Your Jira automation rules break silently. How would you test them?</title>
      <dc:creator>Thomas Kalnik</dc:creator>
      <pubDate>Wed, 30 Sep 2026 03:32:51 +0000</pubDate>
      <link>https://dev.to/skyblueballykid/your-jira-automation-rules-break-silently-how-would-you-test-them-568g</link>
      <guid>https://dev.to/skyblueballykid/your-jira-automation-rules-break-silently-how-would-you-test-them-568g</guid>
      <description>&lt;p&gt;Every Jira admin I've read about has the same story. Someone edits an automation rule, or a field, a permission, or the workflow a rule depends on. The rule keeps "running", the audit log shows green, and it quietly stops doing its job. Nobody notices until an on-call page never goes out, or a customer ticket sits unassigned over a weekend.&lt;/p&gt;

&lt;p&gt;Atlassian's own documentation on testing a rule comes down to this: add a manual trigger and fire it by hand. There's no test mode, no staging copy of a rule, and no "tell me when this stops working".&lt;/p&gt;

&lt;h2&gt;
  
  
  What testing a rule could look like
&lt;/h2&gt;

&lt;p&gt;A rule is a promise about behaviour: &lt;em&gt;when X happens, Y should be true shortly after&lt;/em&gt;. That can be checked from the outside with the plain Jira REST API. Create an issue, wait, then look at it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# jira-checks.yml&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Highest-priority&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;bug&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pages&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on-call"&lt;/span&gt;
  &lt;span class="na"&gt;project&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;OPSTEST&lt;/span&gt;            &lt;span class="c1"&gt;# a test project the rule also applies to&lt;/span&gt;
  &lt;span class="na"&gt;create&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;issuetype&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Bug&lt;/span&gt;
    &lt;span class="na"&gt;priority&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Highest&lt;/span&gt;
    &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[check]&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on-call&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;paging"&lt;/span&gt;
  &lt;span class="na"&gt;expect_within&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;120s&lt;/span&gt;
  &lt;span class="na"&gt;expect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;paged&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;assignee&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;group&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;oncall&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
    &lt;span class="na"&gt;comment_contains&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Paged&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;via&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Opsgenie"&lt;/span&gt;
  &lt;span class="na"&gt;cleanup&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;delete&lt;/span&gt;
  &lt;span class="na"&gt;schedule&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;every&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;6h"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run that on a schedule and you find out about a broken rule within hours, not after the incident. Some details matter more than they first appear:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Use a test project that the production rules really cover.&lt;/strong&gt; If your rules are scoped to one project, a check in a sandbox project proves nothing. Many teams add a small &lt;code&gt;OPSTEST&lt;/code&gt; project to the rule's scope for exactly this.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep side effects inside the test project.&lt;/strong&gt; A check that fires a real webhook, page or customer email is worse than no check. Point notifications for the test project at a dead-letter channel.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automation queues are slow sometimes.&lt;/strong&gt; Give each expectation a generous window, and treat a &lt;em&gt;late&lt;/em&gt; result differently from a &lt;em&gt;missing&lt;/em&gt; one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Record what changed between the last pass and the first failure.&lt;/strong&gt; An API token with admin access can snapshot your rules. A diff of the rule set is the fastest way to answer "who broke it?"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alert when the check itself stops running.&lt;/strong&gt; A monitor that dies silently is the same problem again.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The question
&lt;/h2&gt;

&lt;p&gt;I'm thinking about turning this into a small product. It isn't built yet. The idea:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;you describe your 10 most important rules as checks like the one above;&lt;/li&gt;
&lt;li&gt;it runs them on a schedule, cleans up after itself, and alerts you (Slack or email) the first time one fails, with what changed in your rules since it last passed;&lt;/li&gt;
&lt;li&gt;it keeps a run history, and alerts you if checks stop running;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;$99 per Jira site per month.&lt;/strong&gt; No rule-writing or consulting, just the checks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is roughly what a failure would look like (&lt;strong&gt;a made-up example&lt;/strong&gt;, not a real customer):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;1 of 10 checks failing on acme.atlassian.net since Tue 14:02 UTC&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;✗ &lt;em&gt;Highest-priority bug pages on-call&lt;/em&gt;: fixture OPSTEST-481 created 14:02:11; after 120 s: label &lt;code&gt;paged&lt;/code&gt; missing, assignee unchanged, no comment. Last passed Tue 08:02.&lt;br&gt;
Rules changed between those runs: "P1 paging" edited Tue 11:37 (condition &lt;code&gt;priority = Highest&lt;/code&gt; → &lt;code&gt;priority = Critical&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;✓ 9 other checks passing · next run 20:02 UTC&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you run Jira automation that matters, I'd like to hear from you, in the comments or at &lt;strong&gt;&lt;a href="mailto:checks@apibreak.dev"&gt;checks@apibreak.dev&lt;/a&gt;&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Which rule broke on you last, and how did you find out?&lt;/li&gt;
&lt;li&gt;Would you pay $99/month per site for this? If not, what would it have to do, or cost?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A "no" is as useful as a "yes". If enough admins say yes, I'll build it and the first ones to reply get it free for three months.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: I run Changefeeds Tools, which builds small monitoring tools. This post was drafted with an AI assistant and approved by me.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>jira</category>
      <category>atlassian</category>
      <category>devops</category>
      <category>automation</category>
    </item>
    <item>
      <title>I diffed GitHub's, Stripe's and OpenAI's OpenAPI specs. Three vendors, three totally different change regimes.</title>
      <dc:creator>Thomas Kalnik</dc:creator>
      <pubDate>Wed, 23 Sep 2026 08:34:55 +0000</pubDate>
      <link>https://dev.to/skyblueballykid/i-diffed-githubs-stripes-and-openais-openapi-specs-three-vendors-three-totally-different-3pe1</link>
      <guid>https://dev.to/skyblueballykid/i-diffed-githubs-stripes-and-openais-openapi-specs-three-vendors-three-totally-different-3pe1</guid>
      <description>&lt;p&gt;Every vendor with a public API publishes, somewhere, an OpenAPI document describing its exact shape. Fewer people actually diff it commit to commit. I did — for three vendors a lot of us depend on, GitHub, Stripe and OpenAI — over the most recent stretch each has on record. The three specs don't just change at different speeds. They change in three genuinely different ways, and only one of those ways is safe to ignore.&lt;/p&gt;

&lt;p&gt;Method: for each vendor, two commits of their spec repository — a baseline and a current one — counted operations, and looked at what was removed or newly marked deprecated between them. The commit shas are below so you can reproduce this yourself: clone the repo, check out both commits, diff.&lt;/p&gt;

&lt;h2&gt;
  
  
  GitHub: continuous erosion, no version to hide behind
&lt;/h2&gt;

&lt;p&gt;Repository: &lt;a href="https://github.com/github/rest-api-description" rel="noopener noreferrer"&gt;&lt;code&gt;github/rest-api-description&lt;/code&gt;&lt;/a&gt;. Baseline &lt;code&gt;f7af1e5&lt;/code&gt; (12 March 2026) against current &lt;code&gt;29bcb55&lt;/code&gt; (16 September 2026) — a bit over six months.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Operations: 1093 → 1239 (154 added)&lt;/li&gt;
&lt;li&gt;8 removed&lt;/li&gt;
&lt;li&gt;7 newly marked deprecated&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The 8 removed operations fall into three groups: the Dependabot repository-access endpoints (list, patch, and set the default level), two Copilot metrics endpoints (org-level and team-level), and the issue-field-values endpoints (create, replace, delete). All eight were present at the baseline commit and absent at the current one. Two sampled commits tell you that much and no more: they don't tell you whether a removal was announced in between, which is exactly the problem if the spec is the only notice you get.&lt;/p&gt;

&lt;p&gt;The 7 newly deprecated operations are almost entirely one thing: six of them are the entire GitHub Classroom surface — &lt;code&gt;GET /classrooms&lt;/code&gt;, &lt;code&gt;GET /classrooms/{classroom_id}&lt;/code&gt;, &lt;code&gt;GET /classrooms/{classroom_id}/assignments&lt;/code&gt;, and the &lt;code&gt;/assignments/{assignment_id}&lt;/code&gt; family (itself, its accepted-assignments list, and its grades). The seventh is &lt;code&gt;GET /repos/{owner}/{repo}/dependency-graph/sbom&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;GitHub doesn't pin callers to a version. There's no dated release you're safely a few behind. The specification itself is the change log: an operation gets marked deprecated, then later it's gone, and the only signal in between is whatever you're watching yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenAI: growing fast, deprecating a whole product in the same breath
&lt;/h2&gt;

&lt;p&gt;Repository: &lt;a href="https://github.com/openai/openai-openapi" rel="noopener noreferrer"&gt;&lt;code&gt;openai/openai-openapi&lt;/code&gt;&lt;/a&gt;. Baseline &lt;code&gt;db3e531&lt;/code&gt; (16 July 2026) against current &lt;code&gt;7de0436&lt;/code&gt; (22 September 2026) — about ten weeks.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Operations: 281 → 352 (71 added) — a 25% increase in ten weeks&lt;/li&gt;
&lt;li&gt;0 removed&lt;/li&gt;
&lt;li&gt;10 newly marked deprecated&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All 10 are the entire &lt;code&gt;/videos&lt;/code&gt; surface: create, list, retrieve, retrieve content, edit, extend, remix, delete, plus the two &lt;code&gt;/videos/characters&lt;/code&gt; operations. Every operation under that path, deprecated in the same window.&lt;/p&gt;

&lt;p&gt;One thing worth being precise about: &lt;code&gt;/assistants&lt;/code&gt; (5 operations) was &lt;em&gt;already&lt;/em&gt; deprecated at the baseline commit, not newly deprecated in this window. It doesn't belong in the count above, and I'd be overstating the finding if I folded it in.&lt;/p&gt;

&lt;p&gt;Another: OpenAI's own spec marks an operation deprecated but doesn't encode a shutdown date, so "deprecated" here means "flagged," not "a date is set." And a vendor-support note — I measured OpenAI's spec with the same commit-diff approach used for the other two, but OpenAI isn't one of the two vendors the CI check described at the end of this piece actually supports yet. The numbers are real and reproducible; the tool doesn't run against this vendor today.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stripe: nothing, on purpose
&lt;/h2&gt;

&lt;p&gt;Repository: &lt;a href="https://github.com/stripe/openapi" rel="noopener noreferrer"&gt;&lt;code&gt;stripe/openapi&lt;/code&gt;&lt;/a&gt;. Baseline &lt;code&gt;cfe95bf&lt;/code&gt; (17 March 2026, API version &lt;code&gt;2026-03-25.dahlia&lt;/code&gt;) against current &lt;code&gt;30d3391&lt;/code&gt; (26 August 2026, API version &lt;code&gt;2026-08-26.dahlia&lt;/code&gt;) — about five months.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Operations: 587 → 594 (7 added)&lt;/li&gt;
&lt;li&gt;0 removed&lt;/li&gt;
&lt;li&gt;0 newly deprecated&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Five months, two dated API versions apart, nothing taken away or flagged for removal. That's not a quiet period — it's the design. Stripe pins every account to a dated API version and keeps serving that version until you explicitly upgrade. A change landing in the spec is a question about what you'd get &lt;em&gt;if you upgraded&lt;/em&gt;, not a warning about what's about to happen to you while you sit still. That's why a Stripe finding should read differently from a GitHub one: it's upgrade impact, not incident risk.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three regimes, one table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;GitHub&lt;/th&gt;
&lt;th&gt;Stripe&lt;/th&gt;
&lt;th&gt;OpenAI&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Window&lt;/td&gt;
&lt;td&gt;~6 months&lt;/td&gt;
&lt;td&gt;~5 months&lt;/td&gt;
&lt;td&gt;~10 weeks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operations&lt;/td&gt;
&lt;td&gt;1093 → 1239&lt;/td&gt;
&lt;td&gt;587 → 594&lt;/td&gt;
&lt;td&gt;281 → 352&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Removed&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Newly deprecated&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Versioning&lt;/td&gt;
&lt;td&gt;none — the spec is the notice&lt;/td&gt;
&lt;td&gt;pinned per account&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;"Watch the vendor's changelog" isn't one practice, it's three. For Stripe it means watching for a version you might one day adopt. For GitHub it means watching a spec that moves under you with no pin to hide behind. For OpenAI, on this evidence, it means watching a surface that's both growing fast and willing to deprecate an entire product category inside a quarter. Only Stripe's version of "watch it" is safe to skip while you stay on your pinned version. The other two aren't.&lt;/p&gt;

&lt;h2&gt;
  
  
  The diff isn't the hard part
&lt;/h2&gt;

&lt;p&gt;None of this needed a new diffing tool. &lt;a href="https://github.com/oasdiff/oasdiff" rel="noopener noreferrer"&gt;&lt;code&gt;oasdiff&lt;/code&gt;&lt;/a&gt; already compares two OpenAPI documents field by field, and it does it well — credit, not a pitch, there's no reason to reinvent that. The hard part, if you actually ship against these APIs, is smaller and more tedious than a diff: of GitHub's 15 removed-or-newly-deprecated operations above, the only question that matters to &lt;em&gt;you&lt;/em&gt; is whether any of them is one you call. If you don't touch Dependabot's repository-access endpoints, Copilot metrics, or GitHub Classroom, those 15 changes are noise. The other 154 additive operations in that same window aren't your problem either — nobody's build breaks because a vendor added something new.&lt;/p&gt;

&lt;p&gt;That's a filter, not a diff. A full spec diff against an active vendor like GitHub runs to hundreds of changes most weeks. The number that matters to a given integration is usually zero, occasionally one, and the gap between "zero" and "did anyone check" is the entire reason to automate this instead of skimming a changelog by eye.&lt;/p&gt;

&lt;p&gt;One more thing worth saying plainly, because a monitor that always finds something is suspicious: I looked at Twilio's &lt;code&gt;api_v2010&lt;/code&gt; spec over a comparable window and it didn't change a single operation in six months. A tool that reported findings against that spec anyway would be manufacturing noise, so it's deliberately left out of what I cover.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built to do the filtering
&lt;/h2&gt;

&lt;p&gt;That filter is &lt;a href="https://apibreak.dev" rel="noopener noreferrer"&gt;APIBreak&lt;/a&gt;. You declare the vendor endpoints your code actually calls in a small JSON manifest, pin a baseline commit or date, and it runs the same kind of comparison as above — but reports only your declared endpoints, with everything else counted and left out of the list. &lt;code&gt;npx apibreak check --manifest apibreak.json&lt;/code&gt; in CI, exit code 2 on a finding at or above &lt;code&gt;--fail-on&lt;/code&gt;, a GitHub Action wrapper, MIT-licensed, no credentials and no repository access beyond the checkout your workflow already has. It supports &lt;code&gt;github&lt;/code&gt; and &lt;code&gt;stripe&lt;/code&gt; today; OpenAI isn't wired in. It ships as the npm package &lt;a href="https://www.npmjs.com/package/apibreak" rel="noopener noreferrer"&gt;&lt;code&gt;apibreak&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;It's new. I built it, and it has no users yet — I'm not going to pretend otherwise.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written with AI assistance. The vendor numbers above are machine-generated from the commit shas cited in each section and are reproducible by diffing those commit pairs yourself.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>api</category>
      <category>openapi</category>
      <category>devops</category>
      <category>github</category>
    </item>
    <item>
      <title>Your eval set is probably in your training set — here's how to check in ten minutes</title>
      <dc:creator>Thomas Kalnik</dc:creator>
      <pubDate>Sun, 13 Sep 2026 18:41:51 +0000</pubDate>
      <link>https://dev.to/skyblueballykid/your-eval-set-is-probably-in-your-training-set-heres-how-to-check-in-ten-minutes-4k52</link>
      <guid>https://dev.to/skyblueballykid/your-eval-set-is-probably-in-your-training-set-heres-how-to-check-in-ten-minutes-4k52</guid>
      <description>&lt;p&gt;You fine-tune a model, run your benchmark, and the score jumps six points.&lt;/p&gt;

&lt;p&gt;Before you write that up, there's one question worth ten minutes: &lt;strong&gt;how many of those benchmark examples were in the training data?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the answer is "some", part of that six points is a measurement of memory rather than capability — and there is no way to separate the two after the fact.&lt;/p&gt;

&lt;p&gt;This is train/test contamination. It's one of the most common and least discussed reasons an offline number fails to reproduce in production, and it is almost never introduced deliberately.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it gets in
&lt;/h2&gt;

&lt;p&gt;Nobody copies their test set into training on purpose. It happens through ordinary steps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Merging public datasets.&lt;/strong&gt; Two datasets that look unrelated often share a source. Instruction-tuning collections are especially prone to this — many are recombinations of the same handful of seed sets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Splitting after augmentation.&lt;/strong&gt; Paraphrase or template-expand first, split second, and variants of one item land on both sides. The split looks random. It isn't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-scraping.&lt;/strong&gt; Your eval set came from a site in March. Your training crawl hit the same site in June.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Synthetic data from a model that saw the benchmark.&lt;/strong&gt; You may be distilling memorised answers straight into your training file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Datasets that grow.&lt;/strong&gt; Eval was frozen a year ago; train has been appended to weekly by three people since, and nobody re-checked.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern: contamination arrives with &lt;em&gt;pipeline changes&lt;/em&gt;. That's why a one-off audit doesn't stay true, and why this belongs in CI rather than in a notebook you ran once.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it's worse than the percentage suggests
&lt;/h2&gt;

&lt;p&gt;If 5% of your eval set is contaminated and the model scores near-perfectly on that slice, it can move the headline number by several points — often the same magnitude as the improvement you're trying to demonstrate.&lt;/p&gt;

&lt;p&gt;Worse, it biases &lt;em&gt;decisions&lt;/em&gt;, not just reporting. You pick checkpoints, hyperparameters and data mixes by comparing eval scores. Contamination rewards whichever run memorised more, which is usually the run that trained longer on the contaminated subset. A genuinely worse model can outscore a better one.&lt;/p&gt;

&lt;p&gt;And you can't correct for it afterwards by subtracting points, because you have no idea how the model would have done on those items unseen.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three levels of overlap
&lt;/h2&gt;

&lt;p&gt;Contamination detection usually gets framed as one number. It's really three questions, with three different costs and three different levels of confidence.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Exact
&lt;/h3&gt;

&lt;p&gt;Join whichever fields define an example, compare byte for byte. A hash lookup: O(n), &lt;strong&gt;complete&lt;/strong&gt; — no false positives, no false negatives, nothing to tune.&lt;/p&gt;

&lt;p&gt;Always run it. If it finds something, you have a definite problem and there's nothing to argue about.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Normalized
&lt;/h3&gt;

&lt;p&gt;Apply Unicode NFKC, lowercase, strip punctuation and symbols, collapse whitespace, then compare. Catches the same example after a reformat: smart quotes, a title-cased prompt, trailing whitespace, a markdown wrapper.&lt;/p&gt;

&lt;p&gt;Still a hash lookup, still complete, still essentially free — and in practice it finds several times more matches than exact alone, because real pipelines reformat text constantly. &lt;strong&gt;Skipping this level is the single most common reason a contamination audit reports "clean" when it isn't.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Near-duplicate
&lt;/h3&gt;

&lt;p&gt;The hard one. Two records are near-duplicates when they share most of their content but not all of it. The standard approach:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Turn each text into a set of &lt;strong&gt;shingles&lt;/strong&gt; — overlapping word n-grams, typically 5-grams. (Character n-grams for texts shorter than the window.)&lt;/li&gt;
&lt;li&gt;Define similarity as the &lt;strong&gt;Jaccard index&lt;/strong&gt; of the two shingle sets: intersection over union.&lt;/li&gt;
&lt;li&gt;All-pairs comparison is O(n²), so approximate it with &lt;strong&gt;MinHash&lt;/strong&gt;: a fixed number of hash permutations (128 is common) reduce each set to a short signature whose agreement rate estimates Jaccard.&lt;/li&gt;
&lt;li&gt;Group signatures into &lt;strong&gt;LSH bands&lt;/strong&gt;. Two records become &lt;em&gt;candidates&lt;/em&gt; if any band matches exactly.&lt;/li&gt;
&lt;li&gt;Score every candidate on the &lt;strong&gt;shingle sets themselves&lt;/strong&gt;, not on the signatures, and keep the pairs above your threshold. (Very long records are first reduced to a bottom-k sketch — the k smallest hashes — a uniform sample that stays stable when the record is edited.)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The detail worth internalising: &lt;strong&gt;the approximation lives in step 4, candidate generation — not in the signatures used for scoring.&lt;/strong&gt; Near-duplicate detection can miss a small number of borderline pairs, but every similarity it reports is measured from the shingle sets rather than read off a MinHash signature — exact for ordinary-length records, and an unbiased bottom-k estimate for very long ones. A tool that reports signature agreement as its similarity is giving you a noisier answer than it needs to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Picking a threshold
&lt;/h2&gt;

&lt;p&gt;The Jaccard threshold is the only judgement call in the whole process. Rough orientation for word 5-grams:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Threshold&lt;/th&gt;
&lt;th&gt;Roughly catches&lt;/th&gt;
&lt;th&gt;Expect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0.95+&lt;/td&gt;
&lt;td&gt;Whitespace, single-token diffs&lt;/td&gt;
&lt;td&gt;Almost no false positives; misses most real duplication&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.85&lt;/td&gt;
&lt;td&gt;A changed sentence or number&lt;/td&gt;
&lt;td&gt;Good conservative default for short records&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.80&lt;/td&gt;
&lt;td&gt;Reworded intro plus small edits&lt;/td&gt;
&lt;td&gt;The usual default&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;td&gt;Same content, substantially rewritten&lt;/td&gt;
&lt;td&gt;Real recall gain, real review burden&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&amp;lt; 0.6&lt;/td&gt;
&lt;td&gt;Shared templates and boilerplate&lt;/td&gt;
&lt;td&gt;Mostly false positives on formatted data&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two adjustments that matter more than the table:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Short records → higher threshold.&lt;/strong&gt; Few shingles means each differing word costs a lot of similarity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Heavy boilerplate → strip it first.&lt;/strong&gt; An identical system prompt on every row means you'll measure the template, not the content.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The reliable method is empirical: run at 0.8, read twenty borderline pairs, and move the threshold based on whether &lt;em&gt;you&lt;/em&gt; would call them the same example. Five minutes, and it beats any rule of thumb — including this one.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to report it
&lt;/h2&gt;

&lt;p&gt;Report contamination rate as the fraction of &lt;strong&gt;eval&lt;/strong&gt; records with at least one match in train. Records, not pairs: a training file that duplicates one eval example forty times is one contaminated eval record.&lt;/p&gt;

&lt;p&gt;And report per level:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;exact:      31 eval records
normalized: 12 eval records
near:       27 eval records   (threshold 0.8)
-----------------------------------------
70 of 1,000 eval records (7.00%)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first two lines are certainties. The third depends on a threshold you chose and a reviewer is entitled to argue with it. Keeping them separate is what makes the number defensible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checking your own files
&lt;/h2&gt;

&lt;p&gt;I built &lt;a href="https://splitcheck-two.vercel.app" rel="noopener noreferrer"&gt;SplitCheck&lt;/a&gt; because I was tired of writing this script slightly differently every time. Drop a train file and an eval file on the page and it runs all three levels in a Web Worker — nothing is uploaded, there's no upload endpoint in the app, and the worker bundle contains no network APIs at all. You get the per-level breakdown, the matched pairs with both snippets, and your training file with the overlap removed.&lt;/p&gt;

&lt;p&gt;It's free and there's no sign-in. There's also a &lt;code&gt;$29&lt;/code&gt; CLI for data that can't go in a browser and for CI, but the browser version isn't crippled to sell it — same three levels, same code.&lt;/p&gt;

&lt;p&gt;If you'd rather build it yourself with &lt;code&gt;datasketch&lt;/code&gt; or &lt;code&gt;text-dedup&lt;/code&gt;, do that. It's an afternoon, the libraries are good, and you'll understand your own pipeline better for it. What you're buying, if you buy anything, is the afternoon.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do when you find overlap
&lt;/h2&gt;

&lt;p&gt;The instinct is to delete from the eval set. &lt;strong&gt;Resist it.&lt;/strong&gt; Removing eval examples changes what your benchmark measures and breaks comparability with every number you've already published.&lt;/p&gt;

&lt;p&gt;Clean the &lt;em&gt;training&lt;/em&gt; set instead: remove every train record that matched, retrain, re-evaluate. Your eval set stays fixed, historical comparisons stay valid, and the new number is honest.&lt;/p&gt;

&lt;p&gt;Then:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Keep the report&lt;/strong&gt; next to the model artifacts. In six months, when someone asks whether the eval was clean, you want a file rather than a memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compare the contaminated and clean scores.&lt;/strong&gt; That delta is the most useful diagnostic you'll get all week.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review the borderline pairs before deleting.&lt;/strong&gt; At 0.8 there will be a few genuinely distinct examples that share phrasing.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What no lexical method will catch
&lt;/h2&gt;

&lt;p&gt;Being clear about the ceiling matters more than the pitch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Semantic paraphrase&lt;/strong&gt; sharing no 5-gram. That needs embedding similarity — slower, less interpretable, non-deterministic across model versions, and it will also flag pairs that merely share a topic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contamination through a third model.&lt;/strong&gt; If your training data came from a model that memorised the benchmark, the leak is in weights, not in matching text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fields you didn't compare.&lt;/strong&gt; Overlap in a column you excluded is invisible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A base model's pretraining corpus.&lt;/strong&gt; Nothing here can see that. The honest response is to say so when you report the score.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that makes the check less worth running. Exact and normalized overlap is common, entirely detectable and completely fixable — and it takes ten minutes to find out.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: this post was drafted with AI assistance and reviewed before publishing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>datascience</category>
      <category>python</category>
      <category>datasets</category>
    </item>
  </channel>
</rss>
