<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: The ABTestly team</title>
    <description>The latest articles on DEV Community by The ABTestly team (@abtestlyteam).</description>
    <link>https://dev.to/abtestlyteam</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4144111%2Facea35da-3814-46e8-9aae-d127ce196742.png</url>
      <title>DEV Community: The ABTestly team</title>
      <link>https://dev.to/abtestlyteam</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/abtestlyteam"/>
    <language>en</language>
    <item>
      <title>Almost every A/B testing vendor refuses to publish a price. Here are ours</title>
      <dc:creator>The ABTestly team</dc:creator>
      <pubDate>Sat, 26 Sep 2026 12:06:07 +0000</pubDate>
      <link>https://dev.to/abtestly/almost-every-ab-testing-vendor-refuses-to-publish-a-price-here-are-ours-156h</link>
      <guid>https://dev.to/abtestly/almost-every-ab-testing-vendor-refuses-to-publish-a-price-here-are-ours-156h</guid>
      <description>&lt;p&gt;$99, $249, $549 a month.&lt;/p&gt;

&lt;p&gt;Those are on our pricing page, and they are the least useful part of publishing a price.&lt;/p&gt;

&lt;p&gt;The useful part is what publishing forces you to do, and what refusing to publish lets a vendor get away with.&lt;/p&gt;

&lt;h2&gt;
  
  
  Go and check
&lt;/h2&gt;

&lt;p&gt;Open the pricing page of the major A/B testing platforms. Optimizely, VWO, AB Tasty.&lt;/p&gt;

&lt;p&gt;You will find tiers with names, feature grids, and a button that says Contact Sales. What you will not find is a number at any traffic volume.&lt;/p&gt;

&lt;p&gt;Convert publishes. Their Growth plan is $399 a month at 100,000 tested visitors, and you can read it without talking to anybody.&lt;/p&gt;

&lt;p&gt;That is close to the whole market, and it is one of the odder norms in B2B software. Nobody hides the price of a database.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a hidden price actually costs you
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;You cannot compare without entering a funnel.&lt;/strong&gt; Evaluating four vendors means four discovery calls before you know whether any of them are in your range. That is not a scheduling annoyance, it is a deliberate cost imposed on comparison.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The price is a function of you, not the product.&lt;/strong&gt; When the number arrives at the end of a call, it has been shaped by your company size, your stack, and how urgent you sounded. Two teams buying the same thing pay differently, and neither knows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Renewal has no anchor.&lt;/strong&gt; If the original number was negotiated rather than published, there is nothing to measure the increase against.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It selects for a particular kind of buyer.&lt;/strong&gt; A team that can run a three-month procurement process is a different customer from a team that wants to try something on Thursday. Hidden pricing is how a vendor quietly stops selling to the second kind.&lt;/p&gt;

&lt;h2&gt;
  
  
  What publishing forces
&lt;/h2&gt;

&lt;p&gt;It forces you to pick a metering unit and defend it.&lt;/p&gt;

&lt;p&gt;Ours is monthly tested users, and the definition has edges we have to state rather than resolve in our favour. Subdomains of one registrable domain share a visitor id and count as one. Two genuinely different domains count as two, because browsers isolate storage per domain and there is no cross-domain identifier to join them.&lt;/p&gt;

&lt;p&gt;That second line costs us money on any customer running a &lt;code&gt;.com&lt;/code&gt; and a &lt;code&gt;.co.uk&lt;/code&gt;. It is still what happens, so it is what the docs say.&lt;/p&gt;

&lt;p&gt;It also forces you to stop selling features you do not ship at that tier. Frequentist statistics are on every plan including the $99 one. Sequential and Bayesian start at Pro, $249. I cannot write "three statistics engines" next to the $99 price, because that would be false, and the price being public is what makes it checkable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Our terms, plainly
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;th&gt;Monthly tested users&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Starter&lt;/td&gt;
&lt;td&gt;$99&lt;/td&gt;
&lt;td&gt;50,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro&lt;/td&gt;
&lt;td&gt;$249&lt;/td&gt;
&lt;td&gt;200,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Business&lt;/td&gt;
&lt;td&gt;$549&lt;/td&gt;
&lt;td&gt;500,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise&lt;/td&gt;
&lt;td&gt;Custom&lt;/td&gt;
&lt;td&gt;Custom&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Annual billing saves 20%.&lt;/p&gt;

&lt;p&gt;Every account starts with a 14-day trial on a monthly plan. &lt;strong&gt;A card is required&lt;/strong&gt;, and nothing is charged until day 15.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;There is no free tier.&lt;/strong&gt; There used to be, and it closed to new signups in September 2026. I would rather write that sentence than let someone discover it at signup.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that is not virtue
&lt;/h2&gt;

&lt;p&gt;We can publish because we are small and we are not defending a renewal number against a board.&lt;/p&gt;

&lt;p&gt;A vendor with a large enterprise base has a real reason to keep prices off the page: their existing contracts are all different, and publishing a list price makes every one of those negotiations legible at once. That is a genuine constraint, not just cowardice.&lt;/p&gt;

&lt;p&gt;So this is a structural advantage rather than a moral one, and it will get harder to keep as we grow. Worth saying now, while it is true.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do with this
&lt;/h2&gt;

&lt;p&gt;If you are evaluating tools, ask every vendor for a number at your actual traffic volume, in writing, before the demo.&lt;/p&gt;

&lt;p&gt;The ones who give it to you have told you something real about how they will behave at renewal. The ones who will not have also told you something.&lt;/p&gt;

&lt;p&gt;And check what the metering unit does with your specific setup, because that is where a published price quietly becomes a different price.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;ABTestly is A/B testing for teams that write their variations in code rather than in a visual editor. The metering rules are documented at &lt;a href="https://docs.abtestly.com" rel="noopener noreferrer"&gt;docs.abtestly.com&lt;/a&gt;, and the prices are at &lt;a href="https://abtestly.com/pricing" rel="noopener noreferrer"&gt;abtestly.com/pricing&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>saas</category>
      <category>webdev</category>
      <category>testing</category>
      <category>startup</category>
    </item>
    <item>
      <title>A rollback plan needs a detection plan</title>
      <dc:creator>The ABTestly team</dc:creator>
      <pubDate>Sat, 26 Sep 2026 11:59:19 +0000</pubDate>
      <link>https://dev.to/abtestly/a-rollback-plan-needs-a-detection-plan-5emc</link>
      <guid>https://dev.to/abtestly/a-rollback-plan-needs-a-detection-plan-5emc</guid>
      <description>&lt;p&gt;Every team running experiments has a rollback plan. Pause the test, ship the fix, move on.&lt;/p&gt;

&lt;p&gt;Far fewer have a detection plan, which is the part that decides how long the broken thing was live before anyone hit the button.&lt;/p&gt;

&lt;p&gt;A rollback you trigger on day nine is not a rollback. It is a post-mortem.&lt;/p&gt;

&lt;h2&gt;
  
  
  The release cadence changed under everyone
&lt;/h2&gt;

&lt;p&gt;Chrome now ships a new Stable version every two weeks.&lt;/p&gt;

&lt;p&gt;That is a two-week window between "this variation works" and "this variation runs on a browser that did not exist when it was written". Multiply by the other engines and their own cadences.&lt;/p&gt;

&lt;p&gt;Most A/B test variations are DOM manipulation against selectors that were correct on the day they were authored. Nothing in that description survives contact with a fast release train automatically.&lt;/p&gt;

&lt;p&gt;So the QA question stopped being "did it work when we built it" and became "how would we know if it stopped working".&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually breaks, and when
&lt;/h2&gt;

&lt;p&gt;The first version of a variation can look completely fine and still fail before the statistics ever run.&lt;/p&gt;

&lt;p&gt;The failure modes are boringly consistent. A selector matches on first load and not after a client-side route change. A variation applies on desktop and silently no-ops on the mobile template. A goal fires on an element the variation replaced, so the exposure is recorded and the conversion never is.&lt;/p&gt;

&lt;p&gt;None of those throw an error. None of them show up as a red thing on a dashboard. The test just quietly collects data that means nothing.&lt;/p&gt;

&lt;p&gt;That is the detection gap. The variation is not down, it is wrong, and wrong looks identical to working until you check.&lt;/p&gt;

&lt;h2&gt;
  
  
  Editing a running test is the sharp edge
&lt;/h2&gt;

&lt;p&gt;Here is the part most tools stay quiet about: a running test can be edited.&lt;/p&gt;

&lt;p&gt;You found a bug in your variant code, you fix it, and now half your sample saw one thing and half saw another, reported as one number.&lt;/p&gt;

&lt;p&gt;There are only two honest ways to handle that, and both of ours are confirmation-gated:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pause, edit, resume.&lt;/strong&gt; The clock resets on the results page and you get a new epoch. The old data is still there, it is just separated from what comes after.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reset data.&lt;/strong&gt; Discards the run so far and starts fresh with the updated code.&lt;/p&gt;

&lt;p&gt;Reset is the heavier one, and it does more than clear counts. It rotates the experiment's bucketing salt, so visitors are assigned from scratch and someone who saw one variation may now see another.&lt;/p&gt;

&lt;p&gt;That is the honest cost of a reset, and it is why it is not a shrug.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two design decisions that follow from this
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Material settings lock at 1,000 visitors.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Once a running experiment has 1,000 visitors since its last reset, six material settings become read-only in the editor:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Locked after 1,000 visitors. Pause the experiment to edit.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Pausing opens them again. The lock is not there to stop you, it is there to make the change deliberate, because changing traffic allocation mid-flight on a live test is how you manufacture a sample ratio mismatch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The change notice does not go away.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If material settings changed during a run, the results page says so, and it keeps saying so until you reset the data.&lt;/p&gt;

&lt;p&gt;That is deliberate. A notice that disappears once you have seen it is a notice that stops doing its job at the exact moment it matters, which is when a colleague reads the result three weeks later with no memory of what happened on day two.&lt;/p&gt;

&lt;p&gt;Clearing it requires a reset, which advances the baseline so earlier changes fall outside the window. Nothing is deleted. The change events stay in the table, they just stop being in scope.&lt;/p&gt;

&lt;p&gt;You cannot keep the numbers and lose the notice about them. That is the whole point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building an actual detection plan
&lt;/h2&gt;

&lt;p&gt;Preview every variant on the live site before launching. Not in a staging clone, on the real page, because the thing that breaks is usually the interaction with production markup you did not write.&lt;/p&gt;

&lt;p&gt;Watch the exposure count per arm in the first hour, not the conversion rate. Conversions are too sparse to tell you anything early. Exposures tell you immediately whether the variation is reaching people at the rate you expected.&lt;/p&gt;

&lt;p&gt;Run the sample ratio check continuously, not at the end. An SRM that appears on day one is a bug you can still fix cheaply.&lt;/p&gt;

&lt;p&gt;Decide in advance what number triggers a rollback, and write it down next to the hypothesis. "We will pause if exposures in variant B fall below 40% of control" is a detection plan. "We will keep an eye on it" is not.&lt;/p&gt;

&lt;p&gt;And assume something will change underneath you in the next fortnight, because on the current release cadence, something will.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;ABTestly is A/B testing for teams that write their variations in code rather than in a visual editor. The reset and locking behaviour is documented at &lt;a href="https://docs.abtestly.com" rel="noopener noreferrer"&gt;docs.abtestly.com&lt;/a&gt;, and the prices are published at &lt;a href="https://abtestly.com/pricing" rel="noopener noreferrer"&gt;abtestly.com/pricing&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>testing</category>
      <category>devops</category>
      <category>javascript</category>
    </item>
    <item>
      <title>Your A/B testing tool and your analytics will never agree, and that is fine</title>
      <dc:creator>The ABTestly team</dc:creator>
      <pubDate>Sat, 26 Sep 2026 11:57:42 +0000</pubDate>
      <link>https://dev.to/abtestly/your-ab-testing-tool-and-your-analytics-will-never-agree-and-that-is-fine-3a9l</link>
      <guid>https://dev.to/abtestly/your-ab-testing-tool-and-your-analytics-will-never-agree-and-that-is-fine-3a9l</guid>
      <description>&lt;p&gt;"GA4 doesn't match, so I don't trust the test."&lt;/p&gt;

&lt;p&gt;I hear this constantly, and the instinct behind it is right. Two systems reporting the same experiment should not disagree.&lt;/p&gt;

&lt;p&gt;But they are not reporting the same thing, and once you see why, the disagreement becomes useful rather than alarming.&lt;/p&gt;

&lt;h2&gt;
  
  
  They are counting different populations
&lt;/h2&gt;

&lt;p&gt;An A/B testing tool counts people it bucketed into an experiment. Analytics counts people who loaded a page and fired a tag.&lt;/p&gt;

&lt;p&gt;Those two sets overlap. They are not equal, and nothing you configure will make them equal.&lt;/p&gt;

&lt;p&gt;A visitor with an ad blocker that eats your analytics but not your testing snippet is in one and not the other. A visitor who leaves before your tag manager finishes is in neither, or in one. A bot filter that treats the two scripts differently splits them further.&lt;/p&gt;

&lt;p&gt;So the question is never "why do the totals differ". It is "does the difference move with the variant", which is a completely different and much more answerable question.&lt;/p&gt;

&lt;h2&gt;
  
  
  The identity problem underneath it
&lt;/h2&gt;

&lt;p&gt;Here is the mechanism that decides what "the same visitor" even means, and almost nobody checks it before reading a result.&lt;/p&gt;

&lt;p&gt;The visitor id lives in one domain's localStorage and cookie jar. That single fact produces this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Where the same person browses&lt;/th&gt;
&lt;th&gt;Counted as&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;example.com&lt;/code&gt; and &lt;code&gt;shop.example.com&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;1 visitor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Two experiments on the same site&lt;/td&gt;
&lt;td&gt;1 visitor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;site-A.com&lt;/code&gt; and &lt;code&gt;site-B.com&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;2 visitors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The same domain in two different organisations&lt;/td&gt;
&lt;td&gt;2 visitors&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Subdomains of one registrable domain share a visitor id. Two genuinely different domains cannot, because browsers give each domain its own isolated storage.&lt;/p&gt;

&lt;p&gt;The same human arrives as two unrelated visitors, and is metered twice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;There is no cross-domain identifier, and we are not planning to add one.&lt;/strong&gt; That is a deliberate position, not a gap in the roadmap. Stitching identity across domains means shipping something that follows people between sites, and we would rather explain the limitation than build that.&lt;/p&gt;

&lt;p&gt;If your brand runs &lt;code&gt;example.com&lt;/code&gt; and &lt;code&gt;example.co.uk&lt;/code&gt; as separate registrable domains, your testing tool sees two audiences. So does your analytics, differently. That is most of the gap right there.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to actually reconcile them
&lt;/h2&gt;

&lt;p&gt;Stop comparing totals. Join on the event.&lt;/p&gt;

&lt;p&gt;GA4 emits an &lt;code&gt;experience_impression&lt;/code&gt; event carrying &lt;code&gt;exp_variant_string&lt;/code&gt;. Pipe GA4 to BigQuery and join that event to your conversion events on &lt;code&gt;user_pseudo_id&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That is the same join GA4 does internally, except in raw SQL, with no Custom Dimension cap, no Audience limits, and no &lt;code&gt;(other)&lt;/code&gt; rollup swallowing your variant names.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;(other)&lt;/code&gt; bucket is the one that catches people out. GA4 has cardinality limits, and when you exceed them it quietly rolls values up. Your variant B stops being variant B and becomes part of a bucket, and the report still renders, and nobody notices.&lt;/p&gt;

&lt;p&gt;For concurrent experiments this stops being optional. Custom Dimensions run out fast, and the in-product GA4 reporting cannot express "bucketed into experiment 1 variant B &lt;em&gt;and&lt;/em&gt; experiment 2 control".&lt;/p&gt;

&lt;h2&gt;
  
  
  What we keep, and for how long
&lt;/h2&gt;

&lt;p&gt;The ledger is the canonical store for an experiment. Every exposure and every goal that was accepted, with its payload.&lt;/p&gt;

&lt;p&gt;Behind it sits a raw compressed event archive, and &lt;strong&gt;that archive is purged nightly once it passes 24 months.&lt;/strong&gt; It is a data-retention policy, not a billing lever, and it is unaffected by your plan or your subscription state.&lt;/p&gt;

&lt;p&gt;The ledger the results are computed from survives the purge and stays on the results page for the life of the experiment. Only the raw events age out.&lt;/p&gt;

&lt;p&gt;If you need raw events beyond 24 months, run the async export inside the window and keep the CSV in your own storage. We would rather tell you that now than when you go looking for month 25.&lt;/p&gt;

&lt;h2&gt;
  
  
  The practical rule
&lt;/h2&gt;

&lt;p&gt;Pick one system as the system of record for the decision, before the test starts. For an experiment, that should be the testing tool, because it is the only one that knows who was bucketed.&lt;/p&gt;

&lt;p&gt;Use analytics for the diagnosis, not the verdict. It is excellent at telling you &lt;em&gt;where&lt;/em&gt; a difference came from and poor at telling you whether the difference is real.&lt;/p&gt;

&lt;p&gt;And when the numbers disagree, check the direction of the gap rather than its size. A gap that is the same shape in both arms is measurement. A gap that differs by variant is a bug, and it is the most important thing on your screen.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;ABTestly is A/B testing for teams that write their variations in code rather than in a visual editor. The GA4 and BigQuery join is documented at &lt;a href="https://docs.abtestly.com" rel="noopener noreferrer"&gt;docs.abtestly.com&lt;/a&gt;, and the prices are published at &lt;a href="https://abtestly.com/pricing" rel="noopener noreferrer"&gt;abtestly.com/pricing&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>analytics</category>
      <category>testing</category>
      <category>javascript</category>
    </item>
    <item>
      <title>How long should an A/B test run? Not a number of weeks</title>
      <dc:creator>The ABTestly team</dc:creator>
      <pubDate>Sat, 26 Sep 2026 11:50:44 +0000</pubDate>
      <link>https://dev.to/abtestly/how-long-should-an-ab-test-run-not-a-number-of-weeks-51ce</link>
      <guid>https://dev.to/abtestly/how-long-should-an-ab-test-run-not-a-number-of-weeks-51ce</guid>
      <description>&lt;p&gt;Two quantities decide how long a test runs, and the calendar is not one of them.&lt;/p&gt;

&lt;p&gt;How many visitors per variation the test needs, and how fast you supply them. That is the whole projection.&lt;/p&gt;

&lt;p&gt;Our duration estimator returns 15 days for a particular test. The schedule you should commit to is 21.&lt;/p&gt;

&lt;p&gt;The gap between those two numbers is the interesting part, and it is where most test plans go wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the projection actually computes
&lt;/h2&gt;

&lt;p&gt;Required sample per arm, divided by arrival rate. That is it.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;required&lt;/code&gt; is always computed at 95% confidence and 80% power. Not at whatever you set in the experiment. At 95 and 80, every time.&lt;/p&gt;

&lt;p&gt;So if you configured the test at 99% confidence, or you are running four variations under a Bonferroni or Sidak correction, you need more than the projection says.&lt;/p&gt;

&lt;p&gt;Treat the number as a floor. It is the arithmetic, not the plan.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four things it does not know
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;It does not round to whole weeks.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Behaviour varies across the week. A run that stops mid-week hands one variation an extra Saturday.&lt;/p&gt;

&lt;p&gt;Round your own end date up to a whole number of weeks so every day of the cycle appears the same number of times in both arms. That is why 15 becomes 21.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It does not hold a business-cycle floor.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If traffic is heavy enough that the sample arrives in two days, the projection will happily say two days.&lt;/p&gt;

&lt;p&gt;Two days is one narrow slice of your audience and one mood of the market. Hold a floor of at least one full cycle, preferably two, so you can watch the effect survive a second week.&lt;/p&gt;

&lt;p&gt;If your purchase cycle is longer than a week, stretch the floor to match it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It does not model novelty or primacy.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Returning visitors react to a change because it is new. Some engage more than they will once it is ordinary. Some are thrown by an unfamiliar layout.&lt;/p&gt;

&lt;p&gt;Both effects fade. Neither is visible to a projection that only counts visitors.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It does not know your confidence level or your correction.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Covered above, and it is the one people are most surprised by.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trap at the other end
&lt;/h2&gt;

&lt;p&gt;There is a status on the runway line called &lt;code&gt;alreadySignificant&lt;/code&gt;, and it short-circuits the projection to &lt;code&gt;reached&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The moment the leading variation clears 95% on the gated p-value, the line stops showing a date and starts saying you have enough.&lt;/p&gt;

&lt;p&gt;On day 4 of a planned three-week run, that is exactly the reading a fixed-horizon test cannot support.&lt;/p&gt;

&lt;p&gt;So we wrote the caveat into the product rather than leaving it in a blog post. The status describes the state of the evidence right now. It is not permission to stop.&lt;/p&gt;

&lt;p&gt;If you want a runway you are allowed to watch continuously, that is what a sequential test is for. A fixed-horizon test plus a live dashboard is not the same thing, however much it looks like it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why we gate small numbers instead of hiding them
&lt;/h2&gt;

&lt;p&gt;There is a floor of 100 visitors and 5 conversions before certain figures are judged at all.&lt;/p&gt;

&lt;p&gt;But the rate, the interval and the lift are never withheld, at any sample size.&lt;/p&gt;

&lt;p&gt;Those three are honest at any n, because the interval is the thing that says so. A 95% interval of &lt;code&gt;[-19%, +44%]&lt;/code&gt; is not a weak result being hidden from you. It is the result.&lt;/p&gt;

&lt;p&gt;What we refuse to do is print a bare percentage off a handful of visitors, because a bare percentage is the thing people screenshot.&lt;/p&gt;

&lt;p&gt;The revenue-per-visitor interval uses the same bar, since a revenue interval is at its noisiest exactly where that bar is not met.&lt;/p&gt;

&lt;h2&gt;
  
  
  One thing a confidence interval is not
&lt;/h2&gt;

&lt;p&gt;It is not the probability that the true value lies inside it.&lt;/p&gt;

&lt;p&gt;Under the frequentist reading, the true value is fixed and the interval is what is random. A 95% interval is a procedure that brackets the truth in 95% of repetitions.&lt;/p&gt;

&lt;p&gt;The sentence people actually want sounds like "there is an 87% probability this variation is better". That is a posterior probability with a credible interval beside it, and it comes from a Bayesian engine, not from a frequentist CI with the words swapped.&lt;/p&gt;

&lt;p&gt;It is also not a guard against bias. An interval describes sampling noise and nothing else. A test with a sample ratio mismatch can have a beautifully tight interval around a meaningless number.&lt;/p&gt;

&lt;h2&gt;
  
  
  The practical version
&lt;/h2&gt;

&lt;p&gt;Before you launch, compute the required sample at the confidence and power you are actually using, including any correction.&lt;/p&gt;

&lt;p&gt;Divide by your real arrival rate, not your best week.&lt;/p&gt;

&lt;p&gt;Round up to whole weeks. Apply a floor of one full business cycle, two if you can afford it.&lt;/p&gt;

&lt;p&gt;Write the end date down before the test starts, and treat every look before it as monitoring rather than deciding.&lt;/p&gt;

&lt;p&gt;Then, when someone asks on day 4 whether it is winning, the answer is that the test is still running, and that is a complete answer.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;ABTestly is A/B testing for teams that write their variations in code rather than in a visual editor. The duration model and its caveats are documented at &lt;a href="https://docs.abtestly.com" rel="noopener noreferrer"&gt;docs.abtestly.com&lt;/a&gt;, and the prices are published at &lt;a href="https://abtestly.com/pricing" rel="noopener noreferrer"&gt;abtestly.com/pricing&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>testing</category>
      <category>statistics</category>
      <category>programming</category>
    </item>
    <item>
      <title>2.2 ms, 19.9 ms, 529 ms: benchmarking an A/B testing runtime and publishing the caveats</title>
      <dc:creator>The ABTestly team</dc:creator>
      <pubDate>Sat, 26 Sep 2026 11:24:22 +0000</pubDate>
      <link>https://dev.to/abtestly/22-ms-199-ms-529-ms-benchmarking-an-ab-testing-runtime-and-publishing-the-caveats-6g3</link>
      <guid>https://dev.to/abtestly/22-ms-199-ms-529-ms-benchmarking-an-ab-testing-runtime-and-publishing-the-caveats-6g3</guid>
      <description>&lt;p&gt;Most A/B testing vendors quote one performance number, and it is the wrong one.&lt;/p&gt;

&lt;p&gt;The number is snippet size. Ours is about 31 KB gzipped. That figure tells a visitor nothing at all about what they saw, or when.&lt;/p&gt;

&lt;p&gt;So here is the measurement we actually take, the method behind it, and the reasons it is a floor rather than a result.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;p&gt;p75 time from navigation start to a variation landing in the DOM, over 200 page loads in each condition:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;p75&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Single page app route change&lt;/td&gt;
&lt;td&gt;2.2 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repeat view, warm cache&lt;/td&gt;
&lt;td&gt;19.9 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;First visit, empty cache, throttled&lt;/td&gt;
&lt;td&gt;529 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The throttle is 1.6 Mbps down with a 150 ms round trip. The test runs in headless Chromium against the real built runtime, with one variation that rewrites a heading above the fold.&lt;/p&gt;

&lt;p&gt;Separately, because it is the thing a visitor actually experiences: on that throttled first visit the original heading was on screen first on all 200 loads, for a median of 326 ms.&lt;/p&gt;

&lt;p&gt;On a cached repeat view, no painted frame showed the original on 199 of the 200 loads. On the route change, none did at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why 529 ms is a floor and not a result
&lt;/h2&gt;

&lt;p&gt;The measurement runs against a local origin.&lt;/p&gt;

&lt;p&gt;That means it carries the throttle and the transfer, but none of the DNS, TLS or edge latency a real first visit also pays.&lt;/p&gt;

&lt;p&gt;It also throttles the network without throttling the processor. A first visit on a mid range phone is slower than 529 ms, not faster.&lt;/p&gt;

&lt;p&gt;We would rather say that than let someone quote 529 ms at a prospect as a field figure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the page flickers at all
&lt;/h2&gt;

&lt;p&gt;The runtime is injected as a dynamic &lt;code&gt;&amp;lt;script&amp;gt;&lt;/code&gt;. It does not block the parser.&lt;/p&gt;

&lt;p&gt;That is a deliberate choice, and the cost of it is the 326 ms on that first visit. The original renders, then the variation replaces it.&lt;/p&gt;

&lt;p&gt;The usual fix is an anti-flicker overlay that hides the page while the runtime decides. Ours exists, and it is &lt;strong&gt;unchecked by default&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Checked, the snippet adds a rule and puts the class on &lt;code&gt;&amp;lt;html&amp;gt;&lt;/code&gt; at the top of the page:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight css"&gt;&lt;code&gt;&lt;span class="nc"&gt;.abtestly-async-hide&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;opacity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="m"&gt;0&lt;/span&gt; &lt;span class="cp"&gt;!important&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The class comes off once variants apply, or after two seconds, whichever is first. The page never stays hidden.&lt;/p&gt;

&lt;p&gt;The reason it is off by default is the arithmetic of who pays for it.&lt;/p&gt;

&lt;p&gt;A page-wide hide costs every visitor on every page, including everyone who was never bucketed into any experiment. If the config fetch is ever slow, that is a visibly empty site.&lt;/p&gt;

&lt;p&gt;Flicker only shows on the pages a variant actually changes. So the box is there for the test that needs it, and not before.&lt;/p&gt;

&lt;p&gt;There is a second reason, and it is the one that makes this more than a preference. Our own speed guardrail flags a variation as slower when its p75 largest contentful paint sits 400 ms or more above control. An always-on overlay would ship exactly the regression that guardrail exists to catch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the variant decision happens
&lt;/h2&gt;

&lt;p&gt;In the browser, after the runtime arrives. There is no server hop in the path the benchmark measures.&lt;/p&gt;

&lt;p&gt;Assignment is a deterministic hash with no shared state, so every visitor's bucket is computed from their own id alone:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;hash1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;murmur3&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;salt&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;userId&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/Exposure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;mod&lt;/span&gt; &lt;span class="mi"&gt;10000&lt;/span&gt;
&lt;span class="n"&gt;enter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;hash1&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;trafficAllocationBp&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;salt&lt;/code&gt; is the experiment's immutable UUID, minted at creation and never changed. &lt;code&gt;userId&lt;/code&gt; is a UUIDv7 held in localStorage and a cookie. &lt;code&gt;trafficAllocationBp&lt;/code&gt; is allocation in basis points, so 10000 is 100%.&lt;/p&gt;

&lt;p&gt;Entry is decided first, variant second. The same two-step model Amplitude uses.&lt;/p&gt;

&lt;p&gt;Because it is a pure function of those inputs, the same hash on the same inputs reproduces the same bucket in a Node script. You can check our assignment against your own implementation without asking us anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  The guardrail admits what it is not
&lt;/h2&gt;

&lt;p&gt;This is the part I want to hold up, because it is written in our own docs about our own feature rather than left for a prospect to discover:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;There is no confidence interval, no bootstrap, and no significance test. The panel is a descriptive guardrail, not an inferential one.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So it reports a measurement. It does not prove a variation is safe.&lt;/p&gt;

&lt;p&gt;The docs volunteer the consequence too. A difference just over 400 ms on 100 page loads per arm is not the same evidence as the same difference on 100,000. Same flag, very different weight.&lt;/p&gt;

&lt;p&gt;The panel does not weight them for you. The page load count sits next to the row so that you can.&lt;/p&gt;

&lt;p&gt;Four checks run before a row is judged at all, and they stop at the first failure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;capture rate not above 100%&lt;/li&gt;
&lt;li&gt;at least 100 page loads on both arms&lt;/li&gt;
&lt;li&gt;a capture rate gap of no more than 20 percentage points&lt;/li&gt;
&lt;li&gt;at least 50% capture on each arm&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fail one and the row reads "collecting speed data" or "comparison unavailable", depending on which check it was, rather than guessing.&lt;/p&gt;

&lt;p&gt;There is also no correction across devices or variations. A four variant test on two devices produces six comparisons, each judged on its own threshold. Six independent calls, not one corrected verdict.&lt;/p&gt;

&lt;p&gt;That is the discipline a descriptive guardrail asks of you. Read the page load count behind a flag before acting on it, and read "not flagged" as nothing visible at this volume, rather than as clearance.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to ask your own vendor
&lt;/h2&gt;

&lt;p&gt;Not snippet size.&lt;/p&gt;

&lt;p&gt;Ask for p75 time from navigation start to the variation landing in the DOM, separated by cached and uncached. Ask how long the original was actually painted. Ask for the method, and for the caveats.&lt;/p&gt;

&lt;p&gt;If a vendor cannot produce that, the honest reading is that nobody there has measured it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;ABTestly is A/B testing for teams that write their variations in code rather than in a visual editor. The bucketing algorithm is documented at &lt;a href="https://docs.abtestly.com" rel="noopener noreferrer"&gt;docs.abtestly.com&lt;/a&gt; and the prices are published at &lt;a href="https://abtestly.com/pricing" rel="noopener noreferrer"&gt;abtestly.com/pricing&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>performance</category>
      <category>javascript</category>
      <category>testing</category>
    </item>
    <item>
      <title>Why our A/B testing tool refuses to call a winner</title>
      <dc:creator>The ABTestly team</dc:creator>
      <pubDate>Sat, 26 Sep 2026 11:10:39 +0000</pubDate>
      <link>https://dev.to/abtestly/why-our-ab-testing-tool-refuses-to-call-a-winner-37k9</link>
      <guid>https://dev.to/abtestly/why-our-ab-testing-tool-refuses-to-call-a-winner-37k9</guid>
      <description>&lt;p&gt;Two readings of the same experiment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 4.&lt;/strong&gt; Variant B is up 12.7% on conversion rate. 12,506 visitors in the variant, 12,418 in control. The number is green. Somebody screenshots it and drops it in Slack.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day 11.&lt;/strong&gt; Variant B is up 12.7%. Same test, same split, same traffic mix.&lt;/p&gt;

&lt;p&gt;Same experiment, same lift. The only thing that changed is how long it was allowed to run. One of those two readings is a result. The other is a coin that has come up heads four times.&lt;/p&gt;

&lt;p&gt;Every dashboard I have used will render both of them identically, with a green arrow.&lt;/p&gt;

&lt;h2&gt;
  
  
  The thing that goes wrong is peeking
&lt;/h2&gt;

&lt;p&gt;A standard frequentist A/B test is a fixed-horizon test. You commit to a sample size before you start, you look once when you reach it, and the α you chose (usually 0.05) is the probability of a false positive &lt;strong&gt;for that one look&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The moment you look more than once, that guarantee is gone.&lt;/p&gt;

&lt;p&gt;This is not a subtle effect and it is not our finding; it is the classic result on repeated significance testing. Each additional look is another chance for the random walk of the difference between two proportions to wander across your threshold. Look ten times at α = 0.05 and the real false-positive rate lands far closer to 20% than to 5%. Monitor continuously and, in the limit, you will cross the line eventually with probability approaching 1, whether or not there is any effect at all.&lt;/p&gt;

&lt;p&gt;Now think about how experimentation actually works in a company. The test is live. The dashboard is open in a tab. There is a standup every morning and a roadmap review on Thursday. Nobody looks once.&lt;/p&gt;

&lt;p&gt;So the tooling is not neutral here. A dashboard that renders a day-4 lift the same way it renders a day-11 lift is not reporting a result, it is handing someone a screenshot to win an argument with.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we show instead
&lt;/h2&gt;

&lt;p&gt;Three decisions, all of which make our product look worse in a demo.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The verdict panel says "Still collecting" and means it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not "leading". Not "trending positive". Until the decision boundary is crossed, the headline on the results page is that the test has not concluded, and the lift figure sits underneath it rather than above it. Whoever opens that page has to read the word before they read the number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The confidence interval is always shown next to the sample size.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A lift is meaningless on its own. &lt;code&gt;+12.7%&lt;/code&gt; with a 95% interval of &lt;code&gt;[-2.1%, +28.4%]&lt;/code&gt; at n = 12,506 is a different object from &lt;code&gt;+12.7%&lt;/code&gt; with &lt;code&gt;[+8.9%, +16.6%]&lt;/code&gt; at n = 180,000, and the only way to stop people conflating the two is to refuse to print one without the other. No bare percentages anywhere in the UI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Sample ratio mismatch is on the results page, not in a settings drawer.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you asked for 50/50 and you are getting 50.4/49.6 across 200,000 visitors, something upstream is broken: a redirect that fires before the assignment, a bot filter that treats variants differently, a cache that serves control to a subset. The check is a chi-square goodness of fit against the expected allocation, and a low p-value there invalidates everything above it on the page.&lt;/p&gt;

&lt;p&gt;SRM is the most useful check in experimentation and the one most often buried. If the split is broken, the lift is not wrong, it is meaningless. We put the flag at the top of the result, in red, before the number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sequential is the real answer to peeking
&lt;/h2&gt;

&lt;p&gt;Telling people "don't look" loses to reality every time. The better answer is to use a method designed for looking.&lt;/p&gt;

&lt;p&gt;Sequential tests spend your error budget across the monitoring period rather than at a single endpoint, so continuous monitoring is valid by construction rather than in spite of the maths. You give up some power relative to a perfectly executed fixed-horizon test, and in exchange you get a test that survives the way teams actually behave.&lt;/p&gt;

&lt;p&gt;Our default engine is frequentist because that is what most teams already reason in. Sequential and Bayesian are available on the higher plans for teams that want to monitor continuously or reason about probability-to-beat-control directly. Whichever engine you pick, the method version is locked when the test starts, so nobody can switch engines halfway and pick the flattering one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why any vendor would build this
&lt;/h2&gt;

&lt;p&gt;Here is the part that is uncomfortable to write.&lt;/p&gt;

&lt;p&gt;Every incumbent's renewal conversation depends, in some form, on the customer believing their testing programme is working. A dashboard that surfaces wins is commercially aligned with that. A dashboard that says "still collecting" for nine days is not. "Still collecting" is a worse demo than a trophy, and I have watched it be a worse demo.&lt;/p&gt;

&lt;p&gt;We can afford to build it because we are small, we publish our prices, and we are not defending a renewal number against a board. That is a structural advantage rather than a virtuous one, and it will not last forever.&lt;/p&gt;

&lt;p&gt;But if you sell a statistics product, refusing to overclaim is the product. Everything else is a rendering detail.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;ABTestly is A/B testing for teams that write their variations in code rather than in a visual editor. Docs at &lt;a href="https://docs.abtestly.com" rel="noopener noreferrer"&gt;docs.abtestly.com&lt;/a&gt;, and the pricing is published at &lt;a href="https://abtestly.com/pricing" rel="noopener noreferrer"&gt;abtestly.com/pricing&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>javascript</category>
      <category>testing</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
