<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Anton</title>
    <description>The latest articles on DEV Community by Anton (@kharlamov).</description>
    <link>https://dev.to/kharlamov</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4160070%2Fb985379f-e0f8-416d-afa3-d40202fd5728.png</url>
      <title>DEV Community: Anton</title>
      <link>https://dev.to/kharlamov</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kharlamov"/>
    <language>en</language>
    <item>
      <title>Statistical Tests in TypeScript Without a Separate Python Service</title>
      <dc:creator>Anton</dc:creator>
      <pubDate>Sat, 03 Oct 2026 16:10:45 +0000</pubDate>
      <link>https://dev.to/kharlamov/statistical-tests-in-typescript-without-a-separate-python-service-36c8</link>
      <guid>https://dev.to/kharlamov/statistical-tests-in-typescript-without-a-separate-python-service-36c8</guid>
      <description>&lt;p&gt;Imagine an internal tool that compares two variants of a workflow.&lt;br&gt;
The application already has:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the measurements;&lt;/li&gt;
&lt;li&gt;    filters;&lt;/li&gt;
&lt;li&gt;    tables;&lt;/li&gt;
&lt;li&gt;    charts;&lt;/li&gt;
&lt;li&gt;    user interactions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now you want to add one more thing:&lt;br&gt;
a proper statistical comparison between the two groups.&lt;/p&gt;

&lt;p&gt;One common architecture is to send the data to a separate Python service.&lt;br&gt;
Sometimes that is absolutely the right decision.&lt;br&gt;
But sometimes it introduces a network boundary for a calculation that could have remained local to the application.&lt;br&gt;
That made me wonder:&lt;br&gt;
How much statistics can reasonably stay inside a TypeScript codebase?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The gap between mean() and an analytics platform&lt;/strong&gt;&lt;br&gt;
JavaScript already makes simple descriptive statistics easy.&lt;br&gt;
Calculating an average is trivial.&lt;br&gt;
The interesting part starts when you need things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;confidence intervals;&lt;/li&gt;
&lt;li&gt;hypothesis tests;&lt;/li&gt;
&lt;li&gt;ANOVA;&lt;/li&gt;
&lt;li&gt;regression diagnostics;&lt;/li&gt;
&lt;li&gt;control charts;&lt;/li&gt;
&lt;li&gt;power calculations;&lt;/li&gt;
&lt;li&gt;time-series methods.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At that point, many teams immediately leave the JavaScript ecosystem.&lt;br&gt;
Columna takes a different approach.&lt;br&gt;
It has a separate columna/advanced entry point for statistical procedures.&lt;br&gt;
For example, here is a small two-sample comparison.&lt;br&gt;
The data below is synthetic and exists only to demonstrate the API.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;DataFrame&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;columna&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;formatReport&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;columna/advanced&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;measurements&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromRows&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
  &lt;span class="p"&gt;...[&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;38&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;45&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;41&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;39&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;44&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;seconds&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;variant&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;A&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;seconds&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;})),&lt;/span&gt;

  &lt;span class="p"&gt;...[&lt;/span&gt;&lt;span class="mi"&gt;40&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;39&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;38&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;41&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;43&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;seconds&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;variant&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;B&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;seconds&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;})),&lt;/span&gt;
&lt;span class="p"&gt;]);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;measurements&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ttest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;seconds&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;by&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;variant&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;equalVar&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;formatReport&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;markdown&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This uses Welch's two-sample t-test.&lt;br&gt;
The API returns more than a p-value.&lt;br&gt;
The result includes values such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;test statistic;&lt;/li&gt;
&lt;li&gt;degrees of freedom;&lt;/li&gt;
&lt;li&gt;estimate;&lt;/li&gt;
&lt;li&gt;confidence interval;&lt;/li&gt;
&lt;li&gt;sample information.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And because the result is still a regular JavaScript object, the UI can decide how much of it to present.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why this can simplify application architecture&lt;/strong&gt;&lt;br&gt;
Suppose a user changes a department filter.&lt;br&gt;
The application could:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;filter the observations;&lt;/li&gt;
&lt;li&gt;run the statistical procedure;&lt;/li&gt;
&lt;li&gt;update the result table.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;No HTTP request is inherently required for that flow.&lt;br&gt;
That does not mean every statistics workload belongs in the browser or Node.js.&lt;br&gt;
It means the service boundary can be chosen for architectural reasons rather than because the calculation happens to require another language.&lt;br&gt;
That distinction matters.&lt;br&gt;
If the workload is already handled by a mature Python analytics service with established validation, there may be no reason to move it.&lt;br&gt;
But if you only need a small, interactive statistical feature inside an existing TypeScript application, introducing a second runtime can feel disproportionately expensive.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keeping data preparation and analysis together&lt;/strong&gt;&lt;br&gt;
There is another practical benefit.&lt;br&gt;
The code that decides which observations enter the analysis can stay next to the code that performs the analysis.&lt;br&gt;
For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;filtered&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;measurements&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;seconds&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;gt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;collect&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;filtered&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ttest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;seconds&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;by&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;variant&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;equalVar&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;During review, the team can inspect the entire path:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;which rows are included;&lt;/li&gt;
&lt;li&gt;which rows are excluded;&lt;/li&gt;
&lt;li&gt;how groups are defined;&lt;/li&gt;
&lt;li&gt;which statistical method is used.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is often easier than tracing the same calculation across frontend code, an HTTP contract, another service, and a separate statistics layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The short API call is not the difficult part&lt;/strong&gt;&lt;br&gt;
This is important:&lt;br&gt;
A statistical function returning a number does not make the analysis correct.&lt;br&gt;
The hard questions still exist.&lt;br&gt;
Are observations independent?&lt;br&gt;
Are the samples paired?&lt;br&gt;
Is the chosen test appropriate?&lt;br&gt;
How are missing values handled?&lt;br&gt;
What effect size actually matters?&lt;br&gt;
What should the UI show when the result is uncertain?&lt;br&gt;
For example, if the same people were measured before and after a change, treating those values as two independent samples would be questionable.&lt;br&gt;
You would want a paired analysis instead.&lt;br&gt;
No library can infer the experimental design from a column name.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Please don't turn everything into "p &amp;lt; 0.05"&lt;/strong&gt;&lt;br&gt;
If statistics becomes easier to integrate into an application, there is also a risk that the UI reduces everything to:&lt;br&gt;
significant / not significant&lt;br&gt;
That is usually not a good interface.&lt;br&gt;
At minimum, I would prefer to show:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;sample sizes;&lt;/li&gt;
&lt;li&gt;descriptive statistics;&lt;/li&gt;
&lt;li&gt;estimated difference;&lt;/li&gt;
&lt;li&gt;confidence interval;&lt;/li&gt;
&lt;li&gt;the chosen method.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Users should be able to understand what was compared.&lt;br&gt;
Developers should be able to reproduce the calculation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What about correctness?&lt;/strong&gt;&lt;br&gt;
If I am going to use statistics inside an application, the first thing I care about is not the number of supported functions.&lt;br&gt;
It is whether the implementation has been tested against established references.&lt;br&gt;
Columna's advanced module is built around cross-checks against SciPy, NumPy, NIST references, property tests, and Monte Carlo tests.&lt;br&gt;
That is a useful starting point.&lt;br&gt;
But I would still validate any important production workflow against a known reference dataset before shipping it.&lt;br&gt;
A statistical library should reduce implementation work.&lt;br&gt;
It should not eliminate verification.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where I think this approach makes sense&lt;/strong&gt;&lt;br&gt;
This is especially interesting for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;internal tools;&lt;/li&gt;
&lt;li&gt;QA dashboards;&lt;/li&gt;
&lt;li&gt;experiment explorers;&lt;/li&gt;
&lt;li&gt;educational applications;&lt;/li&gt;
&lt;li&gt;local analysis utilities;&lt;/li&gt;
&lt;li&gt;interactive engineering reports.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a large data platform, I would still expect heavy processing to live in something like DuckDB, Polars, a warehouse, or an existing analytics service.&lt;br&gt;
But small statistical features do not always need to inherit that architecture.&lt;br&gt;
Sometimes the most useful thing a library can do is not replace a platform.&lt;br&gt;
It is to keep a small feature small.&lt;/p&gt;

&lt;p&gt;The examples use columna/advanced:&lt;br&gt;
&lt;a href="https://github.com/ankhitlab/columna" rel="noopener noreferrer"&gt;https://github.com/ankhitlab/columna&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I'm curious where people draw this boundary in real projects: which statistical operations would you keep inside a TypeScript application, and which would you always move to a dedicated analytics stack?&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>opensource</category>
      <category>typescript</category>
      <category>datascience</category>
    </item>
  </channel>
</rss>
