Imagine an internal tool that compares two variants of a workflow.
The application already has:
- the measurements;
- filters;
- tables;
- charts;
- user interactions.
Now you want to add one more thing:
a proper statistical comparison between the two groups.
One common architecture is to send the data to a separate Python service.
Sometimes that is absolutely the right decision.
But sometimes it introduces a network boundary for a calculation that could have remained local to the application.
That made me wonder:
How much statistics can reasonably stay inside a TypeScript codebase?
The gap between mean() and an analytics platform
JavaScript already makes simple descriptive statistics easy.
Calculating an average is trivial.
The interesting part starts when you need things like:
- confidence intervals;
- hypothesis tests;
- ANOVA;
- regression diagnostics;
- control charts;
- power calculations;
- time-series methods.
At that point, many teams immediately leave the JavaScript ecosystem.
Columna takes a different approach.
It has a separate columna/advanced entry point for statistical procedures.
For example, here is a small two-sample comparison.
The data below is synthetic and exists only to demonstrate the API.
import { DataFrame } from 'columna';
import { formatReport } from 'columna/advanced';
const measurements = DataFrame.fromRows([
...[42, 38, 45, 41, 39, 44].map((seconds) => ({
variant: 'A',
seconds,
})),
...[40, 39, 42, 38, 41, 43].map((seconds) => ({
variant: 'B',
seconds,
})),
]);
const result = measurements.ttest('seconds', {
by: 'variant',
equalVar: false,
});
console.log(formatReport(result, 'markdown'));
This uses Welch's two-sample t-test.
The API returns more than a p-value.
The result includes values such as:
- test statistic;
- degrees of freedom;
- estimate;
- confidence interval;
- sample information.
And because the result is still a regular JavaScript object, the UI can decide how much of it to present.
Why this can simplify application architecture
Suppose a user changes a department filter.
The application could:
- filter the observations;
- run the statistical procedure;
- update the result table.
No HTTP request is inherently required for that flow.
That does not mean every statistics workload belongs in the browser or Node.js.
It means the service boundary can be chosen for architectural reasons rather than because the calculation happens to require another language.
That distinction matters.
If the workload is already handled by a mature Python analytics service with established validation, there may be no reason to move it.
But if you only need a small, interactive statistical feature inside an existing TypeScript application, introducing a second runtime can feel disproportionately expensive.
Keeping data preparation and analysis together
There is another practical benefit.
The code that decides which observations enter the analysis can stay next to the code that performs the analysis.
For example:
const filtered = await measurements
.filter((c) => c.seconds.gt(0))
.collect();
const result = filtered.ttest('seconds', {
by: 'variant',
equalVar: false,
});
During review, the team can inspect the entire path:
- which rows are included;
- which rows are excluded;
- how groups are defined;
- which statistical method is used.
That is often easier than tracing the same calculation across frontend code, an HTTP contract, another service, and a separate statistics layer.
The short API call is not the difficult part
This is important:
A statistical function returning a number does not make the analysis correct.
The hard questions still exist.
Are observations independent?
Are the samples paired?
Is the chosen test appropriate?
How are missing values handled?
What effect size actually matters?
What should the UI show when the result is uncertain?
For example, if the same people were measured before and after a change, treating those values as two independent samples would be questionable.
You would want a paired analysis instead.
No library can infer the experimental design from a column name.
Please don't turn everything into "p < 0.05"
If statistics becomes easier to integrate into an application, there is also a risk that the UI reduces everything to:
significant / not significant
That is usually not a good interface.
At minimum, I would prefer to show:
- sample sizes;
- descriptive statistics;
- estimated difference;
- confidence interval;
- the chosen method.
Users should be able to understand what was compared.
Developers should be able to reproduce the calculation.
What about correctness?
If I am going to use statistics inside an application, the first thing I care about is not the number of supported functions.
It is whether the implementation has been tested against established references.
Columna's advanced module is built around cross-checks against SciPy, NumPy, NIST references, property tests, and Monte Carlo tests.
That is a useful starting point.
But I would still validate any important production workflow against a known reference dataset before shipping it.
A statistical library should reduce implementation work.
It should not eliminate verification.
Where I think this approach makes sense
This is especially interesting for:
- internal tools;
- QA dashboards;
- experiment explorers;
- educational applications;
- local analysis utilities;
- interactive engineering reports.
For a large data platform, I would still expect heavy processing to live in something like DuckDB, Polars, a warehouse, or an existing analytics service.
But small statistical features do not always need to inherit that architecture.
Sometimes the most useful thing a library can do is not replace a platform.
It is to keep a small feature small.
The examples use columna/advanced:
https://github.com/ankhitlab/columna
I'm curious where people draw this boundary in real projects: which statistical operations would you keep inside a TypeScript application, and which would you always move to a dedicated analytics stack?
Top comments (1)
Keeping the row selection next to the test is a good argument, and returning the CI and degrees of freedom rather than a bare p-value matters a lot in a UI.
Two things from doing statistics in TypeScript for a production tool (validators for trading backtests):
The interactive flow you describe, change a filter and rerun, is also a multiple-testing machine. A user who tries ten department filters and stops at the first p < 0.05 will find one about 40% of the time even when nothing differs (1 - 0.95^10). Since the test now runs locally, the app knows how many comparisons were made in the session. Showing "test 7 of this session" next to the result, or offering a Holm adjustment over the filters tried, would turn the architecture advantage into a correctness advantage too.
Cross-check against a second implementation in another language. When we wrote an independent Python verifier, working only from the spec, for one of our JavaScript data formats, it found six places where the two read the same input differently: timestamp precision (Python compares microseconds, JS milliseconds), numbers of 1e15 or more, NaN literals (Python's json accepts them, JS rejects them), and the sort order of non-ASCII keys, among others. A statistics library has the same exposure at its edges, and a small fixture file of inputs and expected outputs computed with scipy, run in CI, catches that drift cheaply.
Our implementation is open source if it's a useful reference: github.com/arhancanli/canli-valida...