DEV Community

Cover image for DBench86 — A Style Exercise in PHP 8.6
Pascal CESCATO
Pascal CESCATO Subscriber Community Curator

Posted on

DBench86 — A Style Exercise in PHP 8.6

A benchmark as an excuse to build a PHP 8.6 app, and PHP 8.6 as an excuse to build a benchmark

French readers may prefer to read this article in their native language → version française.

It all started with a benchmark

Databases have always been my playground. From dBASE III+ in the 90s to MyScaleDB today, by way of XML databases in the 2000s… I've always looked for the best solution, the best approach, depending on the project. I've written quite a bit about this over the years, and I launched dbgrade.tech to audit database setups.

Recently, I read a report comparing several engines: MariaDB 10.3, MySQL 8.0, and PostgreSQL 12.

And the further I got into the report, the more something bothered me.

The engines weren't compared under the same conditions, and the versions were years old — 2018 and 2019, while we're in 2026.

At that point, what are you actually measuring?

The database engine? The amount of memory? The CPU? The storage? An outdated version of the software?

Probably a bit of all of it at once.

That kind of benchmark can make sense in a specific context. If I'm preparing a migration away from an old PostgreSQL version, for instance, comparing that version against a recent release of another engine is perfectly legitimate.

But that wasn't the question being asked, and the benchmark left me with an unfinished feeling.

Then PHP 8.6 showed up. In no hurry.

A bit later, I came across the PHPenomenal 8.6 challenge on daily.dev, put together by Damien Séguy for the PHP 8.6 release.

Damien Seguy is a well-known figure in the French PHP community. And PHP is a language I know pretty well — I've been using it in my work for a long time.

The challenge amused me: seven new language features to bring to life in a real application, not in seven isolated demo files. Things like this:

interface DatabaseDriver {
    public string $name { get; }
}

final class MariaDbDriver implements DatabaseDriver {
    public readonly string $name = 'MariaDB';
}
Enter fullscreen mode Exit fullscreen mode

A readonly property satisfying an interface contract, with no constructor, no getter. Nice idea. Now I just needed to find where it would actually be useful, not merely present.

I never do a challenge just to cram a few new language features into a pointless program. Counting how many times my dog barks in a day, or how many times my neighbor turns their light on at night, doesn't interest me much.

If I'm going to write code, it might as well be good for something. And that's where the two topics met.

Why not use the PHP 8.6 challenge to build a performance-comparison tool that actually tries to compare things that are comparable?

That's how DBench86 was born.

Comparing what's actually comparable

On paper, it's simple: latest stable versions, same machine, same data, same queries.

In practice, I was about to discover that "fair" is a much more complicated word than it looks.

First, the versions

I didn't want to compare 2019 or even 2020 stable releases. What would be the point? They're end of life.

DBench86 therefore uses, at the time of the benchmark, the latest stable release available for each of the three engines.

That choice obviously doesn't mean all three versions are the same age, or that the three projects follow the same release cadence. What I care about is much simpler: if I have to pick a database engine for a new application today, which versions would I actually install?

The current stable ones.

Comparing an old PostgreSQL release against the latest MariaDB release can make sense if I'm preparing a migration. But that would be a different benchmark, answering a different question.

Here, I want to compare what MariaDB, PostgreSQL and SQLite offer today.

Then, the machine

The second rule was just as obvious: DBench86 had to run every engine on the same machine.

Not a 2 GB VPS for MariaDB and a 4 GB one for PostgreSQL. Not two vCPUs here and four there. Not an SSD on one side and an NVMe drive on the other.

One single machine.

Same CPU, same memory, same storage, same operating system, same PHP version.

That way, when PostgreSQL takes 200 ms where MariaDB takes 300, I know the 100 ms gap doesn't come from a CPU that's twice as fast on one side.

Put like that, it sounds obvious.

Apparently it isn't, not for everyone. Some people are still adding apples and oranges when they should be adding fruit — or not adding at all.

Three engines, but not three different ways of testing them

I ended up with MariaDB, PostgreSQL and SQLite.

The first two are fairly naturally comparable: they're full database servers, installed on the same machine, queried by the same PHP application.

SQLite is a bit different. There's no SQLite server to configure, no network connection, no background service. In DBench86, the database is even created in memory.

That also changes the durability picture: MariaDB and PostgreSQL honor fsync even with default settings, while SQLite in memory writes nothing to disk at all. That's not a bias to fix — it's a real characteristic of this mode, and it's precisely a legitimate reason to use it — but it's worth stating clearly rather than letting readers wonder why SQLite starts with a structural advantage on certain operations.

So, is this actually fair? Yes… and no.

If the goal was to crown a winner by adding up the times and handing out medals at the end, probably not.

But that's not the goal.

SQLite is one of the options I can choose for storing an application's data. In some cases it'll be a perfect fit. In others, the very idea of using it would be absurd.

I'm not trying to find out which is the best database engine. That question doesn't mean much, honestly.

I'm trying to see how each one behaves under the same operations, and above all how that behavior changes as the data volume grows.

At 10,000 rows, SQLite can easily humiliate everyone else on some operations.

At a million?

We'll see.

Making everyone work on the same data

I chose to generate the dataset myself, deterministically.

Same starting point, same generator, same logical data.

A primary-key lookup has to search for the same keys. A range selection has to use the same bounds. An update has to be comparable to the one run against the other engines.

DBench86 runs ten workloads:

  • primary-key lookup;
  • lookup on an indexed column;
  • range selection;
  • join;
  • aggregation;
  • single insert;
  • batch insert;
  • transactional insert;
  • update by primary key;
  • delete by primary key.

Nothing particularly exotic.

And that's deliberate.

I'm not hunting for the one SQL query twisted enough to bring an engine to its knees, just so I can triumphantly announce that its neighbor is twenty times faster.

I want operations you'd actually find in real applications.

What about the SQL itself?

PostgreSQL, MariaDB and SQLite obviously don't share the same engine. They don't share the same optimizer, and not necessarily the same way of executing a query.

I could start adapting the queries: a small optimization for PostgreSQL here, another for MariaDB there, special handling for SQLite.

And a few hours later, I'd have three different benchmarks.

That's not what I wanted.

DBench86 describes the work to be done with comparable SQL queries, then lets each engine do what it was built to do: optimize and execute it.

I could run an EXPLAIN and analyze the execution plans. That would be interesting — but that's a different topic.

The moment I start patching my code to work around one engine's or another's weaknesses, I'm no longer really comparing engines. I'm comparing my own ability to optimize for each of them individually.

Raw first, tuned second

There are actually two interesting questions here:

“
What do I get when I install the engine and use it with its default configuration?
”

And:

“
What do I get when I take the time to properly configure that same engine for the machine I have?
”

So I needed two campaigns: stock, then tuned.

With one rule that never changes: DBench86 itself stays exactly the same.

Same versions. Same PHP code. Same PDO configuration. Same schema. Same data generation. Same queries. Same workloads. Same machine.

The only variable between the two campaigns is the server configuration.

Simple on paper.

And yet, a bit later, I managed to fall straight into the exact trap I was trying to avoid.

At a million rows, everything stops

The first runs go fine.

10,000 rows, no problem. 100,000 rows, still no problem.

So naturally, I bump it up to a million. And then:

Fatal error: Allowed memory size of 134217728 bytes exhausted
Enter fullscreen mode Exit fullscreen mode

128 MB.

The easiest fix would have been to bump memory_limit.

512 MB? 1 GB? I had the memory available, after all.

But wait…

I'd started this project precisely because a benchmark that changes its resources whenever they become inconvenient had annoyed me.

I wasn't about to do the exact same thing.

And more to the point — why would a database benchmark need several hundred megabytes of PHP memory just to test a database with a million rows?

That data isn't something PHP needs to store.

It lives in the database.

The problem wasn't the database

The problem was in how the workloads were prepared.

My first implementation kept far too much information in memory. The bigger the dataset grew, the more data PHP was holding onto alongside it.

But how many values did I actually need to keep?

The benchmark runs 5 warmups, then 30 measured runs. 35 executions. Not a million.

Preparation was effectively running at:

O(number of rows)
Enter fullscreen mode Exit fullscreen mode

when it could have been reduced to:

O(warmups + runs)
Enter fullscreen mode Exit fullscreen mode

With my usual parameters:

O(35)
Enter fullscreen mode Exit fullscreen mode

A bit better.

Fixing the benchmark instead of raising the memory limit

So I rewrote the dataset generation.

Data is now generated progressively and inserted in batches. DBench86 only keeps in memory what it will actually need for upcoming executions.

Everything else is sent to the database and then forgotten by PHP.

I run it again.

10,000 rows — fine.

100,000 — fine.

A million… also fine.

I didn't raise the limit so the benchmark could survive a million rows.

I removed the reason it needed all that memory in the first place.

range_select and the trap of false fairness

Looking more closely at the tests, another one caught my attention: range_select.

Over large ranges, the result set can represent tens or hundreds of thousands of rows.

And a question comes up:

what am I actually measuring here?

The time the engine needs to run the query? To send back the results? For PDO to fetch them? Some mix of all three?

It would have been very easy to just keep the numbers.

They looked serious.

There was a median, a P95, three decimal places…

So surely it had to be scientific.

Obviously not.

I audited this workload separately: deterministic bounds, identical cardinalities, an order-independent checksum to verify the content.

There's deliberately no ORDER BY: the order isn't used by the benchmark, so why ask the engine to perform a sort I don't need?

A better technical solution… answering the wrong question

For PostgreSQL, I experimented with streaming retrieval. Technically, it was clean and efficient.

But DBench86 was starting to know it was talking to PostgreSQL specifically, and to apply a particular tweak just for it.

I was doing exactly what I'd forbidden myself from doing.

Not with RAM this time. With the driver.

The optimization would have been a good idea if my goal had been to build the most efficient possible PHP application for PostgreSQL.

But I was building a benchmark. So I pulled the optimization back out.

MariaDB uses PDO's default behavior.

PostgreSQL uses PDO's default behavior.

SQLite uses PDO's default behavior.

And in all three cases, DBench86 consumes the entire result set inside the measured window.

Fair doesn't mean identical

A fair benchmark isn't about making the three engines behave identically.

It's about giving them the same problem under the same conditions, then accepting that they won't solve it the same way.

The rule became:

“
DBench86 must not make the engines identical. It must avoid artificially favoring any one of them.
”

Stock vs. tuned: this time, touching the settings is allowed

For the first campaign, I keep the configuration exactly as installed.

First, I record what's actually there.

And I archive it.

In my case, MariaDB started with a 128 MiB InnoDB buffer pool and a 96 MiB redo log, with a limit of 151 connections.

PostgreSQL likewise had 128 MiB of shared_buffers, 4 MiB of work_mem, and 100 possible connections.

Those are the values I tested.

Not the ones from the documentation.

The ones actually in use on the servers.

Then we tune

For the second campaign, I allow myself to touch the servers. The machine has just under 4 GiB of RAM, no swap.

I had to think in terms of a shared budget, not tune each server as if it were alone on the machine.

For MariaDB, the InnoDB buffer pool goes from 128 to 768 MiB. The redo log goes from 96 to 256 MiB. I add a small 16 MiB MyISAM cache and lower the maximum connection count from 151 to 32.

For PostgreSQL, shared_buffers goes from 128 to 512 MiB. I raise some of the maintenance and WAL margins, space checkpoints out to ten minutes, bring connections down to 32 as well, and reduce the parallelism limits.

And work_mem? 4 MiB. I leave it alone.

Because work_mem isn't "the amount of memory PostgreSQL can use." It's an allowance that can be consumed multiple times over.

The full details of both configurations, value by value, are available for anyone who wants to verify or reproduce them: stock configuration and tuned configuration.

And SQLite?

Nothing.

DBench86 uses a private in-memory SQLite database. I could tweak some PRAGMA settings or try to optimize SQLite specifically.

But that's exactly what I don't want to do.

Tuned doesn't mean I have to change something for every single participant.

Tuning, yes. Touching durability, no.

I can make a database a lot faster if I'm willing to make it less safe.

So I left the fundamental durability settings alone.

I want to measure the effect of a configuration better suited to the machine.

Not answer the question:

“
How fast can a database go if keeping my data becomes optional?
”

Except a benchmark isn't a stopwatch

A database has caches. The operating system has caches. Storage has its own behavior.

A benchmark result of 0.154 ms versus 0.168 ms doesn't necessarily mean the first engine is 9% faster.

That's why DBench86 starts with warmups, excluded from the statistics, then runs 30 measured runs.

I mostly care about the median and the P95. The median gives a good sense of typical behavior. The P95 starts telling the story of what happens when things go a bit less smoothly.

Now, the numbers.

At this point, I no longer had any room to change the rules.

The benchmark was written. The workloads were defined. The versions were chosen. The stock configurations had been preserved, and the tuned ones prepared.

And above all, I'd decided all of this before knowing the results.

Three scales: 10,000, 100,000 and 1,000,000 rows. Ten workloads. Thirty measured runs, five warmups. Always 128 MB for PHP.

Test environment

OS Ubuntu 26.04 LTS
CPU 2 vCPUs, Intel Core Processor (Haswell, no TSX)
RAM 3.73 GiB, no swap
Storage 40 GB NVMe
PHP 8.6.0-dev
MariaDB 11.8.6-MariaDB-5ubuntu0.1
PostgreSQL 18.6 (Ubuntu 18.6-0ubuntu0.26.04.1)
SQLite 3.46.1

Tuning doesn't mean "making it faster"

At 10,000 rows, tuning doesn't work miracles. It can even backfire.

On MariaDB, pk_select goes from a 0.117 ms median in stock to 0.194 ms after tuning. indexed_select goes from 0.140 to 0.208 ms.

On PostgreSQL, single_insert goes from 0.408 to 0.660 ms, and batch_insert from 6.772 to 9.739 ms.

Bumping up buffers is not a magic incantation.

Let's move on to a million rows.

MariaDB workload Stock Tuned Ratio
pk_select 0.753 ms 0.158 ms ×4.77
indexed_select 0.948 ms 0.185 ms ×5.12
join 1.190 ms 0.322 ms ×3.70
range_select 1,061.993 ms 353.871 ms ×3.00
transactional_insert 4.499 ms 2.325 ms ×1.94
aggregation 2,107.764 ms 2,067.632 ms ×1.02

Depending on the workload, tuning ranges from barely anything to a factor well above five.

That's exactly why I'm wary of a sentence like "MariaDB is three times faster after tuning."

Which query? At what scale? Three times faster at what, exactly?

PostgreSQL politely refuses to follow the script

At a million rows, range_select comes in at a median of 132.538 ms on PostgreSQL stock.

After tuning: 144.518 ms. Slightly worse.

And I quite like this result, because it stops me from telling too tidy a story.

On aggregation, we go from 1,535.921 to 1,505.447 ms.

An improvement of about 2%. Not exactly champagne-popping territory. It's barely enough to even interpret.

Then there's SQLite

At a million rows, tuned configuration:

Workload MariaDB PostgreSQL SQLite
pk_select 0.158 0.122 0.004
indexed_select 0.185 0.126 0.005
join 0.322 0.219 0.016
range_select 353.871 144.518 208.464
aggregation 2,067.632 1,505.447 848.521
batch_insert 6.681 7.529 5.157
transactional_insert 2.325 2.746 0.437

Median times in milliseconds. SQLite received no tuning at all — the two measurements above and in the scaling table further down come from different sessions, hence a slight natural run-to-run variance.

At first glance, you could write:

SQLite demolishes everyone else.

And I'd have immediately reproduced the exact kind of benchmark that made me want to write DBench86 in the first place.

SQLite runs in-process, with no client/server exchange at all. These numbers are real.

But how you interpret them matters more than how you rank them.

The moment you look at range_select, PostgreSQL takes the lead. Then aggregation reshuffles the deck again.

There's still no universal winner.

There are workloads.

What I care about most: what happens as the database grows

Take range_select. With median times in milliseconds, here's the table:

Engine 10k 100k 1M
MariaDB stock 2.530 35.317 1,061.993
MariaDB tuned 2.544 34.852 353.871
PostgreSQL stock 1.684 13.807 132.538
PostgreSQL tuned 1.592 14.142 144.518
SQLite 1.207 15.716 226.528

At 10,000 rows, MariaDB stock and tuned are basically indistinguishable.

At 100,000, still basically nothing.

Then, at a million:

1,061.993 vs. 353.871 ms.

The tuning that seemed to do nothing suddenly becomes decisive.

PostgreSQL tells nearly the opposite story.

This, I think, is the main finding:

the differences between engines, and the effect of tuning, show up at a given workload and at a given scale.

That's exactly what a benchmark reduced to a single final score hides.

The 128 MB that was never actually a problem

At a million rows, DBench86 hits a PHP peak of 12 MiB, still with a 128 MiB limit in place.

That's obviously not the memory MariaDB, PostgreSQL or SQLite themselves are consuming. This measures the PHP harness's own memory.

But that was precisely the point.

The dataset can scale from 10,000 to a million rows without DBench86 needing to load a million PHP objects into memory just to pretend it's testing scalability.

So, which one's the best?

I could end this article with a ranking.

MariaDB wins here. PostgreSQL wins there. SQLite crushes both elsewhere.

That would be simple. It would also be exactly what I was trying to avoid.

A benchmark doesn't answer the question "which DBMS is the best?" — it answers much more precise questions.

How do these three engines behave on the same machine, with the same resources, on the same SQL operations?

What happens as the volume goes from 10,000 to 100,000, then to a million rows?

What changes when you tune the configuration for the machine, without changing the benchmark itself?

And above all: do the conclusions stay the same as the scale changes?

The answer is no.

There are three winners

SQLite is formidable when you stay in its natural territory.

PostgreSQL shows a particularly interesting behavior on certain reads as the volume grows.

MariaDB tells yet another story: on some operations, its stock configuration becomes a real liability as the dataset grows, and a well-tuned configuration changes the result dramatically.

And on others, almost nothing changes at all.

There isn't one winner. There are three — or none, depending on how you want to look at it.

A benchmark is only valid under the conditions it was run in

DBench86 doesn't claim to have eliminated every bias. I simply tried to eliminate as many as I could.

Same machine, same resources.

Current stable versions of each engine at the time of testing.

Same deterministic dataset, same operations, same number of runs, same warmup, same 128 MB PHP limit.

And when I wanted to measure the effect of tuning, I changed the tuning, and nothing else.

Nothing spectacular about any of it. Just discipline.

And where does PHP 8.6 fit into all this?

It's worth remembering that this whole story started there. Well, almost. It was more of a catalyst, really.

I came across the PHPenomenal 8.6 challenge organized by Damien Séguy.

I've been using PHP for a long time — I even hold the Zend Certified Engineer PHP 5 certification, from 2008. I use it daily in business applications and to extend WordPress.

I just needed to find a lightweight application that would actually justify using PHP 8.6's new features.

Not an artificial demo where I'd use clamp() purely because clamp() had to go somewhere. You could, in principle, build a class that chains all seven features one after another, or nests them inside each other — but where's the interest in that, beyond the exercise of style itself?

DBench86 uses the challenge's seven mandatory features in actual application behavior: Partial Function Application, Time\Duration, clamp(), #[\Override] on constants, writes on const-held objects, readonly property defaults, and the SortDirection enum — plus an eighth, optional feature from the bonus list: parameter DocComments.

And each of them ended up finding its place naturally.

Partial Function Application: prepare the run once

Workload specialization, for instance, comes down to this expression:

$runOnThisDriver = $this->runMeasuredWorkload($timer, ?, $pdo);

$measurements = [];
foreach ($workloads as $workload) {
    $measurement = $runOnThisDriver($workload);
    $measurements[$workload->name()] = $measurement;
}
Enter fullscreen mode Exit fullscreen mode

The ? here is the Partial Function Application placeholder. The timer and the PDO connection are supplied once; the resulting partial function only waits for the workload.

The method being called stays very simple:

private function runMeasuredWorkload(
    WorkloadTimer $timer,
    Workload $workload,
    PDO $pdo,
): WorkloadMeasurement {
    return $timer->run($workload, $pdo);
}
Enter fullscreen mode Exit fullscreen mode

So this isn't PFA bolted on to please the checker — it's what expresses how DBench86 applies the same measurement mechanism to every workload.

Time\Duration: gives the timeout a proper unit

The timeout arrives from the command line as seconds:

$executor = new WorkloadExecutor(
    Duration::fromSeconds($timeoutSeconds)
);
Enter fullscreen mode Exit fullscreen mode

But execution time is measured with hrtime(true), in nanoseconds. DBench86 converts that duration before comparing it to the timeout:

$endTime = hrtime(true);
$elapsedNanoseconds = $endTime - $startTime;

$elapsedDuration = Duration::fromNanoseconds($elapsedNanoseconds);
$isTimeout = !$executionFailed &&
             Duration::compare($elapsedDuration, $this->timeout) > 0;
Enter fullscreen mode Exit fullscreen mode

This is exactly the kind of spot where a dedicated type is more meaningful than a raw number you have to remember the unit of — seconds, milliseconds, nanoseconds.

The timeout is deliberately after the fact: it classifies a run that took too long, it doesn't cancel an in-flight SQL query.

clamp(): respects the engine's limits

Batch inserts raise a very concrete problem: not every engine accepts the same operational limits.

DBench86 computes the selected driver's limit, then bounds the requested size:

$driverLimit = BatchLimitCalculator::driverLimit(
    $this->driver,
    $this->parametersPerRow
);

$this->effectiveBatchSize = BatchLimitCalculator::effectiveBatchSize(
    $this->requestedBatchSize,
    $driverLimit
);
Enter fullscreen mode Exit fullscreen mode

And effectiveBatchSize() ultimately boils down to:

return clamp($requestedBatchSize, 1, $commonLimit);
Enter fullscreen mode Exit fullscreen mode

Here again, the new feature cleanly replaces a combination of min(), max(), or a few branches. And it actually affects the size of the batches sent to the database.

#[\Override]: checks constants too

Each driver has its own limits. MariaDB, for instance, specializes the ones defined by the abstract driver:

#[\Override]
const int MAX_BATCH_SIZE = 2_000;

#[\Override]
const int MAX_BOUND_PARAMETERS = 50_000;
Enter fullscreen mode Exit fullscreen mode

PostgreSQL and SQLite do the same with their own values.

The point of #[\Override] on a constant isn't to make the program faster. It's to make the intent explicit and verifiable: these constants are meant to replace the ones inherited from the parent. If the contract changes, PHP can detect that the override no longer matches anything.

A constant can hold an object… that stays mutable

DBench86 also keeps its execution counters in an object referenced by a namespace constant:

const EXECUTION_METRICS = new ExecutionMetrics();
Enter fullscreen mode Exit fullscreen mode

The constant's binding never changes, but the object's properties can:

EXECUTION_METRICS->queries += 1;

if (!$executionFailed && !$isTimeout) {
    EXECUTION_METRICS->runsCompleted += 1;
}

if ($executionFailed) {
    EXECUTION_METRICS->failures += 1;
}

if ($isTimeout) {
    EXECUTION_METRICS->timeouts += 1;
}
Enter fullscreen mode Exit fullscreen mode

The distinction is a neat one: the binding stays constant; the object it refers to hasn't thereby become immutable.

Default values for readonly properties

A driver's identity never changes over its lifetime. For MariaDB, it can therefore be declared right where it belongs:

class MariaDbDriver extends AbstractDriver
{
    public readonly string $name = 'MariaDB';
    public readonly string $pdoDriver = 'mysql';
}
Enter fullscreen mode Exit fullscreen mode

Same approach for PostgreSQL and SQLite.

No need for a constructor just to assign two fixed values. The immutability of these properties stays explicit, and their value is visible right in the driver's declaration.

SortDirection: no more carrying asc and desc around as plain strings

Finally, the sort direction supplied on the command line is converted immediately into the new global SortDirection enum:

private function validateDirection(string $value): SortDirection
{
    return match ($value) {
        'asc' => SortDirection::Ascending,
        'desc' => SortDirection::Descending,
        default => throw new InvalidArgumentException(
            "Invalid value for --direction: {$value}. Accepted values: asc, desc"
        ),
    };
}
Enter fullscreen mode Exit fullscreen mode

Further down, result display then works against a typed value:

$comparison = $a['statistics'][$sort] <=> $b['statistics'][$sort];

return $direction === SortDirection::Ascending
    ? $comparison
    : -$comparison;
Enter fullscreen mode Exit fullscreen mode

Nothing flashy about it. Which is exactly what I like: the new feature almost disappears behind the problem it solves.

An optional feature that earned its place: parameter DocComments

The challenge also offered a second list of PHP 8.6 features, this one optional.

I didn't try to use all of it.

Adding a feature just to be able to say it's there wouldn't have made much sense — and some of them simply had no relevance to DBench86.

Parameter DocComments, though, had a genuine use case.

DBench86 generates its command-line help from its available parameters. The descriptions placed directly on those parameters can be retrieved through reflection and used to build that help text.

So they don't just document the source code — they become part of the CLI interface's description itself.

That's exactly the bar I'd set for myself with PHP 8.6's new features: don't go looking for somewhere to put them, use them when they genuinely add something to the application.

Among the challenge's optional features, this one passed the test.

And how do you get all this checked in a multi-file application?

One particular constraint of the challenge remained.

The official checker takes in a single PHP file and tokenizes it. But DBench86 is a real little application, spread across several classes and files. The entry point, dbench86.php, obviously doesn't contain all seven features itself — it loads the bootstrap and launches the application.

I had zero interest in turning the source code into a monolith just for the challenge.

So DBench86 builds its own standalone version:

if (in_array('--build-monolith', array_slice($argv, 1), true)) {
    if (count($argv) !== 2) {
        fwrite(STDERR, "Error: --build-monolith must be used alone.\n");
        return 1;
    }

    (new MonolithBuilder($projectDirectory))
        ->write($projectDirectory . '/dist/dbench86-monolith.php');

    echo "Built dist/dbench86-monolith.php\n";
    return 0;
}
Enter fullscreen mode Exit fullscreen mode

The builder walks the canonical source tree, assembles it in a deterministic order, and appends the application's actual entry point:

foreach ($this->sourceFiles() as $relative) {
    $output .= "\n// Source: {$relative}\n"
        . $this->namespaceBlock(
            file_get_contents($this->projectDirectory . '/' . $relative)
        );
}
Enter fullscreen mode Exit fullscreen mode

The result isn't a second implementation written specifically for the challenge. It's the same application, assembled into a self-contained PHP file the checker can analyze.

It started out as a constraint of the challenge. It ended up as a feature of the application.

I quite like that kind of detour.

The challenge stays open until November 19, 2026, the official PHP 8.6 release date. If this kind of exercise appeals to you — bringing new language features to life in a real application instead of seven isolated files — the exakat.io article explains the rules, and the challenge repository is open for entries.

This benchmark isn't an answer

DBench86 is now on GitHub, and its pull request is open for the PHPenomenal 8.6 challenge.

But I don't treat the numbers in this article as gospel.

They describe this machine, these versions, these configurations, this dataset, and these workloads.

On your machine, with your storage, your amount of RAM, your data, and above all your queries, you'll get something else.

And that's exactly how it should be.

The whole point of the program is that it can be run again.

A useful benchmark shouldn't tell you which product to pick — it should help you ask the right questions before you pick one.

I started out just wanting to take part in a PHP 8.6 challenge with something fun to build.

In the end, I took part in the challenge, built a tool that supports comparisons that actually mean something, and got to share the whole process with you.

Not a bad outcome.

Top comments (0)