DEV Community

Nicolas VANDENBOGAERDE
Nicolas VANDENBOGAERDE

Posted on

From RESP to Background Jobs: Building a Ractor-Native Ruby 4 Stack

From RESP to Background Jobs: Building a Ractor-Native Ruby 4 Stack

Ruby has traditionally scaled CPU-bound workloads by adding processes.

That model works extremely well, but it comes with a cost: every additional process has its own Ruby runtime, heap, loaded application code, connection pools, and associated memory overhead.

With Ruby 4 and Ractors, another architecture becomes increasingly interesting:

What if a single Ruby process could efficiently use multiple CPU cores while keeping mutable state isolated?

Over the last few months, I have been experimenting with this idea at several layers of the stack.

That work resulted in three gems:

SolidRESPRactor → SolidRedis → SolidJobs

They form a Ractor-oriented stack going from the Redis serialization protocol all the way to reliable background job execution.

The goal is not simply to make another Redis client or another job processor.

The goal is to explore what Ruby infrastructure can look like when Ractor is treated as an architectural primitive rather than something added afterwards.


The architecture

The stack can be represented simply:

SolidJobs
│
├── Supervisor
├── Scheduler
└── Worker Ractors
│
▼
SolidRedis
│
▼
SolidRESPRactor
│
▼
Redis

The fundamental rule throughout the stack is:

«Share immutable configuration. Keep mutable runtime state local to its owning Ractor.»

Instead of attempting to make complex mutable objects globally shareable, each Ractor owns the state it needs.

That applies to Redis connections, pools, RESP readers, buffers, cluster state and worker execution state.

This considerably simplifies the ownership model.


Layer 1: SolidRESPRactor

Everything starts with RESP.

A Ractor-native Redis client needs a protocol implementation whose mutable parsing and I/O state does not accidentally cross Ractor boundaries.

That became SolidRESPRactor.

The architecture separates shareable protocol configuration from mutable Reader state.

Conceptually:

Shareable
────────────
Codec
Encoder
Limits

Ractor-local
────────────
Reader
Buffer
Cursor
Socket / IO

This gives every Ractor ownership of its parsing state while allowing immutable protocol components to be shared.

But architecture alone doesn't tell us whether something is efficient.

So I built a dedicated benchmark suite.


Benchmark first, optimize second

The current Reader benchmark was run with:

Ruby: 4.0.1
Architecture: arm64-darwin25
Redis: 8.10.0
Ractors: 1 / 2 / 4 / 8
Warmup: 1 second
Measurement: 3 seconds
Runs: 3
Result: median

Reader-only benchmarks consume an in-memory repeating 16 KiB source.

Reader + TCP benchmarks use an isolated loopback Redis server.

Allocation measurements are performed separately in one Ractor with GC disabled.

That last detail matters.

A process-wide allocation counter such as "GC.stat(:total_allocated_objects)" cannot cleanly attribute concurrent allocations to individual Ractors.

So throughput/scaling and intrinsic allocation measurements are deliberately separated.

That decision led to an interesting discovery.


The parser wasn't the bottleneck I expected

A small RESP bulk value decoded directly by the Reader required:

bulk 16 B
787,426 ops/s
3 allocations/op
120.8 bytes/op

But when measuring a Redis GET through the complete TCP path:

GET over TCP
40,447 ops/s
6 allocations/op
16,609.8 bytes/op

That difference immediately looked suspicious.

The Reader itself wasn't allocating 16 KiB.

The socket-to-buffer path was.

The non-blocking socket read path was creating a chunk-sized String even when the Redis response was tiny.

Instead of adding more parser fast paths, I changed the I/O path so the read buffer could be reused.

Then I ran exactly the same benchmark again.


16,609 bytes → 120 bytes per GET

The result was much larger than I expected.

Metric| Before| After| Change
GET bytes/op| 16,609.8| 120.8| -99.27%
GET allocations/op| 6.0| 4.0| -33.3%
GET 1R| 40,447 ops/s| 40,173 ops/s| -0.68%
GET 2R| 62,692 ops/s| 66,235 ops/s| +5.65%
GET 4R| 83,028 ops/s| 89,146 ops/s| +7.37%
GET 8R| 94,345 ops/s| 102,060 ops/s| +8.18%

The 1-Ractor result remained effectively neutral while throughput improved as concurrency increased.

Scaling efficiency at eight Ractors increased from:

29.2% → 31.8%

The same optimization also affected pipelines.

Pipeline metric| Before| After| Change
bytes/op| 472.3| 120.0| -74.59%
allocations/op| 3.1| 3.0| -3.2%
throughput @ 8R| 1,999,445| 2,044,122 ops/s| +2.23%

This is why I prefer profiling the complete path before optimizing code that merely looks expensive.

The obvious candidate was the RESP parser.

The real opportunity was one layer below it.


Layer 2: SolidRedis

Once the RESP layer had explicit Ractor ownership, I could build a Redis client around the same principle.

That became SolidRedis.

The central architectural distinction is again between configuration and runtime state.

For example:

SentinelConfig
│
│ shareable
▼
Ractor A ──> SentinelState A
Ractor B ──> SentinelState B
Ractor C ──> SentinelState C
Ractor D ──> SentinelState D

The same principle applies to connection pools and cluster state.

A pool belongs to its Ractor.

Mutable connection state isn't silently shared between Ractors.

This means the architecture follows the actor model rather than attempting to put locks around an object graph originally designed for threads.

SolidRedis supports the Redis functionality required by the next layer of the experiment:

background jobs.


Layer 3: SolidJobs

SolidJobs asks a simple question:

«What happens if a Ruby background job processor is designed around Ractors from day one?»

A simplified configuration can look conceptually like:

SolidJobs.configure do |config|
config.ractors = 8
end

Worker Ractors can execute Ruby code in parallel across CPU cores.

Threads can still be useful inside a Ractor for I/O-heavy work.

So the model is not necessarily:

Ractors OR Threads

It can be:

Ractors
├── threads
├── threads
└── threads

Ractors provide CPU parallelism and ownership boundaries.

Threads provide inexpensive concurrency for I/O.


Reliability matters more than raw throughput

A fast job processor that loses jobs isn't particularly useful.

SolidJobs therefore uses an at-least-once delivery model.

The lifecycle is conceptually:

READY
│
▼
atomic reserve
│
▼
RESERVED
│
▼
perform
│
├──── failure ────> RETRY / FAILED
│
▼
ACK
│
▼
DONE

The complete path includes:

reserve
↓
journal
↓
perform
↓
fenced ACK

If the process dies after reservation, the job can be recovered.

If execution succeeds but acknowledgement fails because Redis becomes unavailable, the system does not pretend the job never executed.

The reservation remains recoverable.

This deliberately means at-least-once, not exactly-once.

Jobs therefore still need to be idempotent.

Reliability tests cover situations such as process crashes, ambiguous acknowledgements, recovery, Redis failures and worker failures.

Performance measurements include this reliable execution path rather than benchmarking only the call to the job method.


Ruby 4 changes the multicore picture

I then benchmarked SolidJobs on Ruby 4.0.1 with a CPU-bound workload.

The interesting result isn't simply its absolute throughput.

It's the scaling curve.

SolidJobs currently scales as follows:

Ractors| SolidJobs| Scaling efficiency
1| 282 jobs/s| 100.0%
2| 550 jobs/s| 97.7%
4| 1,099 jobs/s| 97.6%
8| 2,054 jobs/s| 91.2%

From one to eight Ractors:

282 → 550 → 1,099 → 2,054 jobs/s

That's approximately 7.28× throughput from 8× concurrency.

For this workload, Ruby 4 is clearly executing useful Ruby work across the available cores.


Comparing with Sidekiq

I also measured Sidekiq under the same CPU-bound benchmark.

At eight-way concurrency:

Metric| Sidekiq| SolidJobs
Throughput| 298 jobs/s| 2,054 jobs/s
CPU-s / 1,000 jobs| 3.36| 3.69
Peak RSS| 42.4 MiB| 42.2 MiB
Allocations/job| 108.1| 68.7

The throughput ratio in this particular benchmark is approximately:

6.89×

But that number needs an important qualification.

This is not a claim that SolidJobs is universally “6.9× faster than Sidekiq.”

The workload is CPU-bound.

A single Sidekiq process is fundamentally being compared with a SolidJobs process capable of executing Ruby work across multiple Ractors and therefore multiple CPU cores.

That is precisely the property being investigated.

For I/O-heavy applications, extremely short jobs, different hardware, multiple Sidekiq processes or different deployment models, the result can be very different.

The next useful comparison is therefore a fixed machine / fixed CPU budget benchmark comparing one multi-Ractor SolidJobs process against enough Sidekiq processes to use the same cores.

That is a much more interesting infrastructure comparison than a headline throughput number.


Something else changed: memory

An earlier SolidJobs baseline had significantly higher RSS:

8 Ractors:
74.1 MiB

After improvements in the underlying RESP/TCP path, the current measurement is:

8 Ractors:
42.2 MiB

The previous and current SolidJobs results are:

Metric @ 8R| Previous| Current
Throughput| 2,053 jobs/s| 2,054 jobs/s
CPU-s/1k jobs| 3.67| 3.69
Peak RSS| 74.1 MiB| 42.2 MiB
Allocations/job| 74.3| 68.7
jobs/s/MiB| 27.71| 48.67

Throughput remained essentially unchanged.

The interesting improvement happened elsewhere:

less memory and fewer allocations for the same amount of useful work.

The large RSS reduction deserves continued A/B validation before attributing all of it to a single RESP optimization, but the reduction in allocations is consistent across the Ractor matrix.

This is also a good example of why infrastructure optimization should not focus only on requests or jobs per second.


Why Ractors could reduce infrastructure costs

This is where the experiment becomes particularly interesting to me.

The traditional way to exploit eight CPU cores with CPU-bound Ruby workloads is generally to run multiple processes.

Conceptually:

8 CPU cores

Traditional process scaling

Process 1 ── Ruby VM ── heap ── connections
Process 2 ── Ruby VM ── heap ── connections
Process 3 ── Ruby VM ── heap ── connections
Process 4 ── Ruby VM ── heap ── connections
...

A Ractor-oriented architecture offers another model:

One Ruby process
│
├── Ractor 1
├── Ractor 2
├── Ractor 3
├── Ractor 4
├── Ractor 5
├── Ractor 6
├── Ractor 7
└── Ractor 8

That doesn't make infrastructure free.

And it certainly doesn't mean that every application should replace process-based scaling with Ractors.

But it creates an interesting possibility:

«more useful CPU work from each Ruby process.»

If the same workload can be handled with fewer Ruby processes, that can potentially mean:

  • fewer duplicated Ruby heaps;
  • fewer duplicated application runtimes;
  • fewer worker processes;
  • fewer Redis connection pools;
  • lower memory requirements;
  • better utilization of multicore instances;
  • potentially fewer containers or VMs for the same workload.

And ultimately:

lower infrastructure cost per completed job.

That last metric is the one I ultimately want to measure.

Not:

Which library wins a microbenchmark?

But:

How many dollars does it cost
to reliably execute N jobs?


Throughput per dollar is more interesting than throughput alone

Imagine two architectures that both need to process the same number of CPU-heavy jobs.

Architecture A requires multiple Ruby processes across several containers.

Architecture B can exploit the available cores efficiently inside fewer processes.

Even if both eventually achieve the same throughput, B may require:

less RAM
+
fewer processes
+
fewer containers
+
less orchestration overhead

That potentially changes cloud economics.

But this needs to be measured, not assumed.

The next stage of the SolidJobs benchmarks will therefore need to compare:

Same hardware
Same number of CPU cores
Same Redis
Same workload
Same reliability guarantees

Then compare:

SolidJobs:
1 process × N Ractors

versus

Sidekiq:
N processes / equivalent CPU utilization

And measure:

jobs/s
CPU-seconds/job
RSS
allocations/job
Redis connections
latency
recovery behavior
and ultimately cost per million jobs

Only then can infrastructure savings be quantified properly.


Ractors aren't a free performance button

There are important trade-offs.

Ractor-oriented code requires much stricter thinking about ownership.

Objects cannot simply be shared everywhere.

Mutable state has to be deliberately located.

Libraries in the dependency graph need to behave correctly.

And Ractor behavior itself has evolved substantially between Ruby versions.

During SolidJobs development I encountered a particularly interesting difference between Ruby 3.4.4 and Ruby 4.0.1.

A concurrent TCP + Reader reproducer using eight Ractors could trigger a native SIGTRAP under Ruby 3.4.4.

The same concurrent reproducer completed 100/100 runs under Ruby 4.0.1.

That experience resulted in additional startup isolation and compatibility work in SolidJobs.

It also reinforced an important lesson:

Ractor benchmarks must include the complete runtime path, not just isolated CPU loops.


What I learned

Building these three gems changed how I think about Ruby performance.

The biggest optimization in the RESP layer did not come from making parsing cleverer.

It came from measuring allocations and discovering an unnecessary 16 KiB I/O allocation.

The biggest SolidJobs result isn't simply that eight Ractors are faster than one.

It's that Ruby 4 can sustain:

1R: 282 jobs/s
2R: 550 jobs/s
4R: 1,099 jobs/s
8R: 2,054 jobs/s

while retaining 91.2% scaling efficiency at eight Ractors on this workload.

And the most interesting question now isn't whether Ractors can run Ruby code in parallel.

They can.

The interesting question is:

«Can Ractor-native infrastructure let us run Ruby applications with better CPU and memory density — and therefore reduce the infrastructure required for a given workload?»

SolidRESPRactor, SolidRedis and SolidJobs are my attempt to explore that question from the protocol layer upward.


Try it and reproduce the benchmarks

All three projects are open source:

SolidRESPRactor
https://github.com/nicolasva/solid-resp-ractor

SolidRedis
https://github.com/nicolasva/solid-redis

SolidJobs
https://github.com/nicolasva/solid-jobs

RubyGems:

https://rubygems.org/gems/solid-resp-ractor

https://rubygems.org/gems/solid-redis

https://rubygems.org/gems/solid-jobs

The benchmark methodology and results are published with the projects.

What would be particularly useful now is independent reproduction.

If you're running Ruby 4 on Linux x86-64, Linux ARM64, Apple Silicon or a large multicore server, I'd love to see how the Ractor scaling behaves on your machine.

Positive or negative results are equally useful.

Because the goal isn't to prove that Ractors always win.

The goal is to find out where this architecture actually makes sense.


SolidRESPRactor, SolidRedis and SolidJobs are experimental open-source projects. Benchmark results above describe the specified Ruby 4.0.1 environment and workload; they should not be interpreted as universal production-performance claims.

Top comments (0)