<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Nicolas VANDENBOGAERDE</title>
    <description>The latest articles on DEV Community by Nicolas VANDENBOGAERDE (@nicolasva).</description>
    <link>https://dev.to/nicolasva</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4156523%2F473592fa-c58d-4547-89f8-c99d06e0225f.jpg</url>
      <title>DEV Community: Nicolas VANDENBOGAERDE</title>
      <link>https://dev.to/nicolasva</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/nicolasva"/>
    <language>en</language>
    <item>
      <title>From RESP to Background Jobs: Building a Ractor-Native Ruby 4 Stack</title>
      <dc:creator>Nicolas VANDENBOGAERDE</dc:creator>
      <pubDate>Fri, 02 Oct 2026 06:23:46 +0000</pubDate>
      <link>https://dev.to/nicolasva/from-resp-to-background-jobs-building-a-ractor-native-ruby-4-stack-51p5</link>
      <guid>https://dev.to/nicolasva/from-resp-to-background-jobs-building-a-ractor-native-ruby-4-stack-51p5</guid>
      <description>&lt;p&gt;From RESP to Background Jobs: Building a Ractor-Native Ruby 4 Stack&lt;/p&gt;

&lt;p&gt;Ruby has traditionally scaled CPU-bound workloads by adding processes.&lt;/p&gt;

&lt;p&gt;That model works extremely well, but it comes with a cost: every additional process has its own Ruby runtime, heap, loaded application code, connection pools, and associated memory overhead.&lt;/p&gt;

&lt;p&gt;With Ruby 4 and Ractors, another architecture becomes increasingly interesting:&lt;/p&gt;

&lt;p&gt;What if a single Ruby process could efficiently use multiple CPU cores while keeping mutable state isolated?&lt;/p&gt;

&lt;p&gt;Over the last few months, I have been experimenting with this idea at several layers of the stack.&lt;/p&gt;

&lt;p&gt;That work resulted in three gems:&lt;/p&gt;

&lt;p&gt;SolidRESPRactor → SolidRedis → SolidJobs&lt;/p&gt;

&lt;p&gt;They form a Ractor-oriented stack going from the Redis serialization protocol all the way to reliable background job execution.&lt;/p&gt;

&lt;p&gt;The goal is not simply to make another Redis client or another job processor.&lt;/p&gt;

&lt;p&gt;The goal is to explore what Ruby infrastructure can look like when Ractor is treated as an architectural primitive rather than something added afterwards.&lt;/p&gt;




&lt;p&gt;The architecture&lt;/p&gt;

&lt;p&gt;The stack can be represented simply:&lt;/p&gt;

&lt;p&gt;SolidJobs&lt;br&gt;
    │&lt;br&gt;
    ├── Supervisor&lt;br&gt;
    ├── Scheduler&lt;br&gt;
    └── Worker Ractors&lt;br&gt;
             │&lt;br&gt;
             ▼&lt;br&gt;
         SolidRedis&lt;br&gt;
             │&lt;br&gt;
             ▼&lt;br&gt;
      SolidRESPRactor&lt;br&gt;
             │&lt;br&gt;
             ▼&lt;br&gt;
           Redis&lt;/p&gt;

&lt;p&gt;The fundamental rule throughout the stack is:&lt;/p&gt;

&lt;p&gt;«Share immutable configuration. Keep mutable runtime state local to its owning Ractor.»&lt;/p&gt;

&lt;p&gt;Instead of attempting to make complex mutable objects globally shareable, each Ractor owns the state it needs.&lt;/p&gt;

&lt;p&gt;That applies to Redis connections, pools, RESP readers, buffers, cluster state and worker execution state.&lt;/p&gt;

&lt;p&gt;This considerably simplifies the ownership model.&lt;/p&gt;




&lt;p&gt;Layer 1: SolidRESPRactor&lt;/p&gt;

&lt;p&gt;Everything starts with RESP.&lt;/p&gt;

&lt;p&gt;A Ractor-native Redis client needs a protocol implementation whose mutable parsing and I/O state does not accidentally cross Ractor boundaries.&lt;/p&gt;

&lt;p&gt;That became SolidRESPRactor.&lt;/p&gt;

&lt;p&gt;The architecture separates shareable protocol configuration from mutable Reader state.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;/p&gt;

&lt;p&gt;Shareable&lt;br&gt;
────────────&lt;br&gt;
Codec&lt;br&gt;
Encoder&lt;br&gt;
Limits&lt;/p&gt;

&lt;p&gt;Ractor-local&lt;br&gt;
────────────&lt;br&gt;
Reader&lt;br&gt;
Buffer&lt;br&gt;
Cursor&lt;br&gt;
Socket / IO&lt;/p&gt;

&lt;p&gt;This gives every Ractor ownership of its parsing state while allowing immutable protocol components to be shared.&lt;/p&gt;

&lt;p&gt;But architecture alone doesn't tell us whether something is efficient.&lt;/p&gt;

&lt;p&gt;So I built a dedicated benchmark suite.&lt;/p&gt;




&lt;p&gt;Benchmark first, optimize second&lt;/p&gt;

&lt;p&gt;The current Reader benchmark was run with:&lt;/p&gt;

&lt;p&gt;Ruby:        4.0.1&lt;br&gt;
Architecture: arm64-darwin25&lt;br&gt;
Redis:       8.10.0&lt;br&gt;
Ractors:     1 / 2 / 4 / 8&lt;br&gt;
Warmup:      1 second&lt;br&gt;
Measurement: 3 seconds&lt;br&gt;
Runs:        3&lt;br&gt;
Result:      median&lt;/p&gt;

&lt;p&gt;Reader-only benchmarks consume an in-memory repeating 16 KiB source.&lt;/p&gt;

&lt;p&gt;Reader + TCP benchmarks use an isolated loopback Redis server.&lt;/p&gt;

&lt;p&gt;Allocation measurements are performed separately in one Ractor with GC disabled.&lt;/p&gt;

&lt;p&gt;That last detail matters.&lt;/p&gt;

&lt;p&gt;A process-wide allocation counter such as "GC.stat(:total_allocated_objects)" cannot cleanly attribute concurrent allocations to individual Ractors.&lt;/p&gt;

&lt;p&gt;So throughput/scaling and intrinsic allocation measurements are deliberately separated.&lt;/p&gt;

&lt;p&gt;That decision led to an interesting discovery.&lt;/p&gt;




&lt;p&gt;The parser wasn't the bottleneck I expected&lt;/p&gt;

&lt;p&gt;A small RESP bulk value decoded directly by the Reader required:&lt;/p&gt;

&lt;p&gt;bulk 16 B&lt;br&gt;
787,426 ops/s&lt;br&gt;
3 allocations/op&lt;br&gt;
120.8 bytes/op&lt;/p&gt;

&lt;p&gt;But when measuring a Redis GET through the complete TCP path:&lt;/p&gt;

&lt;p&gt;GET over TCP&lt;br&gt;
40,447 ops/s&lt;br&gt;
6 allocations/op&lt;br&gt;
16,609.8 bytes/op&lt;/p&gt;

&lt;p&gt;That difference immediately looked suspicious.&lt;/p&gt;

&lt;p&gt;The Reader itself wasn't allocating 16 KiB.&lt;/p&gt;

&lt;p&gt;The socket-to-buffer path was.&lt;/p&gt;

&lt;p&gt;The non-blocking socket read path was creating a chunk-sized String even when the Redis response was tiny.&lt;/p&gt;

&lt;p&gt;Instead of adding more parser fast paths, I changed the I/O path so the read buffer could be reused.&lt;/p&gt;

&lt;p&gt;Then I ran exactly the same benchmark again.&lt;/p&gt;




&lt;p&gt;16,609 bytes → 120 bytes per GET&lt;/p&gt;

&lt;p&gt;The result was much larger than I expected.&lt;/p&gt;

&lt;p&gt;Metric| Before| After| Change&lt;br&gt;
GET bytes/op| 16,609.8| 120.8| -99.27%&lt;br&gt;
GET allocations/op| 6.0| 4.0| -33.3%&lt;br&gt;
GET 1R| 40,447 ops/s| 40,173 ops/s| -0.68%&lt;br&gt;
GET 2R| 62,692 ops/s| 66,235 ops/s| +5.65%&lt;br&gt;
GET 4R| 83,028 ops/s| 89,146 ops/s| +7.37%&lt;br&gt;
GET 8R| 94,345 ops/s| 102,060 ops/s| +8.18%&lt;/p&gt;

&lt;p&gt;The 1-Ractor result remained effectively neutral while throughput improved as concurrency increased.&lt;/p&gt;

&lt;p&gt;Scaling efficiency at eight Ractors increased from:&lt;/p&gt;

&lt;p&gt;29.2% → 31.8%&lt;/p&gt;

&lt;p&gt;The same optimization also affected pipelines.&lt;/p&gt;

&lt;p&gt;Pipeline metric| Before| After| Change&lt;br&gt;
bytes/op| 472.3| 120.0| -74.59%&lt;br&gt;
allocations/op| 3.1| 3.0| -3.2%&lt;br&gt;
throughput @ 8R| 1,999,445| 2,044,122 ops/s| +2.23%&lt;/p&gt;

&lt;p&gt;This is why I prefer profiling the complete path before optimizing code that merely looks expensive.&lt;/p&gt;

&lt;p&gt;The obvious candidate was the RESP parser.&lt;/p&gt;

&lt;p&gt;The real opportunity was one layer below it.&lt;/p&gt;




&lt;p&gt;Layer 2: SolidRedis&lt;/p&gt;

&lt;p&gt;Once the RESP layer had explicit Ractor ownership, I could build a Redis client around the same principle.&lt;/p&gt;

&lt;p&gt;That became SolidRedis.&lt;/p&gt;

&lt;p&gt;The central architectural distinction is again between configuration and runtime state.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;SentinelConfig&lt;br&gt;
     │&lt;br&gt;
     │ shareable&lt;br&gt;
     ▼&lt;br&gt;
Ractor A ──&amp;gt; SentinelState A&lt;br&gt;
Ractor B ──&amp;gt; SentinelState B&lt;br&gt;
Ractor C ──&amp;gt; SentinelState C&lt;br&gt;
Ractor D ──&amp;gt; SentinelState D&lt;/p&gt;

&lt;p&gt;The same principle applies to connection pools and cluster state.&lt;/p&gt;

&lt;p&gt;A pool belongs to its Ractor.&lt;/p&gt;

&lt;p&gt;Mutable connection state isn't silently shared between Ractors.&lt;/p&gt;

&lt;p&gt;This means the architecture follows the actor model rather than attempting to put locks around an object graph originally designed for threads.&lt;/p&gt;

&lt;p&gt;SolidRedis supports the Redis functionality required by the next layer of the experiment:&lt;/p&gt;

&lt;p&gt;background jobs.&lt;/p&gt;




&lt;p&gt;Layer 3: SolidJobs&lt;/p&gt;

&lt;p&gt;SolidJobs asks a simple question:&lt;/p&gt;

&lt;p&gt;«What happens if a Ruby background job processor is designed around Ractors from day one?»&lt;/p&gt;

&lt;p&gt;A simplified configuration can look conceptually like:&lt;/p&gt;

&lt;p&gt;SolidJobs.configure do |config|&lt;br&gt;
  config.ractors = 8&lt;br&gt;
end&lt;/p&gt;

&lt;p&gt;Worker Ractors can execute Ruby code in parallel across CPU cores.&lt;/p&gt;

&lt;p&gt;Threads can still be useful inside a Ractor for I/O-heavy work.&lt;/p&gt;

&lt;p&gt;So the model is not necessarily:&lt;/p&gt;

&lt;p&gt;Ractors OR Threads&lt;/p&gt;

&lt;p&gt;It can be:&lt;/p&gt;

&lt;p&gt;Ractors&lt;br&gt;
   ├── threads&lt;br&gt;
   ├── threads&lt;br&gt;
   └── threads&lt;/p&gt;

&lt;p&gt;Ractors provide CPU parallelism and ownership boundaries.&lt;/p&gt;

&lt;p&gt;Threads provide inexpensive concurrency for I/O.&lt;/p&gt;




&lt;p&gt;Reliability matters more than raw throughput&lt;/p&gt;

&lt;p&gt;A fast job processor that loses jobs isn't particularly useful.&lt;/p&gt;

&lt;p&gt;SolidJobs therefore uses an at-least-once delivery model.&lt;/p&gt;

&lt;p&gt;The lifecycle is conceptually:&lt;/p&gt;

&lt;p&gt;READY&lt;br&gt;
  │&lt;br&gt;
  ▼&lt;br&gt;
atomic reserve&lt;br&gt;
  │&lt;br&gt;
  ▼&lt;br&gt;
RESERVED&lt;br&gt;
  │&lt;br&gt;
  ▼&lt;br&gt;
perform&lt;br&gt;
  │&lt;br&gt;
  ├──── failure ────&amp;gt; RETRY / FAILED&lt;br&gt;
  │&lt;br&gt;
  ▼&lt;br&gt;
ACK&lt;br&gt;
  │&lt;br&gt;
  ▼&lt;br&gt;
DONE&lt;/p&gt;

&lt;p&gt;The complete path includes:&lt;/p&gt;

&lt;p&gt;reserve&lt;br&gt;
   ↓&lt;br&gt;
journal&lt;br&gt;
   ↓&lt;br&gt;
perform&lt;br&gt;
   ↓&lt;br&gt;
fenced ACK&lt;/p&gt;

&lt;p&gt;If the process dies after reservation, the job can be recovered.&lt;/p&gt;

&lt;p&gt;If execution succeeds but acknowledgement fails because Redis becomes unavailable, the system does not pretend the job never executed.&lt;/p&gt;

&lt;p&gt;The reservation remains recoverable.&lt;/p&gt;

&lt;p&gt;This deliberately means at-least-once, not exactly-once.&lt;/p&gt;

&lt;p&gt;Jobs therefore still need to be idempotent.&lt;/p&gt;

&lt;p&gt;Reliability tests cover situations such as process crashes, ambiguous acknowledgements, recovery, Redis failures and worker failures.&lt;/p&gt;

&lt;p&gt;Performance measurements include this reliable execution path rather than benchmarking only the call to the job method.&lt;/p&gt;




&lt;p&gt;Ruby 4 changes the multicore picture&lt;/p&gt;

&lt;p&gt;I then benchmarked SolidJobs on Ruby 4.0.1 with a CPU-bound workload.&lt;/p&gt;

&lt;p&gt;The interesting result isn't simply its absolute throughput.&lt;/p&gt;

&lt;p&gt;It's the scaling curve.&lt;/p&gt;

&lt;p&gt;SolidJobs currently scales as follows:&lt;/p&gt;

&lt;p&gt;Ractors| SolidJobs| Scaling efficiency&lt;br&gt;
1| 282 jobs/s| 100.0%&lt;br&gt;
2| 550 jobs/s| 97.7%&lt;br&gt;
4| 1,099 jobs/s| 97.6%&lt;br&gt;
8| 2,054 jobs/s| 91.2%&lt;/p&gt;

&lt;p&gt;From one to eight Ractors:&lt;/p&gt;

&lt;p&gt;282 → 550 → 1,099 → 2,054 jobs/s&lt;/p&gt;

&lt;p&gt;That's approximately 7.28× throughput from 8× concurrency.&lt;/p&gt;

&lt;p&gt;For this workload, Ruby 4 is clearly executing useful Ruby work across the available cores.&lt;/p&gt;




&lt;p&gt;Comparing with Sidekiq&lt;/p&gt;

&lt;p&gt;I also measured Sidekiq under the same CPU-bound benchmark.&lt;/p&gt;

&lt;p&gt;At eight-way concurrency:&lt;/p&gt;

&lt;p&gt;Metric| Sidekiq| SolidJobs&lt;br&gt;
Throughput| 298 jobs/s| 2,054 jobs/s&lt;br&gt;
CPU-s / 1,000 jobs| 3.36| 3.69&lt;br&gt;
Peak RSS| 42.4 MiB| 42.2 MiB&lt;br&gt;
Allocations/job| 108.1| 68.7&lt;/p&gt;

&lt;p&gt;The throughput ratio in this particular benchmark is approximately:&lt;/p&gt;

&lt;p&gt;6.89×&lt;/p&gt;

&lt;p&gt;But that number needs an important qualification.&lt;/p&gt;

&lt;p&gt;This is not a claim that SolidJobs is universally “6.9× faster than Sidekiq.”&lt;/p&gt;

&lt;p&gt;The workload is CPU-bound.&lt;/p&gt;

&lt;p&gt;A single Sidekiq process is fundamentally being compared with a SolidJobs process capable of executing Ruby work across multiple Ractors and therefore multiple CPU cores.&lt;/p&gt;

&lt;p&gt;That is precisely the property being investigated.&lt;/p&gt;

&lt;p&gt;For I/O-heavy applications, extremely short jobs, different hardware, multiple Sidekiq processes or different deployment models, the result can be very different.&lt;/p&gt;

&lt;p&gt;The next useful comparison is therefore a fixed machine / fixed CPU budget benchmark comparing one multi-Ractor SolidJobs process against enough Sidekiq processes to use the same cores.&lt;/p&gt;

&lt;p&gt;That is a much more interesting infrastructure comparison than a headline throughput number.&lt;/p&gt;




&lt;p&gt;Something else changed: memory&lt;/p&gt;

&lt;p&gt;An earlier SolidJobs baseline had significantly higher RSS:&lt;/p&gt;

&lt;p&gt;8 Ractors:&lt;br&gt;
74.1 MiB&lt;/p&gt;

&lt;p&gt;After improvements in the underlying RESP/TCP path, the current measurement is:&lt;/p&gt;

&lt;p&gt;8 Ractors:&lt;br&gt;
42.2 MiB&lt;/p&gt;

&lt;p&gt;The previous and current SolidJobs results are:&lt;/p&gt;

&lt;p&gt;Metric @ 8R| Previous| Current&lt;br&gt;
Throughput| 2,053 jobs/s| 2,054 jobs/s&lt;br&gt;
CPU-s/1k jobs| 3.67| 3.69&lt;br&gt;
Peak RSS| 74.1 MiB| 42.2 MiB&lt;br&gt;
Allocations/job| 74.3| 68.7&lt;br&gt;
jobs/s/MiB| 27.71| 48.67&lt;/p&gt;

&lt;p&gt;Throughput remained essentially unchanged.&lt;/p&gt;

&lt;p&gt;The interesting improvement happened elsewhere:&lt;/p&gt;

&lt;p&gt;less memory and fewer allocations for the same amount of useful work.&lt;/p&gt;

&lt;p&gt;The large RSS reduction deserves continued A/B validation before attributing all of it to a single RESP optimization, but the reduction in allocations is consistent across the Ractor matrix.&lt;/p&gt;

&lt;p&gt;This is also a good example of why infrastructure optimization should not focus only on requests or jobs per second.&lt;/p&gt;




&lt;p&gt;Why Ractors could reduce infrastructure costs&lt;/p&gt;

&lt;p&gt;This is where the experiment becomes particularly interesting to me.&lt;/p&gt;

&lt;p&gt;The traditional way to exploit eight CPU cores with CPU-bound Ruby workloads is generally to run multiple processes.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;/p&gt;

&lt;p&gt;8 CPU cores&lt;/p&gt;

&lt;p&gt;Traditional process scaling&lt;/p&gt;

&lt;p&gt;Process 1 ── Ruby VM ── heap ── connections&lt;br&gt;
Process 2 ── Ruby VM ── heap ── connections&lt;br&gt;
Process 3 ── Ruby VM ── heap ── connections&lt;br&gt;
Process 4 ── Ruby VM ── heap ── connections&lt;br&gt;
...&lt;/p&gt;

&lt;p&gt;A Ractor-oriented architecture offers another model:&lt;/p&gt;

&lt;p&gt;One Ruby process&lt;br&gt;
       │&lt;br&gt;
       ├── Ractor 1&lt;br&gt;
       ├── Ractor 2&lt;br&gt;
       ├── Ractor 3&lt;br&gt;
       ├── Ractor 4&lt;br&gt;
       ├── Ractor 5&lt;br&gt;
       ├── Ractor 6&lt;br&gt;
       ├── Ractor 7&lt;br&gt;
       └── Ractor 8&lt;/p&gt;

&lt;p&gt;That doesn't make infrastructure free.&lt;/p&gt;

&lt;p&gt;And it certainly doesn't mean that every application should replace process-based scaling with Ractors.&lt;/p&gt;

&lt;p&gt;But it creates an interesting possibility:&lt;/p&gt;

&lt;p&gt;«more useful CPU work from each Ruby process.»&lt;/p&gt;

&lt;p&gt;If the same workload can be handled with fewer Ruby processes, that can potentially mean:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;fewer duplicated Ruby heaps;&lt;/li&gt;
&lt;li&gt;fewer duplicated application runtimes;&lt;/li&gt;
&lt;li&gt;fewer worker processes;&lt;/li&gt;
&lt;li&gt;fewer Redis connection pools;&lt;/li&gt;
&lt;li&gt;lower memory requirements;&lt;/li&gt;
&lt;li&gt;better utilization of multicore instances;&lt;/li&gt;
&lt;li&gt;potentially fewer containers or VMs for the same workload.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And ultimately:&lt;/p&gt;

&lt;p&gt;lower infrastructure cost per completed job.&lt;/p&gt;

&lt;p&gt;That last metric is the one I ultimately want to measure.&lt;/p&gt;

&lt;p&gt;Not:&lt;/p&gt;

&lt;p&gt;Which library wins a microbenchmark?&lt;/p&gt;

&lt;p&gt;But:&lt;/p&gt;

&lt;p&gt;How many dollars does it cost&lt;br&gt;
to reliably execute N jobs?&lt;/p&gt;




&lt;p&gt;Throughput per dollar is more interesting than throughput alone&lt;/p&gt;

&lt;p&gt;Imagine two architectures that both need to process the same number of CPU-heavy jobs.&lt;/p&gt;

&lt;p&gt;Architecture A requires multiple Ruby processes across several containers.&lt;/p&gt;

&lt;p&gt;Architecture B can exploit the available cores efficiently inside fewer processes.&lt;/p&gt;

&lt;p&gt;Even if both eventually achieve the same throughput, B may require:&lt;/p&gt;

&lt;p&gt;less RAM&lt;br&gt;
+&lt;br&gt;
fewer processes&lt;br&gt;
+&lt;br&gt;
fewer containers&lt;br&gt;
+&lt;br&gt;
less orchestration overhead&lt;/p&gt;

&lt;p&gt;That potentially changes cloud economics.&lt;/p&gt;

&lt;p&gt;But this needs to be measured, not assumed.&lt;/p&gt;

&lt;p&gt;The next stage of the SolidJobs benchmarks will therefore need to compare:&lt;/p&gt;

&lt;p&gt;Same hardware&lt;br&gt;
Same number of CPU cores&lt;br&gt;
Same Redis&lt;br&gt;
Same workload&lt;br&gt;
Same reliability guarantees&lt;/p&gt;

&lt;p&gt;Then compare:&lt;/p&gt;

&lt;p&gt;SolidJobs:&lt;br&gt;
1 process × N Ractors&lt;/p&gt;

&lt;p&gt;versus&lt;/p&gt;

&lt;p&gt;Sidekiq:&lt;br&gt;
N processes / equivalent CPU utilization&lt;/p&gt;

&lt;p&gt;And measure:&lt;/p&gt;

&lt;p&gt;jobs/s&lt;br&gt;
CPU-seconds/job&lt;br&gt;
RSS&lt;br&gt;
allocations/job&lt;br&gt;
Redis connections&lt;br&gt;
latency&lt;br&gt;
recovery behavior&lt;br&gt;
and ultimately cost per million jobs&lt;/p&gt;

&lt;p&gt;Only then can infrastructure savings be quantified properly.&lt;/p&gt;




&lt;p&gt;Ractors aren't a free performance button&lt;/p&gt;

&lt;p&gt;There are important trade-offs.&lt;/p&gt;

&lt;p&gt;Ractor-oriented code requires much stricter thinking about ownership.&lt;/p&gt;

&lt;p&gt;Objects cannot simply be shared everywhere.&lt;/p&gt;

&lt;p&gt;Mutable state has to be deliberately located.&lt;/p&gt;

&lt;p&gt;Libraries in the dependency graph need to behave correctly.&lt;/p&gt;

&lt;p&gt;And Ractor behavior itself has evolved substantially between Ruby versions.&lt;/p&gt;

&lt;p&gt;During SolidJobs development I encountered a particularly interesting difference between Ruby 3.4.4 and Ruby 4.0.1.&lt;/p&gt;

&lt;p&gt;A concurrent TCP + Reader reproducer using eight Ractors could trigger a native SIGTRAP under Ruby 3.4.4.&lt;/p&gt;

&lt;p&gt;The same concurrent reproducer completed 100/100 runs under Ruby 4.0.1.&lt;/p&gt;

&lt;p&gt;That experience resulted in additional startup isolation and compatibility work in SolidJobs.&lt;/p&gt;

&lt;p&gt;It also reinforced an important lesson:&lt;/p&gt;

&lt;p&gt;Ractor benchmarks must include the complete runtime path, not just isolated CPU loops.&lt;/p&gt;




&lt;p&gt;What I learned&lt;/p&gt;

&lt;p&gt;Building these three gems changed how I think about Ruby performance.&lt;/p&gt;

&lt;p&gt;The biggest optimization in the RESP layer did not come from making parsing cleverer.&lt;/p&gt;

&lt;p&gt;It came from measuring allocations and discovering an unnecessary 16 KiB I/O allocation.&lt;/p&gt;

&lt;p&gt;The biggest SolidJobs result isn't simply that eight Ractors are faster than one.&lt;/p&gt;

&lt;p&gt;It's that Ruby 4 can sustain:&lt;/p&gt;

&lt;p&gt;1R:   282 jobs/s&lt;br&gt;
2R:   550 jobs/s&lt;br&gt;
4R: 1,099 jobs/s&lt;br&gt;
8R: 2,054 jobs/s&lt;/p&gt;

&lt;p&gt;while retaining 91.2% scaling efficiency at eight Ractors on this workload.&lt;/p&gt;

&lt;p&gt;And the most interesting question now isn't whether Ractors can run Ruby code in parallel.&lt;/p&gt;

&lt;p&gt;They can.&lt;/p&gt;

&lt;p&gt;The interesting question is:&lt;/p&gt;

&lt;p&gt;«Can Ractor-native infrastructure let us run Ruby applications with better CPU and memory density — and therefore reduce the infrastructure required for a given workload?»&lt;/p&gt;

&lt;p&gt;SolidRESPRactor, SolidRedis and SolidJobs are my attempt to explore that question from the protocol layer upward.&lt;/p&gt;




&lt;p&gt;Try it and reproduce the benchmarks&lt;/p&gt;

&lt;p&gt;All three projects are open source:&lt;/p&gt;

&lt;p&gt;SolidRESPRactor&lt;br&gt;
&lt;a href="https://github.com/nicolasva/solid-resp-ractor" rel="noopener noreferrer"&gt;https://github.com/nicolasva/solid-resp-ractor&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SolidRedis&lt;br&gt;
&lt;a href="https://github.com/nicolasva/solid-redis" rel="noopener noreferrer"&gt;https://github.com/nicolasva/solid-redis&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SolidJobs&lt;br&gt;
&lt;a href="https://github.com/nicolasva/solid-jobs" rel="noopener noreferrer"&gt;https://github.com/nicolasva/solid-jobs&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;RubyGems:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://rubygems.org/gems/solid-resp-ractor" rel="noopener noreferrer"&gt;https://rubygems.org/gems/solid-resp-ractor&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://rubygems.org/gems/solid-redis" rel="noopener noreferrer"&gt;https://rubygems.org/gems/solid-redis&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://rubygems.org/gems/solid-jobs" rel="noopener noreferrer"&gt;https://rubygems.org/gems/solid-jobs&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The benchmark methodology and results are published with the projects.&lt;/p&gt;

&lt;p&gt;What would be particularly useful now is independent reproduction.&lt;/p&gt;

&lt;p&gt;If you're running Ruby 4 on Linux x86-64, Linux ARM64, Apple Silicon or a large multicore server, I'd love to see how the Ractor scaling behaves on your machine.&lt;/p&gt;

&lt;p&gt;Positive or negative results are equally useful.&lt;/p&gt;

&lt;p&gt;Because the goal isn't to prove that Ractors always win.&lt;/p&gt;

&lt;p&gt;The goal is to find out where this architecture actually makes sense.&lt;/p&gt;




&lt;p&gt;SolidRESPRactor, SolidRedis and SolidJobs are experimental open-source projects. Benchmark results above describe the specified Ruby 4.0.1 environment and workload; they should not be interpreted as universal production-performance claims.&lt;/p&gt;

</description>
      <category>programming</category>
      <category>ruby</category>
      <category>rails</category>
    </item>
  </channel>
</rss>
