DEV Community

Cover image for I Built a Concurrent Resource Scheduler in Go Using Sharded Priority Heaps
Feroz
Feroz

Posted on

I Built a Concurrent Resource Scheduler in Go Using Sharded Priority Heaps

Support on GitHub: github.com/phero20/concurrent-resource-scheduler (Give it a star if you find it useful!)

View Docs: pkg.go.dev/github.com/phero20/concurrent-resource-scheduler


What happens when thousands of concurrent requests compete for a small pool of reusable resources?

You can put a mutex around a slice and hope for the best.

Or you can design the scheduler around concurrency from the beginning.

I chose the second option.

I built Concurrent Resource Scheduler (CRS), a domain-agnostic Go library for selecting, prioritizing, routing, and maintaining reusable resources under heavy concurrent load.

The core idea is simple:

                  MANY CONCURRENT REQUESTS
                           │
                           ▼
                  ┌───────────────────┐
                  │ Resource Scheduler│
                  └─────────┬─────────┘
                            │
             ┌──────────────┼──────────────┐
             │              │              │
             ▼              ▼              ▼
        Priority         Acquire        State
          Heap           Strategy      Management
             │              │              │
             └──────────────┼──────────────┘
                            │
                            ▼
                    BEST AVAILABLE
                       RESOURCE
Enter fullscreen mode Exit fullscreen mode

But making that work correctly under concurrency is where things get interesting.

CRS is designed for use cases such as:

  • LLM/API gateways
  • API key pools
  • proxy rotation
  • database replicas
  • GPU workers
  • backend pools
  • worker resources
  • connection pools
  • rate-limited providers
  • reusable compute resources

The scheduler itself does not know what a resource means.

It only knows:

"I have resources. I need to safely maintain them, prioritize them, and return an appropriate one to a concurrent caller."


Table of Contents

  1. The Problem
  2. The Naive Approach
  3. Why a Global Mutex Becomes a Problem
  4. The Core Idea Behind CRS
  5. Architecture at a Glance
  6. Sharded Priority Heaps
  7. Why Sharding Helps
  8. The O(1) Lookup Map
  9. Priority and Acquire Are Different Problems
  10. Acquire Strategies
  11. Round Robin
  12. Weighted Acquire
  13. Adaptive Acquire
  14. Affinity Routing
  15. Shared vs Exclusive Acquisition
  16. Resource Lifecycle
  17. Atomic State Transitions
  18. The Inactive Store
  19. Batch Operations
  20. Updates Without Destroying Heap Ordering
  21. Cooldowns
  22. Asynchronous Events
  23. Observability
  24. Prometheus Integration
  25. Concurrency Model
  26. Complexity
  27. Testing the Library
  28. Race Detector Validation
  29. Real-World Load Testing
  30. 10,000 Concurrent Workers
  31. Burst Testing
  32. Failure Testing
  33. Cooldown Stress Testing
  34. What the Load Tests Actually Tell Us
  35. A Minimal Example
  36. LLM Gateway Example
  37. Why CRS Is Domain-Agnostic
  38. Project Structure
  39. Design Principles
  40. Lessons Learned
  41. When You Should NOT Use CRS
  42. Future Directions
  43. Final Thoughts

The Problem

Let's start with a realistic scenario.

Imagine an LLM gateway with 100 API keys.

Each key may have different:

  • rate-limit availability
  • priority
  • health
  • cooldown state
  • provider
  • capacity
  • temporary availability

Thousands of requests arrive concurrently.

A simplified system looks like:

                   ┌──────────────┐
Request 1 ────────►              │
Request 2 ────────►              │
Request 3 ────────►   Gateway    │
Request 4 ────────►              │
Request 5 ────────►              │
   ...             │              │
Request N ────────►              │
                   └──────┬───────┘
                          │
                          ▼
                  ┌───────────────┐
                  │ Resource Pool │
                  └───────┬───────┘
                          │
           ┌──────────────┼──────────────┐
           ▼              ▼              ▼
        API Key 1      API Key 2      API Key N
Enter fullscreen mode Exit fullscreen mode

The scheduler now has to answer:

  • Which resource should this request use?
  • Which resource has the best priority?
  • Is the resource currently active?
  • Can multiple requests use it simultaneously?
  • Should this resource temporarily leave the pool?
  • Which shard should we search?
  • Should requests stick to the same shard?
  • What happens when the resource is released?
  • How do we update its priority?
  • How do we observe all of this without slowing down the hot path?

That is the problem CRS tries to solve.


The Naive Approach

The easiest implementation looks something like:

type Scheduler struct {
    mu        sync.Mutex
    resources []*Resource
}
Enter fullscreen mode Exit fullscreen mode

Then:

func (s *Scheduler) Acquire() *Resource {
    s.mu.Lock()
    defer s.mu.Unlock()

    // Scan resources.
    // Find the best one.
    // Return it.

    return best
}
Enter fullscreen mode Exit fullscreen mode

At first glance, this looks perfectly reasonable.

For 10 resources and 2 goroutines, it probably is.

But imagine:

Resources:           10,000
Concurrent requests: 5,000
Enter fullscreen mode Exit fullscreen mode

Now every operation fights over one lock.

                GLOBAL MUTEX
                     │
       ┌─────────────┼─────────────┐
       │             │             │
       ▼             ▼             ▼
    Worker 1      Worker 2      Worker 3
       │             │             │
       └─────────────┼─────────────┘
                     │
                  WAITING
Enter fullscreen mode Exit fullscreen mode

The scheduler becomes serialized around the lock.


Why a Global Mutex Becomes a Problem

There are several problems.

1. Lock contention

Only one goroutine can manipulate the pool at a time.

2. Linear scanning

If resources are stored in an array, finding the best resource can become:

O(N)
Enter fullscreen mode Exit fullscreen mode

per acquisition.

3. Priority maintenance

If resources have priorities that change, the scheduler has to continuously maintain ordering.

4. State transitions

Resources can move between:

ACTIVE
INACTIVE
REMOVED
Enter fullscreen mode Exit fullscreen mode

and those transitions must be synchronized.

5. Observability

Metrics and event callbacks should not block the scheduler.

The challenge is therefore not just:

"How do I build a priority queue?"

It is:

"How do I build a concurrent priority resource manager where priority, acquire, lifecycle, and observability coexist?"


The Core Idea Behind CRS

The central architectural decision was:

Don't put one global lock around the entire priority structure.

Instead, CRS partitions resources into independently locked shards.

Conceptually:

                  CRS
                   │
        ┌──────────┼──────────┐
        │          │          │
        ▼          ▼          ▼
     Shard 1    Shard 2    Shard N
        │          │          │
      Heap       Heap       Heap
        │          │          │
     Mutex      Mutex      Mutex
Enter fullscreen mode Exit fullscreen mode

Each shard owns its own heap and its own lock.

This is the heart of CRS.


Architecture at a Glance

                        APPLICATION
                             │
               Add / Acquire / Release / Update
                             │
                             ▼
                  ┌────────────────────┐
                  │   CRS Scheduler    │
                  └─────────┬──────────┘
                            │
            ┌───────────────┼────────────────┐
            │               │                │
            ▼               ▼                ▼
      Acquire          Lookup          Inactive
       Strategy            Map             Store
            │               │                │
            ▼               ▼                │
     Candidate Shard     O(1) Node          │
            │                                │
            ▼                                │
    ┌─────────────────────────────────┐      │
    │         ACTIVE HEAP SHARDS       │      │
    │                                 │      │
    │  Heap 1     Heap 2     Heap N   │      │
    │  +Mutex     +Mutex     +Mutex   │      │
    └─────────────────────────────────┘      │
            │                                │
            └──────────────┬─────────────────┘
                           │
                           ▼
                    EVENT DISPATCHER
                           │
                 ┌─────────┴──────────┐
                 ▼                    ▼
             Telemetry            Cooldown
                 │                    │
                 ▼                    ▼
             Prometheus          Resource State
Enter fullscreen mode Exit fullscreen mode

This separation is intentional.


Sharded Priority Heaps

Each shard maintains a priority heap.

For example:

Shard 1

        [Priority 10]
        /           \
 [Priority 20]    [Priority 30]
    /     \
[40]      [50]
Enter fullscreen mode Exit fullscreen mode

Another shard:

Shard 2

        [Priority 5]
        /          \
 [Priority 15]   [Priority 25]
Enter fullscreen mode Exit fullscreen mode

Every shard has its own synchronization boundary.

Shard 1                 Shard 2                 Shard 3

┌─────────────┐        ┌─────────────┐        ┌─────────────┐
│   Mutex     │        │   Mutex     │        │   Mutex     │
├─────────────┤        ├─────────────┤        ├─────────────┤
│ Priority    │        │ Priority    │        │ Priority    │
│ Heap        │        │ Heap        │        │ Heap        │
└─────────────┘        └─────────────┘        └─────────────┘
Enter fullscreen mode Exit fullscreen mode

There is no global heap mutex.


Why Sharding Helps

Suppose we have 32 shards.

Instead of:

1 global lock
Enter fullscreen mode Exit fullscreen mode

we have:

32 independently locked heaps
Enter fullscreen mode Exit fullscreen mode

Different goroutines can operate on different shards simultaneously.

Goroutine A ─────► Shard 1 ─────► lock
Goroutine B ─────► Shard 7 ─────► lock
Goroutine C ─────► Shard 19 ────► lock
Goroutine D ─────► Shard 27 ────► lock
Enter fullscreen mode Exit fullscreen mode

The locks are independent.

This doesn't magically eliminate contention.

If every request targets the same shard, that shard can still become contended.

That's why CRS also separates:

acquire strategy

from

priority ordering

This distinction is extremely important.


The O(1) Lookup Map

A heap is excellent at answering:

"What is the best resource?"

But a heap is not ideal for answering:

"Where is resource X?"

Searching a heap can require scanning.

CRS therefore maintains an additional lookup structure.

Conceptually:

ID
 │
 ▼
┌──────────────────────┐
│ Lookup Map            │
│                      │
│ "backend-01" ───────► Node
│ "backend-02" ───────► Node
│ "backend-03" ───────► Node
└──────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The lookup map is protected independently with a read/write mutex.

This gives the scheduler an O(1)-style membership/location lookup by application-defined key.

That is particularly useful for:

Get
Update
Remove
Release
Exclude
Include
Enter fullscreen mode Exit fullscreen mode

without scanning every heap.


Priority and Acquire Are Different Problems

This is one of the most important design ideas in CRS.

A resource can have:

Priority = 10
Enter fullscreen mode Exit fullscreen mode

but that doesn't necessarily tell us:

Which shard should we inspect first?

These are separate decisions.

CRS therefore separates:

                REQUEST
                   │
                   ▼
            ACQUIRE STRATEGY
                   │
                   ▼
              SHARD SELECTION
                   │
                   ▼
             PRIORITY HEAP
                   │
                   ▼
             BEST RESOURCE
Enter fullscreen mode Exit fullscreen mode

This allows different routing strategies to be plugged into the scheduler without changing the underlying heap implementation.


Acquire Strategies

CRS provides several acquire approaches:

  • Round Robin
  • Weighted
  • Adaptive
  • Consistent Hashing for affinity routing

Each solves a different problem.


Round Robin

Round Robin is the simplest.

Request 1 → Shard 1
Request 2 → Shard 2
Request 3 → Shard 3
Request 4 → Shard 4
Request 5 → Shard 1
...
Enter fullscreen mode Exit fullscreen mode

It is simple and predictable.

Use it when:

  • shards are roughly equivalent
  • you want even distribution
  • resource capacity is similar

Weighted Acquire

Not every shard is necessarily equal.

Imagine:

GPU 1 → 24 GB VRAM
GPU 2 → 24 GB VRAM
GPU 3 → 80 GB VRAM
GPU 4 → 80 GB VRAM
Enter fullscreen mode Exit fullscreen mode

You may want larger resources to receive more work.

Weighted acquire lets you express relative capacity.

Conceptually:

Shard 1: weight 1
Shard 2: weight 1
Shard 3: weight 4
Shard 4: weight 4
Enter fullscreen mode Exit fullscreen mode

Traffic can then be distributed proportionally.

This is useful for:

  • heterogeneous GPUs
  • backend instances with different capacity
  • API providers with different quotas
  • worker pools with different performance characteristics

Adaptive Acquire

Round Robin doesn't know anything about current load.

Adaptive acquire attempts to account for shard activity.

Conceptually:

             REQUEST
                │
                ▼
       ┌──────────────────┐
       │ Inspect shard    │
       │ load information  │
       └────────┬─────────┘
                │
       ┌────────┼────────┐
       ▼        ▼        ▼
    Shard A   Shard B   Shard C
      busy      low       busy
                │
                ▼
            choose B
Enter fullscreen mode Exit fullscreen mode

The scheduler uses lightweight shard-level state to favor less-contended shards without introducing another global lock.

This is useful when the resource pool is dynamic and simple round-robin distribution isn't enough.


Affinity Routing

Sometimes you don't want random distribution.

You want:

user-123 → same shard
user-456 → same shard
Enter fullscreen mode Exit fullscreen mode

CRS supports affinity routing through consistent hashing.

Conceptually:

                   HASH RING

             ┌───────────────────┐
             │                   │
        S1   │        S2         │
             │                   │
             │                   │
        S4   │        S3         │
             │                   │
             └───────────────────┘

              ▲
              │
        hash("user-123")
Enter fullscreen mode Exit fullscreen mode

The same affinity identifier deterministically maps to the same shard.

This is useful for:

  • sticky sessions
  • tenant affinity
  • cache locality
  • connection locality
  • stateful workers

Shared vs Exclusive Acquisition

CRS supports two major acquisition semantics.

Shared

The resource remains active.

ACTIVE
  │
  │ Acquire
  ▼
ACTIVE
Enter fullscreen mode Exit fullscreen mode

Multiple callers can acquire the same resource.

This is useful when resources represent things like:

  • API keys
  • read replicas
  • stateless endpoints
  • shared provider capacity

Exclusive

The resource temporarily leaves the active pool.

ACTIVE
   │
   │ Acquire
   ▼
INACTIVE
   │
   │ Release
   ▼
ACTIVE
Enter fullscreen mode Exit fullscreen mode

This is useful when a resource represents something that cannot safely be used by multiple concurrent operations.

Examples:

  • GPU worker
  • exclusive connection
  • physical device
  • single-use worker
  • exclusive job executor

Resource Lifecycle

A CRS resource essentially moves through states.

                  ┌───────────┐
                  │   ADD     │
                  └─────┬─────┘
                        ▼
                 ┌─────────────┐
                 │   ACTIVE    │
                 └──────┬──────┘
                        │
              ┌─────────┼─────────┐
              │         │         │
           Acquire    Exclude   Remove
              │         │         │
              ▼         ▼         ▼
          INACTIVE   INACTIVE   DELETED
              │
            Release
              │
              ▼
           ACTIVE
Enter fullscreen mode Exit fullscreen mode

The important invariant is:

A resource should exist in exactly one state/location at a time.


Atomic State Transitions

Imagine:

Goroutine A: Acquire(resource)
Goroutine B: Remove(resource)
Goroutine C: Update(resource)
Enter fullscreen mode Exit fullscreen mode

all happening at almost the same time.

Without careful synchronization, you can get:

Resource exists in heap
AND
Resource exists in inactive store
Enter fullscreen mode Exit fullscreen mode

or:

Lookup map says ACTIVE
but heap doesn't contain it
Enter fullscreen mode Exit fullscreen mode

Those are catastrophic consistency bugs.

CRS therefore treats state transitions as carefully synchronized operations involving:

  • lookup state
  • heap state
  • inactive state
  • shard locks

The Inactive Store

The inactive store is particularly important for Exclusive acquisition.

Suppose:

backend-01
Enter fullscreen mode Exit fullscreen mode

is acquired exclusively.

It is removed from the active heap and stored as inactive.

ACTIVE HEAP

backend-01
backend-02
backend-03
Enter fullscreen mode Exit fullscreen mode

After acquisition:

ACTIVE HEAP

backend-02
backend-03


INACTIVE STORE

backend-01
Enter fullscreen mode Exit fullscreen mode

When released:

INACTIVE STORE
      │
      │ Release
      ▼
ACTIVE HEAP
Enter fullscreen mode Exit fullscreen mode

The lookup map allows CRS to locate the resource without scanning every heap.


Batch Operations

Adding resources one at a time is easy.

But imagine importing 10,000 resources.

You don't want partial state like:

Batch:
1 ✓
2 ✓
3 ✓
4 ✓
5 ✗
6 ?
7 ?
...
Enter fullscreen mode Exit fullscreen mode

CRS provides BatchAdd with atomic insertion behavior.

Conceptually:

                BatchAdd
                   │
                   ▼
          ┌─────────────────┐
          │ Validate batch  │
          └────────┬────────┘
                   │
                   ▼
          ┌─────────────────┐
          │ Prepare changes │
          └────────┬────────┘
                   │
                   ▼
          ┌─────────────────┐
          │ Insert shards   │
          └────────┬────────┘
                   │
                   ▼
              COMPLETE
Enter fullscreen mode Exit fullscreen mode

The goal is to avoid exposing a partially inserted batch.


Updates Without Destroying Heap Ordering

Suppose:

A = priority 10
B = priority 20
C = priority 30
Enter fullscreen mode Exit fullscreen mode

Now:

C → priority 5
Enter fullscreen mode Exit fullscreen mode

C should move toward the top.

Simply changing the value isn't enough.

The heap must be repaired.

Conceptually:

Before:

        A(10)
       /     \
    B(20)   C(30)


Update:

C(30) → C(5)


After:

        C(5)
       /    \
    A(10)  B(20)
Enter fullscreen mode Exit fullscreen mode

This preserves the priority-queue invariant.


Cooldowns

Real resources sometimes need a cooldown period.

For example:

API key hits rate limit
        │
        ▼
   cooldown 5s
        │
        ▼
available again
Enter fullscreen mode Exit fullscreen mode

CRS includes a cooldown extension.

The architecture is event-driven:

Acquire
   │
   ▼
Release
   │
   ▼
Event Dispatcher
   │
   ▼
Cooldown Manager
   │
   ▼
Exclude resource
   │
   ▼
wait
   │
   ▼
Include resource
Enter fullscreen mode Exit fullscreen mode

The cooldown extension is asynchronous. That means there can be a small eventual-consistency window between a release event and the cooldown observer processing it.

That is intentional.


Asynchronous Events

Observability and extensions should not unnecessarily slow the scheduler's hot path.

CRS therefore uses an event dispatcher.

                 Scheduler
                     │
                     │ emit event
                     ▼
             ┌───────────────┐
             │ Buffered      │
             │ Event Stream  │
             └───────┬───────┘
                     │
              background worker
                     │
          ┌──────────┼──────────┐
          ▼          ▼          ▼
      Observer 1  Observer 2  Observer 3
          │          │          │
          ▼          ▼          ▼
       Metrics   Cooldown   Prometheus
Enter fullscreen mode Exit fullscreen mode

The scheduler doesn't need to execute arbitrary observer logic while holding heap locks.


Observability

A production scheduler should answer:

How many acquisitions happened?
How many releases?
Which resources are being used?
How many failures?
What is the current resource count?
Enter fullscreen mode Exit fullscreen mode

CRS includes telemetry support using atomic counters and asynchronous events.

The goal is:

Scheduler hot path
      │
      ▼
lightweight event
      │
      ▼
asynchronous telemetry
Enter fullscreen mode Exit fullscreen mode

rather than doing expensive monitoring work while holding scheduler locks.


Prometheus Integration

CRS provides a Prometheus exporter extension.

Conceptually:

               CRS
                │
                ▼
          Telemetry
                │
                ▼
      ┌─────────────────┐
      │ Atomic Counters │
      └────────┬────────┘
               │
               ▼
       Prometheus Collector
               │
               ▼
       /metrics endpoint
Enter fullscreen mode Exit fullscreen mode

This makes it possible to expose scheduler activity to an existing monitoring stack.


Concurrency Model

The concurrency model is based on multiple independent synchronization boundaries.

                   Scheduler
                       │
          ┌────────────┼────────────┐
          │            │            │
          ▼            ▼            ▼
      Shard 1      Shard 2       Shard N
       Mutex         Mutex         Mutex
          │            │            │
          ▼            ▼            ▼
        Heap          Heap          Heap
Enter fullscreen mode Exit fullscreen mode

Separately:

Lookup Map
    │
sync.RWMutex
Enter fullscreen mode Exit fullscreen mode

And:

Telemetry
    │
atomic counters / event channel
Enter fullscreen mode Exit fullscreen mode

This is much more granular than one mutex around everything.


Complexity

The scheduler is designed around heap and lookup properties.

Operation Complexity Synchronization
Add O(log N) Single shard lock
BatchAdd O(log N) per insertion Shard locks
Acquire Acquire + heap operation Shard lock
AcquireByAffinity Hash lookup + heap operation Shard lock
Release O(log N) Single shard lock
Update O(log N) active / O(1) inactive Shard lock
Remove O(log N) active / O(1) inactive Shard lock
Get O(1)-style lookup Lookup synchronization
Stats O(Shards) Short shard reads

Where:

N = resources in a shard
Enter fullscreen mode Exit fullscreen mode

Actual performance depends on workload, shard count, resource distribution, acquire policy, and contention.


Testing the Library

A concurrency library cannot be validated with only:

go test ./...
Enter fullscreen mode Exit fullscreen mode

Unit tests answer:

"Does this operation behave correctly?"

Load tests answer:

"What happens when thousands of goroutines continuously hammer it?"

The CRS test strategy includes:

                  TESTING
                     │
       ┌─────────────┼─────────────┐
       │             │             │
       ▼             ▼             ▼
   Unit Tests    Race Tests    Load Tests
       │             │             │
       ▼             ▼             ▼
  correctness    data races    behavior
Enter fullscreen mode Exit fullscreen mode

Tests cover:

  • scheduler construction
  • configuration validation
  • Add
  • BatchAdd
  • Get
  • Len
  • Stats
  • Acquire
  • Release
  • Update
  • Include
  • Exclude
  • Remove
  • Shutdown
  • acquire strategies
  • event observers
  • cooldown
  • telemetry
  • Prometheus collection
  • concurrent stress behavior

Race Detector Validation

One of the most important commands for a concurrent Go library is:

go test -count=1 -race ./...
Enter fullscreen mode Exit fullscreen mode

The race detector doesn't prove that a library is bug-free.

But it can catch a class of extremely dangerous concurrency problems that ordinary tests may miss.

The final release was validated with the race detector without reported data races.


Real-World Load Testing

For CRS, I created a dedicated load-test harness.

The harness models:

  • concurrent workers
  • backend processing delay
  • random request durations
  • backend failures
  • request cancellation
  • different scenarios
  • resource utilization
  • acquisition latency
  • backend latency
  • total latency
  • throughput
  • cooldown behavior
  • burst traffic

The important distinction is that these are load-test results, not universal benchmarks.


10,000 Concurrent Workers

One of the larger tests used:

Concurrency: 10,000 workers
Resources:   4 backends
Duration:    60 seconds
Scenario:    normal
Acquire:     adaptive
Policy:      shared
Race detector: enabled
Enter fullscreen mode Exit fullscreen mode

The test produced approximately:

Total requests:        2,248,610
Successful:            2,226,165
Failed:                22,445

Throughput:            ~37,137.28 req/s

Peak simultaneous:     10,000
Enter fullscreen mode Exit fullscreen mode

Resource utilization was extremely well-balanced:

backend-01    25.01%
backend-02    24.96%
backend-03    25.03%
backend-04    25.00%
Enter fullscreen mode Exit fullscreen mode

The scheduler acquisition latency was remarkably low, and should be clearly distinguished from the simulated backend latency:

Acquire Latency:
mean: 7.59µs
max:  51.893ms

Backend Latency:
p50: 260.334ms
p95: 477.661ms
p99: 496.863ms

Total Request Latency:
p50: 262.430ms
p95: 479.425ms
p99: 498.559ms
Enter fullscreen mode Exit fullscreen mode

Accounting check verified complete consistency:

success + backend failures = 2,248,610
backend requests = 2,248,610
total request attempts = 2,248,610
accounting: OK
Enter fullscreen mode Exit fullscreen mode

Note: The 22,445 failures correspond exactly to the configured 1% simulated backend failure rate, not scheduler acquire failures (which remained at 0). Acquire timeouts and release failures were also 0.

The test environment was:

OS:           Windows
Architecture: amd64
CPU cores:    12
GOMAXPROCS:   12
Go:           1.25.5
Enter fullscreen mode Exit fullscreen mode

The backend latency was simulated and should not be confused with CRS scheduler latency. These are workload-specific test results and should NOT be presented as universal benchmark claims.


Burst Testing

Real systems often receive bursts.

Traffic
  │
  │       ███████████
  │       ███████████
  │       ███████████
  │
  │  ███
  │  ███
  │  ███
  └────────────────────────► time
Enter fullscreen mode Exit fullscreen mode

The burst scenario generated approximately:

Requests:        468,791
Successful:      464,109
Failed:          4,682

Throughput:      ~7,746 req/s

Peak simultaneous: 5,000
Enter fullscreen mode Exit fullscreen mode

The success rate was approximately:

99%
Enter fullscreen mode Exit fullscreen mode

Backend latency remained around:

p50 ≈ 260 ms
p95 ≈ 476 ms
p99 ≈ 495 ms
Enter fullscreen mode Exit fullscreen mode

These backend numbers came from the simulated workload.


Failure Testing

A scheduler should also behave correctly when resources fail.

I tested a failure scenario with:

Concurrency: 1000
Resources:   4
Duration:    60 seconds
Failure rate: 10%
Enter fullscreen mode Exit fullscreen mode

The workload produced approximately:

Requests:       231,443
Successful:     208,293
Failed:          23,150
Success rate:    90%
Enter fullscreen mode Exit fullscreen mode

The scheduler reported:

Acquire failures: 0
Release failures: 0
Backend failures: 23,150
Enter fullscreen mode Exit fullscreen mode

That distinction matters.

The scheduler successfully acquired resources while simulated backend operations failed at the expected rate.


Cooldown Stress Testing

Cooldown was one of the more interesting tests.

With exclusive acquisition, resources temporarily leave the active pool.

                4 resources
                     │
                     ▼
              ┌─────────────┐
              │   ACTIVE    │
              └──────┬──────┘
                     │
                Acquire
                     │
                     ▼
              ┌─────────────┐
              │  INACTIVE   │
              └──────┬──────┘
                     │
                  cooldown
                     │
                     ▼
              ┌─────────────┐
              │   ACTIVE    │
              └─────────────┘
Enter fullscreen mode Exit fullscreen mode

At 1,000 concurrent workers, the scheduler generated a very large number of acquisition attempts while only a limited number of resources were available.

This demonstrated an important property:

High concurrency doesn't mean unlimited successful backend concurrency.

If only four resources exist and the policy is exclusive, four resources are still four resources.


What the Load Tests Actually Tell Us

The load tests gave several useful observations.

1. Sharding worked under heavy concurrency

The scheduler continued operating with thousands of concurrent workers without race-detector failures.

2. Acquire remained balanced

The adaptive workload distributed acquisitions across four backends at approximately 25% each.

3. Resource state remained consistent

The accounting checks showed that acquired resources were not silently lost.

4. Failures remained distinguishable

The harness separated:

Acquire failure
Backend failure
Release failure
Timeout
Enter fullscreen mode Exit fullscreen mode

A scheduler failure and a backend failure are very different operational problems.

5. Cooldown changes the workload completely

With exclusive resources, the bottleneck becomes resource availability rather than CPU.


A Minimal Example

The basic usage pattern is:

package main

import (
    "fmt"
    "log"

    "github.com/phero20/concurrent-resource-scheduler/config"
    "github.com/phero20/concurrent-resource-scheduler/scheduler"
)

type Worker struct {
    ID       string
    Priority int
}

func main() {
    compare := func(a, b *Worker) int {
        if a.Priority < b.Priority {
            return -1
        }

        if a.Priority > b.Priority {
            return 1
        }

        return 0
    }

    keyFunc := func(w *Worker) string {
        return w.ID
    }

    cfg := config.Config[*Worker, string]{
        HeapCount: 1,
        Comparator: compare,
        KeyFunc:    keyFunc,
    }

    sched, err := scheduler.New(cfg)
    if err != nil {
        log.Fatal(err)
    }

    defer sched.Shutdown()

    sched.Add(&Worker{
        ID:       "worker-1",
        Priority: 50,
    })

    sched.Add(&Worker{
        ID:       "worker-2",
        Priority: 10,
    })

    resource, err := sched.Acquire()

    if err != nil {
        log.Fatal(err)
    }

    fmt.Println("Acquired:", resource.ID)
}
Enter fullscreen mode Exit fullscreen mode

The application controls the resource type, key, and comparison logic.

The scheduler handles the concurrent resource management.


LLM Gateway Example

Imagine:

type APIKey struct {
    ID        string
    Provider  string
    Remaining int
}
Enter fullscreen mode Exit fullscreen mode

Your application could define priority based on remaining quota.

Conceptually:

API Key
   │
   ├── Provider
   ├── Remaining quota
   └── Health
Enter fullscreen mode Exit fullscreen mode

The comparator could make a key with more remaining capacity more desirable.

The scheduler doesn't need to know what those fields mean.

The application owns that logic.

CRS maintains the ordering and concurrent lifecycle.

This is the core idea behind being domain-agnostic.


Why CRS Is Domain-Agnostic

The scheduler doesn't contain logic like:

if resource.IsGPU() {
    ...
}
Enter fullscreen mode Exit fullscreen mode

or:

if provider == "openai" {
    ...
}
Enter fullscreen mode Exit fullscreen mode

Instead, the application provides:

Resource type
     +
Key function
     +
Comparator
     +
Acquire strategy
Enter fullscreen mode Exit fullscreen mode

This makes the same scheduler applicable to many domains.


Example: GPU Workers

                Job
                 │
                 ▼
          ┌─────────────┐
          │ CRS         │
          └──────┬──────┘
                 │
       ┌─────────┼─────────┐
       ▼         ▼         ▼
      GPU 1     GPU 2     GPU 3
      24 GB     24 GB     80 GB
Enter fullscreen mode Exit fullscreen mode

Weighted acquire could favor larger GPUs.

Exclusive acquisition could prevent two jobs from taking the same GPU.


Example: Database Replicas

                 Query
                   │
                   ▼
                 CRS
                   │
        ┌──────────┼──────────┐
        ▼          ▼          ▼
      DB-1       DB-2       DB-3
       20%        70%        35%
Enter fullscreen mode Exit fullscreen mode

The application could define priority around current load.

The scheduler remains unaware that the resources happen to be databases.


Example: Proxy Pool

                  HTTP Request
                       │
                       ▼
                     CRS
                       │
        ┌──────────────┼──────────────┐
        ▼              ▼              ▼
    Proxy A         Proxy B        Proxy C
    healthy         cooldown       healthy
Enter fullscreen mode Exit fullscreen mode

Cooldown can temporarily remove a failing proxy.

Affinity can keep a particular tenant or session mapped consistently.


Project Structure

concurrent-resource-scheduler/
│
├── config/
├── scheduler/
├── acquire/
├── internal/
│   ├── heap/
│   ├── lookup/
│   └── node/
├── extensions/
│   ├── cooldown/
│   ├── metrics/
│   └── prometheus/
├── events/
├── errors/
├── stats/
├── examples/
│   ├── basic/
│   ├── batch/
│   ├── affinity/
│   ├── cooldown/
│   ├── exclusive/
│   ├── shared/
│   └── prometheus/
├── docs/
│   ├── API.md
│   ├── ARCHITECTURE.md
│   ├── OVERVIEW.md
├── README.md
├── CHANGELOG.md
├── CONTRIBUTING.md
├── LICENSE
└── go.mod
Enter fullscreen mode Exit fullscreen mode

The main architectural boundary is:

Public API
    │
    ▼
Scheduler
    │
    ├── Acquire
    ├── Heap
    ├── Lookup
    ├── Events
    └── Extensions
Enter fullscreen mode Exit fullscreen mode

Design Principles

Several principles shaped CRS.

1. No global heap lock

The heap state is partitioned.

2. Separate routing from priority

Acquire chooses where to look.

The heap chooses what is best.

3. Keep the scheduler domain-agnostic

The application owns business logic.

4. Keep observability outside the hot path

Telemetry should not become the scheduler bottleneck.

5. Make state transitions explicit

Resources are either active, inactive, or removed.

6. Optimize for concurrent workloads

Concurrency is not an afterthought.

It is part of the architecture.


Lessons Learned

Building a concurrent library taught me something important:

Concurrency problems are usually not caused by one complicated function.

They are caused by interactions between simple functions.

For example:

Acquire
   +
Release
   +
Update
   +
Remove
   +
Cooldown
   +
Events
   +
Concurrent callers
Enter fullscreen mode Exit fullscreen mode

Each function can be correct individually while the combination is broken.

That's why concurrency testing needs to go beyond unit tests.


Another Lesson: Observability Is Part of the Design

When you have:

10,000 requests
4 resources
multiple shards
multiple goroutines
failures
cooldowns
Enter fullscreen mode Exit fullscreen mode

you need to know what actually happened.

That's why CRS exposes events and telemetry around scheduler activity.

Without observability, debugging concurrent systems becomes guesswork.


Another Lesson: Load Tests Need Accounting

One of the most useful parts of the load harness was accounting.

For example:

total requests
      │
      ├── acquire failures
      │
      └── reached backend
              │
              ├── successful
              └── backend failure
Enter fullscreen mode Exit fullscreen mode

If:

success + backend_failure
Enter fullscreen mode Exit fullscreen mode

doesn't match:

backend_requests
Enter fullscreen mode Exit fullscreen mode

something is wrong.

Accounting checks are powerful for detecting concurrency bugs.


When You Should NOT Use CRS

CRS isn't intended to replace every queue or pool implementation.

You probably don't need it if you simply have:

10 objects
+
one goroutine
Enter fullscreen mode Exit fullscreen mode

A simple slice may be better.

You also probably don't need CRS if your problem is purely:

"I need a FIFO queue."

Use a simpler primitive.

CRS becomes interesting when you need combinations of:

concurrency
+
priority
+
acquire
+
resource lifecycle
+
affinity
+
cooldowns
+
observability
Enter fullscreen mode Exit fullscreen mode

Future Directions

Potential future areas include:

  • deeper benchmarking
  • additional acquire strategies
  • richer scheduling policies
  • more advanced resource health models
  • better operational tooling
  • additional observability integrations
  • workload-specific tuning
  • broader real-world validation

The goal is not to optimize for feature count.

The goal is to keep the core scheduler:

small
predictable
composable
concurrent
domain-agnostic
Enter fullscreen mode Exit fullscreen mode

Installation

Install CRS with:

go get github.com/phero20/concurrent-resource-scheduler
Enter fullscreen mode Exit fullscreen mode

Then explore the examples and documentation in the repository.


Final Thoughts

The interesting part of building a scheduler isn't writing:

Acquire()
Enter fullscreen mode Exit fullscreen mode

The interesting part is everything around it.

You need to think about:

                    RESOURCE SCHEDULING

                         ┌─────────┐
                         │ Priority│
                         └────┬────┘
                              │
        ┌─────────────────────┼─────────────────────┐
        │                     │                     │
        ▼                     ▼                     ▼
   Concurrency           Acquire              Lifecycle
        │                     │                     │
        ▼                     ▼                     ▼
   Sharded heaps       Adaptive/Weighted       Active/Inactive
        │               Round Robin             Release
        │                     │                     │
        └─────────────────────┼─────────────────────┘
                              │
                              ▼
                         Observability
                              │
                   ┌──────────┴──────────┐
                   ▼                     ▼
               Telemetry            Prometheus
Enter fullscreen mode Exit fullscreen mode

CRS was built around that complete picture.

It is not just a priority queue.

It is a concurrent resource-management layer that combines:

  • sharded priority heaps
  • concurrent-safe lookup
  • configurable acquire
  • priority ordering
  • shared and exclusive acquisition
  • affinity routing
  • resource lifecycle management
  • cooldown extensions
  • asynchronous events
  • telemetry
  • Prometheus integration

And the project has been tested beyond basic unit tests, including:

Race detector
       +
Concurrent stress tests
       +
1,000 workers
       +
2,000 workers
       +
5,000 workers
       +
10,000 workers
       +
Burst workloads
       +
Failure workloads
       +
Cooldown workloads
Enter fullscreen mode Exit fullscreen mode

The most important result isn't a single throughput number.

It's that the architecture gives me a foundation where concurrency, scheduling policy, resource state, and observability can evolve independently.

That's what I wanted to build.


If You're Building Something Similar

I'd love to hear how you approach resource scheduling.

Especially if you're working on:

  • LLM gateways
  • GPU schedulers
  • proxy pools
  • database routing
  • distributed workers
  • API key rotation
  • connection pools
  • high-concurrency Go services

What would you change about this architecture?

What workloads should I test next?

What concurrency problems have you encountered in production?

I'd genuinely like to hear from people who have operated systems like these in the real world.


Project

Support on GitHub: github.com/phero20/concurrent-resource-scheduler (Give it a star if you find it useful!)

View Docs: pkg.go.dev/github.com/phero20/concurrent-resource-scheduler

Concurrent Resource Scheduler (CRS)

A high-performance, domain-agnostic Go resource scheduler built around sharded priority heaps, concurrent-safe lookup, pluggable acquire strategies, affinity routing, lifecycle management, cooldowns, and observability.

If you find the project useful, consider giving it a star.

Found a bug? Open an issue.

Have an idea? Start a discussion.

The best validation for a concurrency library isn't another local test.

It's seeing it survive workloads you didn't design yourself.

Top comments (1)

Collapse
 
anabolic profile image
Anabolic

Really nice approach. I especially like the separation between shard selection and priority ordering. Sharding the heaps is a clean way to reduce contention under high concurrency.