DEV Community

Rahul Kumar
Rahul Kumar

Posted on

I ran a small Go experiment to see how CPU cores affect performance.

## 1. Understanding CPU cores at the system level

At the system level, the CPU is responsible for executing instructions. We have different numbers of CPU cores, such as 1 core, 2 cores, 4 cores, 8 cores, etc.

Each logical CPU can execute one thread at a time.

For example:

  • 1 logical CPU → 1 thread executing at a time.
  • 2 logical CPUs → 2 threads executing simultaneously.
  • 4 logical CPUs → 4 threads executing simultaneously.

However, having more CPU cores does not automatically mean faster execution. It depends on how the application utilizes those cores and whether the workload can be executed in parallel.

2. Concurrency vs. Parallelism

Concurrency: Multiple tasks are in progress, but they don't necessarily execute simultaneously.

For example, on a single CPU core, the operating system can switch between multiple tasks. To humans, they may appear to be executing simultaneously because the switching happens very quickly.

Parallelism: Two or more tasks execute at the same time using separate execution resources, such as different CPU cores.

For example:

  • Task 1 executes on Core 1.
  • Task 2 executes on Core 2.

Both tasks are executing simultaneously.

Remember: Concurrency is about managing multiple tasks at once, whereas parallelism is about executing multiple tasks at the same instant.

3. How Go handles task execution

Go provides its own concurrency mechanism through goroutines and a runtime scheduler.

Goroutines

A goroutine is a lightweight execution unit managed by the Go runtime.

For example:

  • 1 task → potentially 1 goroutine.
  • 2 independent tasks → potentially 2 goroutines.
  • 10,000 independent tasks → potentially 10,000 goroutines.

However, the number of goroutines depends on how we design our application. We don't necessarily need one goroutine for every individual job.

Go can efficiently manage thousands of goroutines without creating an equivalent number of OS threads.

OS Threads

OS threads are execution contexts managed by the operating system.

They are responsible for executing instructions on available logical CPUs.

Go's runtime manages the relationship between goroutines and OS threads.

GOMAXPROCS

GOMAXPROCS controls how many goroutines can execute Go code simultaneously.

By default, Go generally configures this according to the available CPU resources.

For example, if our system has 2 available logical CPUs:

Available CPUs: 2
GOMAXPROCS: 2
Enter fullscreen mode Exit fullscreen mode

This means Go can execute Go code in parallel across up to 2 logical CPUs.

Important: GOMAXPROCS is not the number of OS threads Go creates. Go's runtime can create and manage more OS threads as needed.

4. Relationship between CPU cores, OS threads, and goroutines

Suppose our system has only 2 logical CPUs.

We can still have:

  • 10 goroutines.
  • 10 OS threads.
  • 100 goroutines.
  • Even thousands of goroutines.

But having more threads or goroutines does not mean we have more physical CPU cores.

For example:

Available logical CPUs: 2
Goroutines: 4
GOMAXPROCS: 2
Enter fullscreen mode Exit fullscreen mode

Go can execute up to 2 goroutines simultaneously, while the remaining goroutines wait for their turn or wait for some operation to complete.

The Go runtime schedules goroutines onto OS threads, and the operating system schedules those threads onto available logical CPUs.

5. Does having more threads make execution faster?

Sometimes yes, but not always.

It depends on the workload.

Case 1: CPU-intensive tasks

Suppose we are calculating prime numbers within a specific range.

This task requires continuous CPU computation.

If we have 2 logical CPUs:

  • With GOMAXPROCS=1, only one goroutine can execute Go code at a time.
  • With GOMAXPROCS=2, two goroutines can execute Go code simultaneously.

Therefore, the second configuration can complete the calculation faster because it utilizes both available CPUs.

However, increasing GOMAXPROCS beyond the available CPU capacity does not provide additional physical processing power.

Case 2: I/O-intensive tasks

Suppose our backend server is handling multiple requests, and some requests require:

  • Fetching data from a database.
  • Reading files from disk.
  • Calling external APIs.
  • Waiting for network responses.

During these operations, a goroutine may become blocked or suspended while waiting for the result.

Instead of keeping the CPU idle, Go can schedule another runnable goroutine.

This allows the application to continue processing other requests while some operations are waiting.

This is one of the major reasons backend applications use many goroutines even when they have relatively few CPU cores.

6. Why can excessive threads introduce overhead?

Suppose we have 2 logical CPUs but many runnable OS threads.

The operating system must schedule those threads onto the available CPUs.

When switching between threads, the operating system may need to:

  • Save the execution state of the currently running thread.
  • Restore the execution state of another thread.
  • Update scheduling information.
  • Handle CPU cache effects and other scheduling costs.

This is called context switching.

Context switching is not the same as moving a thread to disk or physically removing it from memory. Threads generally remain represented in memory while they are waiting to execute.

If the system performs excessive scheduling and context switching, it can introduce overhead and reduce performance.

However, more threads do not automatically mean more overhead. The impact depends on how many threads are runnable, what they are doing, and how the workload is structured.

7. Why does Go use more OS threads when the system has limited CPU cores?

Go's runtime manages goroutines and OS threads separately.

For CPU-intensive workloads, Go generally uses its available parallel execution capacity efficiently.

For I/O-intensive workloads, some goroutines may wait for operations to complete.

Go can schedule other runnable goroutines while those operations are waiting.

For certain blocking system calls, an OS thread may become blocked, and the Go runtime can arrange for other goroutines to continue executing using other available threads.

For network I/O, Go commonly uses runtime network polling to avoid requiring one blocked OS thread for every waiting network connection.

This helps Go handle a large number of concurrent requests without requiring an equivalent number of OS threads.

8. Practical Experiment: Measuring CPU Parallelism in Go

We created a Go program to calculate prime numbers within a specified range using multiple goroutines.

The program used:

  • 2 available logical CPUs.
  • 4 goroutines.
  • Different GOMAXPROCS configurations.
  • Execution time as the performance measurement.

Experiment results

GOMAXPROCS Workers Primes found Execution time
1 4 78,498 388.075 ms
2 4 78,498 205.7652 ms
3 4 78,498 244.6589 ms
4 4 78,498 272.6667 ms

What did we observe?

When GOMAXPROCS = 1:

Only one goroutine could execute Go code at a time.

Execution time: approximately 388 ms.

When GOMAXPROCS = 2:

Two goroutines could execute Go code simultaneously across the available CPUs.

Execution time: approximately 206 ms.

The workload completed substantially faster.

When GOMAXPROCS = 3 or 4:

The configuration permitted more parallel execution capacity, but our machine still had only 2 available logical CPUs.

The measured execution time increased compared with the 2-CPU configuration.

This may involve scheduling overhead and measurement variability. Additional parallelism does not guarantee improved performance.

We should repeat the benchmark multiple times and compare averages before drawing a firm conclusion.

Conclusion from the experiment

More CPU cores can improve the performance of CPU-intensive workloads when the application can utilize them effectively.

However, more threads, goroutines, or parallel execution slots do not automatically make an application faster.

We must identify the workload, available hardware resources, and potential bottlenecks.

CODE:

package main

import (
"flag"
"fmt"
"runtime"
"sync"
"time"
)

func isPrime(n int) bool {
if n < 2 {
return false
}

for i := 2; i*i <= n; i++ {
    if n%i == 0 {
        return false
    }
}
return true
Enter fullscreen mode Exit fullscreen mode

}

func main() {
procs := flag.Int("procs", 1, "Number of CPUs to use")
flag.Parse()

runtime.GOMAXPROCS(*procs)

const limit = 1000000
const workers = 4

var wg sync.WaitGroup
var total int64

var mu sync.Mutex

start := time.Now()

chunk := limit / workers

for w := 0; w < workers; w++ {
    wg.Add(1)

    go func(worker int) {
        defer wg.Done()

        count := 0
        startNum := worker*chunk + 2
        endNum := (worker + 1) * chunk + 1

        for n := startNum; n <= endNum; n++ {
            if isPrime(n) {
                count++
            }
        }

        mu.Lock()
        total += int64(count)
        mu.Unlock()
    }(w)
}

wg.Wait()

elapsed := time.Since(start)

fmt.Println("Available CPUs:", runtime.NumCPU())
fmt.Println("GOMAXPROCS:", runtime.GOMAXPROCS(0))
fmt.Println("Workers:", workers)
fmt.Println("Primes found:", total)
fmt.Println("Execution time:", elapsed)
Enter fullscreen mode Exit fullscreen mode

}

OUTPUT :

D:\Scaling>go run main.go -procs=1
Available CPUs: 2
GOMAXPROCS: 1
Workers: 4
Primes found: 78498
Execution time: 388.075ms

D:\Scaling>go run main.go -procs=2
Available CPUs: 2
GOMAXPROCS: 2
Workers: 4
Primes found: 78498
Execution time: 205.7652ms

D:\Scaling>go run main.go -procs=3
Available CPUs: 2
GOMAXPROCS: 3
Workers: 4
Primes found: 78498
Execution time: 244.6589ms

D:\Scaling>go run main.go -procs=4
Available CPUs: 2
GOMAXPROCS: 4
Workers: 4
Primes found: 78498
Execution time: 272.6667ms

Top comments (0)