<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rahul Kumar</title>
    <description>The latest articles on DEV Community by Rahul Kumar (@coderahul1).</description>
    <link>https://dev.to/coderahul1</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4157616%2Fd07c9ea9-a025-45bd-ad44-8a499b776401.jpg</url>
      <title>DEV Community: Rahul Kumar</title>
      <link>https://dev.to/coderahul1</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/coderahul1"/>
    <language>en</language>
    <item>
      <title>I ran a small Go experiment to see how CPU cores affect performance.</title>
      <dc:creator>Rahul Kumar</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:53:31 +0000</pubDate>
      <link>https://dev.to/coderahul1/i-ran-a-small-go-experiment-to-see-how-cpu-cores-affect-performance-31hi</link>
      <guid>https://dev.to/coderahul1/i-ran-a-small-go-experiment-to-see-how-cpu-cores-affect-performance-31hi</guid>
      <description>&lt;p&gt;&lt;strong&gt;## 1. Understanding CPU cores at the system level&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;At the system level, the CPU is responsible for executing instructions. We have different numbers of CPU cores, such as 1 core, 2 cores, 4 cores, 8 cores, etc.&lt;/p&gt;

&lt;p&gt;Each logical CPU can execute one thread at a time.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;1 logical CPU → 1 thread executing at a time.&lt;/li&gt;
&lt;li&gt;2 logical CPUs → 2 threads executing simultaneously.&lt;/li&gt;
&lt;li&gt;4 logical CPUs → 4 threads executing simultaneously.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, having more CPU cores does not automatically mean faster execution. It depends on how the application utilizes those cores and whether the workload can be executed in parallel.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Concurrency vs. Parallelism
&lt;/h2&gt;

&lt;p&gt;Concurrency: Multiple tasks are in progress, but they don't necessarily execute simultaneously.&lt;/p&gt;

&lt;p&gt;For example, on a single CPU core, the operating system can switch between multiple tasks. To humans, they may appear to be executing simultaneously because the switching happens very quickly.&lt;/p&gt;

&lt;p&gt;Parallelism: Two or more tasks execute at the same time using separate execution resources, such as different CPU cores.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Task 1 executes on Core 1.&lt;/li&gt;
&lt;li&gt;Task 2 executes on Core 2.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both tasks are executing simultaneously.&lt;/p&gt;

&lt;p&gt;Remember: Concurrency is about managing multiple tasks at once, whereas parallelism is about executing multiple tasks at the same instant.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. How Go handles task execution
&lt;/h2&gt;

&lt;p&gt;Go provides its own concurrency mechanism through goroutines and a runtime scheduler.&lt;/p&gt;

&lt;h3&gt;
  
  
  Goroutines
&lt;/h3&gt;

&lt;p&gt;A goroutine is a lightweight execution unit managed by the Go runtime.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;1 task → potentially 1 goroutine.&lt;/li&gt;
&lt;li&gt;2 independent tasks → potentially 2 goroutines.&lt;/li&gt;
&lt;li&gt;10,000 independent tasks → potentially 10,000 goroutines.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, the number of goroutines depends on how we design our application. We don't necessarily need one goroutine for every individual job.&lt;/p&gt;

&lt;p&gt;Go can efficiently manage thousands of goroutines without creating an equivalent number of OS threads.&lt;/p&gt;

&lt;h3&gt;
  
  
  OS Threads
&lt;/h3&gt;

&lt;p&gt;OS threads are execution contexts managed by the operating system.&lt;/p&gt;

&lt;p&gt;They are responsible for executing instructions on available logical CPUs.&lt;/p&gt;

&lt;p&gt;Go's runtime manages the relationship between goroutines and OS threads.&lt;/p&gt;

&lt;h3&gt;
  
  
  GOMAXPROCS
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;GOMAXPROCS&lt;/code&gt; controls how many goroutines can execute Go code simultaneously.&lt;/p&gt;

&lt;p&gt;By default, Go generally configures this according to the available CPU resources.&lt;/p&gt;

&lt;p&gt;For example, if our system has 2 available logical CPUs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Available CPUs: 2
GOMAXPROCS: 2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This means Go can execute Go code in parallel across up to 2 logical CPUs.&lt;/p&gt;

&lt;p&gt;Important: GOMAXPROCS is not the number of OS threads Go creates. Go's runtime can create and manage more OS threads as needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Relationship between CPU cores, OS threads, and goroutines
&lt;/h2&gt;

&lt;p&gt;Suppose our system has only 2 logical CPUs.&lt;/p&gt;

&lt;p&gt;We can still have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;10 goroutines.&lt;/li&gt;
&lt;li&gt;10 OS threads.&lt;/li&gt;
&lt;li&gt;100 goroutines.&lt;/li&gt;
&lt;li&gt;Even thousands of goroutines.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But having more threads or goroutines does not mean we have more physical CPU cores.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Available logical CPUs: 2
Goroutines: 4
GOMAXPROCS: 2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Go can execute up to 2 goroutines simultaneously, while the remaining goroutines wait for their turn or wait for some operation to complete.&lt;/p&gt;

&lt;p&gt;The Go runtime schedules goroutines onto OS threads, and the operating system schedules those threads onto available logical CPUs.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Does having more threads make execution faster?
&lt;/h2&gt;

&lt;p&gt;Sometimes yes, but not always.&lt;/p&gt;

&lt;p&gt;It depends on the workload.&lt;/p&gt;

&lt;h3&gt;
  
  
  Case 1: CPU-intensive tasks
&lt;/h3&gt;

&lt;p&gt;Suppose we are calculating prime numbers within a specific range.&lt;/p&gt;

&lt;p&gt;This task requires continuous CPU computation.&lt;/p&gt;

&lt;p&gt;If we have 2 logical CPUs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;With GOMAXPROCS=1, only one goroutine can execute Go code at a time.&lt;/li&gt;
&lt;li&gt;With GOMAXPROCS=2, two goroutines can execute Go code simultaneously.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Therefore, the second configuration can complete the calculation faster because it utilizes both available CPUs.&lt;/p&gt;

&lt;p&gt;However, increasing GOMAXPROCS beyond the available CPU capacity does not provide additional physical processing power.&lt;/p&gt;

&lt;h3&gt;
  
  
  Case 2: I/O-intensive tasks
&lt;/h3&gt;

&lt;p&gt;Suppose our backend server is handling multiple requests, and some requests require:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Fetching data from a database.&lt;/li&gt;
&lt;li&gt;Reading files from disk.&lt;/li&gt;
&lt;li&gt;Calling external APIs.&lt;/li&gt;
&lt;li&gt;Waiting for network responses.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;During these operations, a goroutine may become blocked or suspended while waiting for the result.&lt;/p&gt;

&lt;p&gt;Instead of keeping the CPU idle, Go can schedule another runnable goroutine.&lt;/p&gt;

&lt;p&gt;This allows the application to continue processing other requests while some operations are waiting.&lt;/p&gt;

&lt;p&gt;This is one of the major reasons backend applications use many goroutines even when they have relatively few CPU cores.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Why can excessive threads introduce overhead?
&lt;/h2&gt;

&lt;p&gt;Suppose we have 2 logical CPUs but many runnable OS threads.&lt;/p&gt;

&lt;p&gt;The operating system must schedule those threads onto the available CPUs.&lt;/p&gt;

&lt;p&gt;When switching between threads, the operating system may need to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Save the execution state of the currently running thread.&lt;/li&gt;
&lt;li&gt;Restore the execution state of another thread.&lt;/li&gt;
&lt;li&gt;Update scheduling information.&lt;/li&gt;
&lt;li&gt;Handle CPU cache effects and other scheduling costs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is called context switching.&lt;/p&gt;

&lt;p&gt;Context switching is not the same as moving a thread to disk or physically removing it from memory. Threads generally remain represented in memory while they are waiting to execute.&lt;/p&gt;

&lt;p&gt;If the system performs excessive scheduling and context switching, it can introduce overhead and reduce performance.&lt;/p&gt;

&lt;p&gt;However, more threads do not automatically mean more overhead. The impact depends on how many threads are runnable, what they are doing, and how the workload is structured.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Why does Go use more OS threads when the system has limited CPU cores?
&lt;/h2&gt;

&lt;p&gt;Go's runtime manages goroutines and OS threads separately.&lt;/p&gt;

&lt;p&gt;For CPU-intensive workloads, Go generally uses its available parallel execution capacity efficiently.&lt;/p&gt;

&lt;p&gt;For I/O-intensive workloads, some goroutines may wait for operations to complete.&lt;/p&gt;

&lt;p&gt;Go can schedule other runnable goroutines while those operations are waiting.&lt;/p&gt;

&lt;p&gt;For certain blocking system calls, an OS thread may become blocked, and the Go runtime can arrange for other goroutines to continue executing using other available threads.&lt;/p&gt;

&lt;p&gt;For network I/O, Go commonly uses runtime network polling to avoid requiring one blocked OS thread for every waiting network connection.&lt;/p&gt;

&lt;p&gt;This helps Go handle a large number of concurrent requests without requiring an equivalent number of OS threads.&lt;/p&gt;

&lt;h1&gt;
  
  
  8. Practical Experiment: Measuring CPU Parallelism in Go
&lt;/h1&gt;

&lt;p&gt;We created a Go program to calculate prime numbers within a specified range using multiple goroutines.&lt;/p&gt;

&lt;p&gt;The program used:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;2 available logical CPUs.&lt;/li&gt;
&lt;li&gt;4 goroutines.&lt;/li&gt;
&lt;li&gt;Different &lt;code&gt;GOMAXPROCS&lt;/code&gt; configurations.&lt;/li&gt;
&lt;li&gt;Execution time as the performance measurement.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Experiment results
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;GOMAXPROCS&lt;/th&gt;
&lt;th&gt;Workers&lt;/th&gt;
&lt;th&gt;Primes found&lt;/th&gt;
&lt;th&gt;Execution time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;78,498&lt;/td&gt;
&lt;td&gt;388.075 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;78,498&lt;/td&gt;
&lt;td&gt;205.7652 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;78,498&lt;/td&gt;
&lt;td&gt;244.6589 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;78,498&lt;/td&gt;
&lt;td&gt;272.6667 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  What did we observe?
&lt;/h3&gt;

&lt;p&gt;When GOMAXPROCS = 1:&lt;/p&gt;

&lt;p&gt;Only one goroutine could execute Go code at a time.&lt;/p&gt;

&lt;p&gt;Execution time: approximately 388 ms.&lt;/p&gt;

&lt;p&gt;When GOMAXPROCS = 2:&lt;/p&gt;

&lt;p&gt;Two goroutines could execute Go code simultaneously across the available CPUs.&lt;/p&gt;

&lt;p&gt;Execution time: approximately 206 ms.&lt;/p&gt;

&lt;p&gt;The workload completed substantially faster.&lt;/p&gt;

&lt;p&gt;When GOMAXPROCS = 3 or 4:&lt;/p&gt;

&lt;p&gt;The configuration permitted more parallel execution capacity, but our machine still had only 2 available logical CPUs.&lt;/p&gt;

&lt;p&gt;The measured execution time increased compared with the 2-CPU configuration.&lt;/p&gt;

&lt;p&gt;This may involve scheduling overhead and measurement variability. Additional parallelism does not guarantee improved performance.&lt;/p&gt;

&lt;p&gt;We should repeat the benchmark multiple times and compare averages before drawing a firm conclusion.&lt;/p&gt;

&lt;h3&gt;
  
  
  Conclusion from the experiment
&lt;/h3&gt;

&lt;p&gt;More CPU cores can improve the performance of CPU-intensive workloads when the application can utilize them effectively.&lt;/p&gt;

&lt;p&gt;However, more threads, goroutines, or parallel execution slots do not automatically make an application faster.&lt;/p&gt;

&lt;p&gt;We must identify the workload, available hardware resources, and potential bottlenecks.&lt;/p&gt;

&lt;p&gt;CODE:&lt;/p&gt;

&lt;p&gt;package main&lt;/p&gt;

&lt;p&gt;import (&lt;br&gt;
    "flag"&lt;br&gt;
    "fmt"&lt;br&gt;
    "runtime"&lt;br&gt;
    "sync"&lt;br&gt;
    "time"&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;func isPrime(n int) bool {&lt;br&gt;
    if n &amp;lt; 2 {&lt;br&gt;
        return false&lt;br&gt;
    }&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;for i := 2; i*i &amp;lt;= n; i++ {
    if n%i == 0 {
        return false
    }
}
return true
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;}&lt;/p&gt;

&lt;p&gt;func main() {&lt;br&gt;
    procs := flag.Int("procs", 1, "Number of CPUs to use")&lt;br&gt;
    flag.Parse()&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;runtime.GOMAXPROCS(*procs)

const limit = 1000000
const workers = 4

var wg sync.WaitGroup
var total int64

var mu sync.Mutex

start := time.Now()

chunk := limit / workers

for w := 0; w &amp;lt; workers; w++ {
    wg.Add(1)

    go func(worker int) {
        defer wg.Done()

        count := 0
        startNum := worker*chunk + 2
        endNum := (worker + 1) * chunk + 1

        for n := startNum; n &amp;lt;= endNum; n++ {
            if isPrime(n) {
                count++
            }
        }

        mu.Lock()
        total += int64(count)
        mu.Unlock()
    }(w)
}

wg.Wait()

elapsed := time.Since(start)

fmt.Println("Available CPUs:", runtime.NumCPU())
fmt.Println("GOMAXPROCS:", runtime.GOMAXPROCS(0))
fmt.Println("Workers:", workers)
fmt.Println("Primes found:", total)
fmt.Println("Execution time:", elapsed)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;}&lt;/p&gt;

&lt;p&gt;OUTPUT :&lt;/p&gt;

&lt;p&gt;D:\Scaling&amp;gt;go run main.go -procs=1&lt;br&gt;
Available CPUs: 2&lt;br&gt;
GOMAXPROCS: 1&lt;br&gt;
Workers: 4&lt;br&gt;
Primes found: 78498&lt;br&gt;
Execution time: 388.075ms&lt;/p&gt;

&lt;p&gt;D:\Scaling&amp;gt;go run main.go -procs=2&lt;br&gt;
Available CPUs: 2&lt;br&gt;
GOMAXPROCS: 2&lt;br&gt;
Workers: 4&lt;br&gt;
Primes found: 78498&lt;br&gt;
Execution time: 205.7652ms&lt;/p&gt;

&lt;p&gt;D:\Scaling&amp;gt;go run main.go -procs=3&lt;br&gt;
Available CPUs: 2&lt;br&gt;
GOMAXPROCS: 3&lt;br&gt;
Workers: 4&lt;br&gt;
Primes found: 78498&lt;br&gt;
Execution time: 244.6589ms&lt;/p&gt;

&lt;p&gt;D:\Scaling&amp;gt;go run main.go -procs=4&lt;br&gt;
Available CPUs: 2&lt;br&gt;
GOMAXPROCS: 4&lt;br&gt;
Workers: 4&lt;br&gt;
Primes found: 78498&lt;br&gt;
Execution time: 272.6667ms&lt;/p&gt;

</description>
      <category>go</category>
      <category>backend</category>
      <category>performance</category>
      <category>systemdesign</category>
    </item>
  </channel>
</rss>
