DEV Community

Cover image for Why Does HPC Need InfiniBand?
Muhammad Zubair Bin Akbar
Muhammad Zubair Bin Akbar

Posted on Originally published at hpcpulse.substack.com

Why Does HPC Need InfiniBand?

When people first learn about HPC clusters, they often notice something unusual.

The compute nodes don't just have Ethernet.

Many HPC clusters also have InfiniBand.

So a natural question is:

Why?

If Ethernet already connects computers together, why do HPC clusters need another network?

The answer becomes clearer when you look at what HPC applications actually do.

HPC applications communicate a lot

Consider an MPI application running across 32 compute nodes.

The application may constantly exchange data between those nodes.

For example:

Node 1  <---->  Node 2
  ↕              ↕
Node 3  <---->  Node 4
Enter fullscreen mode Exit fullscreen mode

The application isn't just sending an occasional file or SSH connection.

The compute nodes can be communicating continuously while the application is running.

That makes network performance extremely important.

It's not just about bandwidth

A common first thought is:

"Just give the cluster a faster network."

But HPC networking is about more than bandwidth.

Two important factors are:

Bandwidth

How much data can be transferred per second.

Latency

How long it takes for a message to travel between nodes.

For many HPC workloads, both matter.

A network can have high bandwidth but still have latency that isn't ideal for tightly coupled workloads.

Where InfiniBand comes in

InfiniBand was designed for high performance communication.

It provides:

  • High bandwidth
  • Low latency
  • RDMA
  • Efficient communication between nodes

One particularly important feature is RDMA.

RDMA allows data to be transferred directly between the memory of different systems with minimal involvement from the CPU and operating system.

That can reduce overhead during communication.

Why does that matter?

Imagine an MPI application running across hundreds of nodes.

The application might repeatedly exchange relatively small messages between processes.

If every communication requires significant CPU and operating system involvement, that overhead can become important.

With technologies such as RDMA, the network can handle communication more efficiently.

This is one reason high performance networks are so important in HPC.

InfiniBand isn't the only option

InfiniBand isn't the only technology used for HPC networking.

You can also find:

  • High speed Ethernet
  • RoCE
  • Other specialised interconnects

Modern Ethernet can provide very high performance, and technologies such as RoCE can provide RDMA capabilities over Ethernet.

So the question isn't simply:

"InfiniBand or Ethernet?"

It's really about what the workload and cluster architecture require.

The bigger picture

An HPC cluster is a collection of computers working together.

The CPUs do the computation.

The memory holds the data being processed.

The storage provides the data.

And the network allows the nodes to communicate.

If the application spends a lot of time waiting for data from another node, having powerful CPUs won't solve the problem.

The network becomes part of the application's performance.

That's why HPC networking is such an important part of cluster design.

And that's also why you'll often hear technologies such as:

InfiniBand → RDMA → MPI → UCX

when talking about HPC communication.

Understanding how these pieces fit together is much more useful than simply knowing that "HPC uses InfiniBand."

Top comments (0)