GPU racks and Ethernet switches beneath the title GB200 & GB300: Building the Ethernet Network.

If you’re planning a GB200 or GB300 deployment, it’s easy to spend most of your time looking at the GPUs. There are 72 of them in an NVL72 rack, after all. The compute numbers are impressive, the memory capacity is enormous, and there’s enough networking involved to make your cable-label printer nervous.

But once you start drawing the deployment, the questions get more familiar. How do we bring the servers online? Where does storage connect? Can we still reach a tray when its operating system stops responding? And who owns the DNS records?

Those are the questions I want to work through here. We’ll start with what changes between GB200 and GB300, then follow the connections out of the rack and into the rest of the data center.

We’re focusing on Ethernet for the GPU network between racks, as well as storage, management, and external services. Inside each NVL72 domain, NVLink still handles the high-bandwidth GPU connections. Both have a role in this design.

NVIDIA’s current GB300 enterprise reference architecture makes a useful starting point because it uses Spectrum-X Ethernet. GB300 also supports Quantum-X800 InfiniBand, so the choice of Ethernet here is intentional. It isn’t a restriction of the platform.

A Note on AI: This and other articles on bodiddely.com are based on my own notes, experiences, research, and opinions. I write the initial drafts myself, then use AI as an editing tool to proofread, improve clarity, and suggest changes. The ideas and perspectives remain my own. Also, if you haven’t already noticed, the images accompanying these articles are AI-generated.

Why focus on Ethernet?

For an enterprise network team, Ethernet offers some welcome familiarity. You already have ways to manage routing, allocate addresses, collect telemetry, and automate changes. Keeping those practices gives you somewhere solid to start while you learn the parts that are specific to GPU infrastructure.

The GB300 enterprise design builds on that foundation with a RoCE-capable, rail-optimized compute fabric. It also includes a converged north-south network and separate out-of-band management. We’ll unpack those terms shortly. For now, the useful idea is that GPU communication, storage access, and recovery access have different needs, even when they all use Ethernet.

The challenge is how GPUs use the network. They can produce bursts of synchronized traffic that expose congestion and uneven path use very quickly. Fast ports help, but getting consistent performance takes coordination between the switches, adapters, and application software.

InfiniBand offers another way to build that compute fabric, with its own management and operating model. NVIDIA’s cloud-partner architecture allows either option for the cluster interconnect. Tenant access and secure management remain Ethernet in that design.

If you’re considering Arista or Cisco switches, the Ethernet focus still applies. Just give that combination its own validation plan. Spectrum-X is an integrated reference solution, and another vendor’s fabric can differ in congestion handling, load balancing, supported features, and support arrangements. Those details are worth settling while the design is still a drawing.

GB200 versus GB300: what changes for the network team?

Let’s keep the comparison to NVL72 rack systems. The Blackwell product family is broad enough that it’s surprisingly easy to compare two specifications that describe different things.

Design pointGB200 NVL72GB300 NVL72Deployment consequence
Compute generationGrace CPUs with Blackwell GPUsGrace CPUs with Blackwell Ultra GPUsRequalify the hardware and software combination
Rack compute count72 GPUs, 36 Grace CPUs72 GPUs, 36 Grace CPUsSimilar rack-level inventory does not mean identical connectivity
Compute trays18 trays, four GPUs per tray18 trays, four GPUs per trayPreserve tray and GPU identity in the cabling plan
GPU memory, published rack specification13.4 TB HBM3EApproximately 20 TB HBM3EMore memory can change model placement, batching, and data movement
NVLink generationFifth generationFifth generationBoth retain a high-bandwidth rack-scale domain
Scale-out adapters in the cited reference configurationsFour 400G ConnectX-7 adapters per trayFour 800G ConnectX-8 adapters per trayThe nominal per-tray scale-out attachment doubles
North-south connectivityBlueField-3, configuration-dependentBlueField-3, configuration-dependentVerify card limits and port assignments rather than copying the older plan
GB200 and GB300 each have 72 GPUs and 36 Grace CPUs, with four 400G versus four 800G scale-out adapters per tray in the cited configurations.
Published reference-configuration adapter rates; rack illustrations are conceptual and not a count of physical cabinets.

The big networking change is the move from 400G ConnectX-7 adapters to 800G ConnectX-8 adapters in these reference configurations. More GPU memory also gives the application team different choices about model placement and batch size, which can change the traffic you need to carry.

The table uses NVIDIA’s published rack specifications and cloud-partner adapter configurations. When you move to an actual order, match them against your OEM’s configuration, including how it reports usable memory.

Both racks use 18 compute trays and nine NVLink switch trays in the configurations discussed here. Those NVLink switches are different devices from the Ethernet leaves outside the scale-up fabric.

Putting the adapter change into rack-level numbers makes the planning impact easier to see:

  • GB200: 18 trays × 4 adapters × 400 Gb/s = 28.8 Tb/s of nominal scale-out attachment per rack.
  • GB300: 18 trays × 4 adapters × 800 Gb/s = 57.6 Tb/s of nominal scale-out attachment per rack.

We’re adding up nominal port rates in one direction here. How much of that an application uses depends on the paths, workload, and rest of the system. You’ll still need to calculate the capacity across the fabric itself.

For a worked example of comparing endpoint capacity with the links carrying it, see my guide to calculating network oversubscription.

You’ll also see an approximately 130 TB/s aggregate NVLink figure for the rack. It describes a different interconnect and counts bandwidth differently. Keep the units and direction beside every number in your spreadsheet. A little lowercase “b” can cause a remarkably large budgeting conversation.

A few terms before we start connecting things

You don’t need to memorize this table. It’s here as a reference, especially for the terms that sound interchangeable until you’re trying to work out which cable goes where.

TermWhat it means here
NVL72A 72-GPU NVLink domain in the rack systems discussed here
Scale-upConnecting GPUs through the high-bandwidth local NVLink fabric
Scale-outConnecting compute resources across racks through Ethernet/RoCE or InfiniBand
NVLink / NVSwitchNVIDIA’s GPU interconnect and the switching technology used to build the NVLink fabric
RoCEv2RDMA over UDP/IP Ethernet, allowing RDMA traffic across a routed network
RDMARemote Direct Memory Access, moving data between registered memory regions with transport work offloaded to adapters
GPUDirect RDMAA supported direct data path between GPU memory and an RDMA-capable network adapter
SuperNIC / ConnectXNVIDIA adapter terminology and product family used for GPU scale-out connectivity
DPU / BlueFieldA programmable infrastructure processor used for networking, storage, and related services
RailA topology grouping that aligns corresponding GPU/NIC positions across compute nodes
PlaneA separate fabric path or topology used to distribute traffic and limit failures
Clos / leaf-spineA multistage topology providing multiple paths between endpoints
Bisection bandwidthCapacity available across a cut that divides the participating endpoints
ECN / CNP / PFCCongestion marking, RoCE congestion feedback, and per-priority Ethernet pause mechanisms
NCCLNVIDIA’s library for GPU collectives and point-to-point communication
BMC / OOBThe hardware management controller and its out-of-band access network
VRF / VLANSeparate routing tables and Layer 2 segments, respectively
BGP EVPN / VXLANA control plane and encapsulation commonly used to build segmented Ethernet/IP fabrics
SU / podA reference design’s repeatable capacity unit; its size varies by architecture and is not a Kubernetes pod

Follow the traffic out of the rack

I find it easier to design these environments by following a few ordinary activities: starting a server, loading a model, running a distributed job, saving its state, and recovering from a failure. Each activity tells us something about the network it needs.

Here’s how those responsibilities fit together:

Network roleMain endpoints and trafficDesign priority
NVLink scale-upGPUs and NVLink switch traysCorrect domain formation and local GPU communication
GPU scale-outConnectX adapters and compute leaves/spinesPredictable RDMA performance and path diversity
StorageCompute/DPU interfaces and file, object, or block servicesRead throughput, checkpoints, metadata, and recovery
In-band managementHost operating systems, DPUs, controllers, and schedulersProvisioning and dependable control-plane access
Out-of-band managementBMCs, switch management, power systemsRecovery when the host or production network fails
Customer and external servicesBorder routers, clients, registries, DNS, identity servicesControlled reachability and adequate ingress/egress capacity
NVLink inside a GPU rack, with connections representing GPU scale-out, storage, in-band management, and out-of-band management.
Conceptual illustration. Logical roles can share physical infrastructure; connections show responsibilities rather than exact wiring.

Several of these roles can share switches. In the enterprise reference design, the north-south fabric brings multiple services together. At larger scales, my preference is to give GPU communication its own backend and preserve an independent recovery path. Then we can decide where sharing equipment makes sense without accidentally making every service depend on the same pair of switches.

Before traffic reaches an Ethernet leaf, the GPUs have a substantial network of their own. NVLink carries GPU-to-GPU communication through the rack’s NVLink switch trays. In the design we’re discussing, that gives each rack its own 72-GPU scale-up domain.

There’s a useful detail here for the network team: managing that domain still involves IP connectivity. NVOS, Fabric Manager, NVLink Subnet Manager, and NMX services configure and observe the fabric. IMEX supports GPU memory mapping across operating-system instances and communicates between compute nodes over TCP/IP.

So even when the application data travels over NVLink, part of the machinery that makes it work depends on ordinary networking.

During planning, work out who owns those services and how their health appears in the monitoring system. If an application can’t use the full rack, you’ll want to know whether the NVLink domain is ready before spending the afternoon examining Ethernet counters.

2. GPU scale-out: where Ethernet earns its keep

Once a workload spreads across racks, the backend becomes a much larger part of the performance story. It carries operations such as AllReduce, AllGather, and All-to-All, allowing the GPUs to combine results and exchange data. The application’s parallelism strategy determines how much of each you see.

NVIDIA’s GB300 enterprise example uses two planes, splitting each GPU’s nominal 800G attachment into two 400G links, one per plane. Rail alignment groups corresponding GPU positions across trays. This is a specific topology with endpoint support, not a suggestion to put every GPU port into an ordinary LACP bond.

Now count the connections. With 72 GPUs and two 400G links per GPU, you’re looking at 144 host-facing link endpoints per rack, before adding any switch-to-switch connections. Eight racks bring that to 1,152. That’s where a diagram starts turning into a serious discussion about switch ports, breakout cables, fiber routes, and spares.

The second plane gives traffic another path, but each plane supplies half the nominal attachment in this example. Losing one therefore leaves less capacity. How the application recovers also depends on the supported NIC and software behavior, so include a plane failure in testing rather than assuming an active RDMA connection moves without interruption.

To make the fabric behave consistently, work through these items as a connected system:

  • End-to-end MTU and RDMA traffic classification.
  • Switch ECN behavior and endpoint congestion response.
  • PFC priorities, headroom, and watchdog behavior where the selected design uses PFC.
  • Routing and load distribution under both normal and degraded conditions.
  • GPU/NIC locality, driver and firmware compatibility, and the communication-library configuration.

Arista’s deployment guide provides a concrete RoCE QoS example and explicitly notes that ECN thresholds vary by deployment. Cisco’s AI/ML blueprint likewise treats congestion management and visibility as design requirements. Their configuration examples are starting points for their stated platforms, not interchangeable recipes.

If performance falls short, my GPU fabric performance troubleshooting guide walks through five common causes, the evidence to collect, and practical fixes.

3. Storage: the other high-bandwidth conversation

Imagine the cluster starting several jobs at once. Models load, datasets move, and local caches fill. Later, those same jobs save checkpoints, potentially at roughly the same time. Both moments can put substantial pressure on storage, but in different directions and with different latency requirements.

That’s why I like to bring the storage team into the design early. We need to agree on the behavior we’re sizing for, including the less convenient moments, such as restarting jobs after a failure.

The rack-scale guide describes local NVMe caching and external shared storage. Local cache reduces repeated network reads; it does not eliminate cold starts or the need to persist checkpoints outside the failed component’s scope.

In our Ethernet design, compute reaches storage over the north-south network. The storage system may have another network for its own replication or backend traffic. Showing both on the drawing helps everyone see where capacity and failures can affect a write.

Let’s put a number on that. Suppose the workloads need to write a combined 8 TB checkpoint within two minutes:

8 TB written within 120 seconds requires about 66.7 GB/s, or 533 Gb/s, of aggregate application write throughput.

That calculation uses decimal units and counts only the payload. We still have to allow for network overhead, storage protection, metadata, other jobs, and operation with a failed component. It gives us a useful starting target to discuss with the storage team, rather than choosing a port speed and hoping it works out.

Watch the DPU limits, too. The current enterprise hardware chapter describes a BlueField-3 B3240 with two 400G ports but approximately 400G aggregate card bandwidth. Two connectors do not promise 800G of usable storage throughput. Reconcile the exact card, host interface, operating mode, and measured storage performance.

4. In-band management: the network that builds the cluster

Before the cluster runs its first job, someone has to install the operating systems, distribute software, register the nodes, and make the scheduler happy. That’s the work supported by in-band management. Once the cluster is running, this network continues carrying configuration, control, and telemetry traffic.

NVIDIA’s DGX terminology separates internalnet for control-plane provisioning and management, dgxnet for compute provisioning and management, storagenet, computenet, and externalnet. These names describe roles in that stack; they do not mandate identical VLAN numbers in every OEM installation. DGX network overview

Give those control services a home that can start independently of the GPU cluster. Otherwise, you can end up needing the cluster to run the service that brings up the cluster. It’s an interesting puzzle, but preferably one you solve on a whiteboard.

This is also a good time to check the software you’re distributing. Grace uses Arm CPUs, so provisioning images, monitoring agents, storage clients, security tools, and containers all need the appropriate architecture support. Finding an x86-only dependency before installation is a much more pleasant experience than finding it halfway through bring-up.

5. Out-of-band management: the way back in

Now picture a compute tray whose operating system stops responding. You need its hardware controller, perhaps a console session, and possibly a power operation. That’s when out-of-band management becomes the most useful network in the room.

Its inventory includes compute BMCs, DPU management, Ethernet switches, NVLink switch management, and the relevant rack power and cooling interfaces.

NVIDIA’s enterprise design uses dedicated management switches and low-speed device ports. Those ports carry a different responsibility from the high-speed workload links: they must remain useful during failure. Enterprise management network

The question I use when reviewing this part of the design is simple: can we still reach the device when its host OS, data NIC, or compute fabric is unavailable?

Follow that question all the way back to the operator. The access path, authentication, management uplinks, and recovery credentials all matter. So does having a way to recover the management switches themselves.

Be precise about the redundancy you have. A management VLAN still depends on the switch carrying it. A BMC with one physical connection still has one access link, even when the rack contains two OOB switches. Writing that down helps the on-call team understand when remote recovery is possible and when someone needs to visit the rack.

Bring the facilities team into this conversation too. Power and cooling telemetry can help explain a rack-wide problem, but access to building-control systems needs agreed boundaries and ownership. You want everyone to know their part before an alarm starts making the decisions for them.

6. External connectivity: where the cluster meets everyone else

Eventually, users need to reach the applications, administrators need access, and the platform needs services beyond its own racks. Those connections normally pass through the north-south fabric toward the customer edge. That gives us a clear place to manage routes to registries, identity systems, corporate networks, and permitted Internet destinations.

Inference adds an interesting wrinkle. A small client request can trigger a much larger exchange inside the cluster. If prefill and decode run on separate workers, for example, transferring the attention state held in the KV cache adds another traffic path to consider.

Walk through a request with the application team and draw where the data actually goes. That conversation often tells you more about capacity requirements than labeling every arrow “north-south” or “east-west.”

Where DNS, DHCP, and the other essentials belong

This is the part of the design that tends to look reassuringly familiar. DNS, DHCP, time services, and identity still do the jobs you expect. The difference is how many systems can need them at once, and how much work stops when they aren’t available.

Here’s how I’d organize the conversation with the teams providing those services. These are planning recommendations; the exact placement depends on your environment.

ServicePlacement and reachabilityWhat to verify
DNSRedundant cluster-local or enterprise resolvers reached through management/service routingForward and reverse records, zone ownership, split views, and operation during upstream loss
DHCP and boot servicesProvisioning services reachable from intended boot segments, directly or through relayMAC reservations, relay behavior, boot options, and authority for each subnet
NTPReliable time sources reachable by hosts and management componentsClock alignment for certificates, logs, and incident correlation
Identity and PKIControlled service paths to identity providers, certificate services, and secrets systemsRenewal, authorization, and recovery when upstream services are unavailable
Registries and repositoriesLocal mirrors or caches where practical, with managed external accessArm-compatible images, image-pull bursts, package availability, and trust
Telemetry and logsCollectors reachable from each required management domainRetention, consistent timestamps, and collection through a fabric incident
IPAM and source of truthManagement infrastructure integrated with provisioning and network automationRack, tray, GPU/NIC, MAC, address, rail, plane, and switch-port relationships
Shared services supporting a GPU cluster: DNS, DHCP and boot, time, identity and certificates, registries and images, and telemetry and logs.
Conceptual illustration of foundational services supporting cluster provisioning and operation.

Take DHCP as an example. Your provisioning system may use it to identify a tray and tell it where to boot. If that tray sits on a different routed segment from the server, the relay configuration becomes part of the boot process. Agree on which system owns each scope, especially when the cluster manager also provides DHCP. Meanwhile, loopbacks, routed links, and critical services still need carefully allocated addresses.

DNS deserves a similar conversation. You can delegate a cluster subdomain to local services or manage its records in enterprise DNS. What matters is that both teams understand the arrangement and know what happens if the upstream connection disappears.

I would test these dependencies during bring-up, while the environment is small enough to understand easily. Having hundreds of nodes report the same name-resolution error gives you more logs, but very little additional insight.

Routing, segmentation, and addressing at scale

Once the traffic roles are clear, the routing plan becomes much easier to discuss. We can identify which services need to talk, where tenants need isolation, and which paths should remain available during a failure. Those decisions lead to the VRFs, provisioning segments, storage networks, service routes, loopbacks, and point-to-point links.

NVIDIA’s Mission Control north-south guide describes an EVPN multihoming VXLAN architecture with symmetric routing and head-end replication. That is a north-south reference design; it is not evidence that the GPU backend must use the same overlay.

For the backend, I’d keep the routing design as straightforward as the validated solution allows. An overlay can be useful when it addresses a real requirement, but it also gives you another layer to account for in MTU, policy, and troubleshooting.

Address planning is much easier before racks start arriving. Reserve room for expansion and compare the proposed ranges with corporate networks, storage appliances, container networks, and remote-access pools. Reference designs are helpful examples, but their addresses don’t know what you already use.

When you build the access policy, record the source, destination, protocol, and reason for each connection. Include the less visible dependencies, such as NVLink control services and the communication library’s bootstrap connections. This gives the network and systems teams something concrete to review together when a job can move data but still struggles to initialize.

Turn the reference design into a buildable port plan

At some point, the architecture has to become an order and a set of installation instructions. This is where being comfortably specific pays off.

Start by agreeing on what one expansion block contains. Reference designs use terms like “scalable unit” and “pod” differently, so write down the rack and node counts behind the name. Then build a small set of documents that stays together as the deployment grows:

  1. A bill of materials tied to an exact document and software revision.
  2. A point-to-point cable schedule, including rail and plane identity.
  3. A switch-port budget for endpoints, inter-switch links, services, spares, and growth.
  4. Normal and degraded bandwidth calculations.
  5. Rack elevation, power, cooling, fiber-path, and maintenance-access drawings.

Keep ASIC bandwidth, physical cages, breakout interfaces, optical channels, and cable assemblies as separate entries in the plan. With breakout connections, counting the holes in the front of the switch won’t tell you how many endpoint links you can build.

There’s a practical reason to reconcile the documents carefully: the current NVIDIA enterprise pages use different switch port presentations and contain some inconsistent component wording. Tie the approved bill of materials to the exact topology and revision you’re deploying. It saves the installation team from trying to resolve the discrepancy with a box of optics in hand.

Factory inventory helps here too. Import serial numbers and MAC addresses into your source of truth, then associate them with trays, adapters, addresses, and switch ports. NVIDIA’s deployment guide calls for that inventory and a point-to-point cabling plan. See more at Factory inventory and preparation requirements

Good labels and accurate records feel like small details until you’re replacing one cable among a thousand. That’s an excellent time to appreciate the person who insists on updating the drawing.

Training, inference, and placement change the network requirement

Before settling the final capacity plan, spend some time with the people who run the workloads. “Training” and “inference” are useful starting descriptions, but neither tells you enough on its own.

A job that fits inside one rack may keep most of its GPU communication on NVLink. Spread it across racks and the Ethernet backend becomes part of that conversation. Model size, batch size, and the chosen parallelism strategy all influence how much traffic crosses the boundary.

For training, look at representative collectives and checkpoint behavior. For inference, ask whether you’re serving independent replicas or splitting a model or execution stages across racks. Independent rack-local replicas can have very different backend needs from a distributed mixture-of-experts workload.

GB300’s larger memory gives you more placement options. It may let a workload stay local, or the application team may use the headroom for larger batches and more demanding models. Either is reasonable. The network plan needs to reflect the choice.

Bring that information into the scheduler as well. When scheduling respects locality and failure domains, the application has a better chance of benefiting from the topology you spend so much time designing.

I explore those workload differences in more detail in the article on training versus inference network design.

How do we know it’s ready?

The first successful job is a good moment. Enjoy it. Then find out whether it still behaves well with another job running, storage busy, and one of the expected failure scenarios in play.

I like to build acceptance around a series of practical questions. Can we recover a host? Can we provision reliably? Does a workload scale across racks? Can storage meet the checkpoint target? Agreeing on the answers and measurements beforehand keeps “looks good to me” from becoming the entire test plan.

GateEvidence to collect
Physical readinessCorrect cabling, optical compatibility, negotiated speed, FEC behavior, cooling, and power
Recovery accessBMC and console reachability with the host or data path unavailable
Foundational servicesDNS, boot services, time, identity, registries, and control-plane failover
Rack readinessExpected GPU inventory, healthy NVLink domain, and required control services
Scale-out readinessGPU-buffer RDMA bandwidth and latency across node pairs, rails, planes, and racks
Application readinessRepresentative collectives and workloads at increasing scale and concurrency
Storage readinessCold reads, checkpoint bursts, restart recovery, and concurrent workload interference
Failure readinessDefined behavior during link, leaf, plane, service, and storage-path failures
Deployment testing progresses through hardware access, services, fabric verification, workloads, and failure testing.
Conceptual illustration. Validate normal operation and recovery before expanding the deployment.

Keep workload sizes and placement consistent when comparing results. Collect link errors, queue behavior, congestion signals, RDMA retries, collective timing, storage latency, and application performance together. That gives you a baseline you can return to after an upgrade or during an incident.

The identifiers matter as much as the counters. If a slow job points to a GPU, you want to follow that GPU to its tray, NIC, rail, plane, and switch port without opening five spreadsheets. This is where the inventory work starts paying for itself.

Record the tested software combination too: switch OS, adapter and DPU firmware, host drivers, CUDA, NCCL, NVOS, and management services. When it’s time to upgrade, test a bounded part of the environment first and compare it with that baseline. It gives you a much clearer answer than trying to remember whether the cluster “feels slower.”

Bringing it all together

What I find interesting about GB200 and GB300 deployments is how much they ask us to connect familiar disciplines. The GPU fabric gets us into high-speed RDMA and rack-scale communication. Getting the platform into service brings us right back to storage, routing, provisioning, naming, and a reliable way to reach a broken host.

GB300 raises the scale-out capacity and gives applications more memory to work with. The design still succeeds through the same careful conversations: what the workload needs, which path carries it, what happens when that path fails, and how the operator finds the problem.

Get those conversations into the design early, and bringing the racks online becomes a much more manageable job. You’ll still have plenty of cables to label. At least you’ll know why each one is there.

If you’re working through the details of your own deployment, these guides take a closer look at performance, capacity, and keeping the inventory useful.

NVIDIA Documentation

Miscellaneous Documentation

Leave a Reply

Quote of the week

“If you change the way you look at things, the things you look at change.”

~ Dr. Wayne Dyer

Discover more from Bodiddely

Subscribe now to keep reading and get access to the full archive.

Continue reading