Why does building an AI model sometimes require thousands of accelerators, while running that model can take far fewer? And if both jobs use GPUs, does the network stay the same?
Training scales around the work required to learn a model. Inference scales around the work required to serve it at the required speed and volume. They can share hardware and networking technology, but their communication patterns and operating priorities differ.
There is an important qualification: one inference deployment can be much smaller than a large training job, while the total fleet serving a popular model can be enormous. “Inference is smaller” only makes sense after you define what you are comparing.
This article focuses on large language models, especially autoregressive transformer models. Image generation, recommendation systems, and other AI workloads have their own execution patterns. Even within language models, pretraining, fine-tuning, interactive chat, and offline batch processing are different sizing problems.
Training changes the model. Inference uses it.
During training, the system runs data through the model, calculates an error or loss, computes gradients through backpropagation, and updates model parameters. The process repeats across a large amount of data.
Ordinary inference uses those parameters to produce outputs without the backward pass and optimizer update. A conversation can build up context without changing the model’s weights.
Training therefore needs additional state: gradients, optimizer information, and saved or recomputed intermediate activations. Inference has its own working-memory needs, but it does not normally carry the full training state. Hugging Face’s GPU memory guide explains the training components.
| Design question | Training | Interactive LLM inference |
|---|---|---|
| Primary objective | Reach a quality target within a time and cost budget | Serve requests within latency, quality, and cost targets |
| Why add accelerators? | Fit training state and shorten the training run | Fit the model and working state, increase throughput, or reduce latency |
| Coordination | Often frequent communication within a distributed job | Ranges from independent replicas to tightly coupled distributed workers |
| Important measurements | Training progress, step time, scaling efficiency, recovery time | Time to first token, token delivery rate, tail latency, successful throughput |
| Typical failure response | Recover or restart affected work from saved state | Route around failed capacity and recover or retry affected requests |
Neither column describes every deployment. Offline inference may prioritize throughput over immediate response, and a small fine-tuning job can fit on one accelerator.
Why a training job can be so much larger

Conceptual illustration: a distributed training job repeatedly coordinates work across accelerators.
Three pressures drive training scale: memory capacity, total computation, and the deadline.
Consider a hypothetical 70-billion-parameter model. At two bytes per parameter, its weights alone occupy about 140 GB, using decimal units. That is not the complete inference memory requirement.
For comparison, one mixed-precision Adam training layout uses six bytes per parameter for weight copies, four for gradients, and eight for optimizer state. That totals 18 bytes per parameter, or roughly 1.26 TB for the same 70 billion parameters, before activations and temporary buffers. This is an illustrative layout, not a universal requirement: precision, optimizer choice, sharding, and offloading change the answer. The accounting follows Hugging Face’s memory breakdown.
Fitting those tensors is only the first hurdle. A configuration that can hold the model may still take too long to train it. More accelerators let the system process more work concurrently, provided the software and network keep them productive.
That last condition matters. Adding GPUs does not guarantee proportional speedup. Communication, uneven work, and synchronization can consume the time the additional hardware is supposed to save.
For a concrete historical example, Meta’s March 2024 infrastructure account describes two clusters with 24,576 H100 GPUs each, using its Llama 3 training cluster design. One uses RoCE Ethernet and the other InfiniBand, both with 400 Gbps endpoints. These are examples of a particular large-scale design, not minimum requirements for training an AI model. Meta’s engineering report.
Inference has two different kinds of scale
First, there is the serving group: the GPU or group of GPUs cooperating to run a model instance. A model might fit on one GPU, across GPUs in a server, or across multiple servers.
Second, there is the fleet: all the serving groups required to handle demand, plus capacity for failures, traffic bursts, and deployment changes. Replicas increase fleet capacity without necessarily joining every GPU into one synchronized job. vLLM’s scaling guidance describes single-GPU, tensor-parallel, and pipeline-parallel deployment choices.
Here is a deliberately simplified capacity example, using invented benchmark numbers:
- The service receives 100 requests per second.
- Each request produces an average of 200 output tokens.
- That creates demand for 20,000 output tokens per second.
- A tested eight-GPU serving group sustains 1,000 output tokens per second for the same prompt mix while meeting the required latency targets.
The arithmetic suggests 20 groups, or 160 GPUs, before additional operational headroom. These figures demonstrate the calculation; they are not performance claims for any product. Real sizing must also validate prompt processing, burst arrivals, concurrency, cache capacity, and failure scenarios.
This explains why the total inference fleet can exceed the resources used for a particular training run. It also explains why comparing lifetime inference spending with the cost of one training run requires a defined traffic volume and time period.
The model weights are not the entire memory budget
Transformer generation commonly maintains a key-value cache, or KV cache, to reuse attention information from preceding tokens. Longer contexts and more simultaneous requests can increase its memory footprint substantially. Cache behavior also depends on the model’s attention architecture and runtime. Hugging Face documents these cache strategies and tradeoffs.
A model that loads successfully may still lack enough memory for the desired concurrency. Smaller weight representations can help, but they do not eliminate cache, activation, and runtime requirements.
Inference itself contains different workloads

Conceptual illustration of inference at scale. Separate prefill and decode pools are one deployment option, not a requirement.
For autoregressive language models, two phases deserve separate attention:
Prefill processes the prompt and builds the initial attention state. With enough prompt work, it often makes strong use of matrix-compute capacity.
Decode generates subsequent tokens using that state. At smaller batch sizes it is often limited by memory bandwidth; with larger batches or different model architectures, compute and communication can become limiting instead. “Inference is memory-bound” is a useful starting hypothesis, not a law. NVIDIA’s explanation of disaggregated serving describes the differing resource demands.
The user experiences these phases through two related metrics: time to first token, which includes waiting before output begins, and time per output token, which describes ongoing delivery. Throughput alone can hide an unresponsive service. MLCommons explains why LLM benchmarks measure both latency and throughput.
Some deployments run both phases on the same workers. Others separate prefill and decode into independently sized pools. Separation introduces a new requirement: moving KV state from the prefill worker to the decode worker. Cross-server cache movement can make the network part of the request’s critical path. NVIDIA Dynamo documents this deployment pattern.
Disaggregation adds scheduling and transfer overhead. NVIDIA’s research finds its benefits depend on the workload and hardware, with stronger results for larger models and prefill-heavy traffic. It deserves measurement against an aggregated baseline, rather than adoption simply because it sounds more advanced. Beyond the Buzz: A Pragmatic Take on Inference Disaggregation.
Does the network design remain the same?
The physical topology can remain similar. The traffic engineering often changes.
Separate three communication domains before selecting the design:
| Domain | What it connects | Why it matters |
|---|---|---|
| Scale-up interconnect | Accelerators within a tightly connected system or rack-scale domain | Frequent model-parallel exchanges can use links such as NVLink or other platform interconnects |
| Scale-out compute fabric | GPU servers participating in distributed work | Carries collective operations, model-parallel traffic, and potentially KV transfers |
| Service, storage, and management networks | Clients, gateways, storage, orchestration, and operators | Carry requests, model loading, datasets, checkpoints, monitoring, and control traffic |
These domains may use separate physical networks or share infrastructure with isolation and capacity controls. NVLink and an Ethernet NIC are different layers of the system; quoting one link’s bandwidth does not describe the entire cluster.
Leaf-spine or Clos-style fabrics remain useful for both workloads. What changes is the required capacity between groups, the degree of oversubscription, and which flows must avoid contention. NVIDIA’s HGX reference architecture provides one example of separate network functions and a nonblocking fabric; it is a reference design, not a requirement for every inference installation. NVIDIA network logical architecture.
Training: communication is part of making progress
In data-parallel training, workers process different data and synchronize gradients. PyTorch’s DistributedDataParallel uses collective communication for this synchronization. Slow communication can extend training steps even when the GPUs have ample compute capacity. PyTorch’s distributed training documentation.
Other parallelism strategies distribute model operations, layers, or experts. They produce different traffic, so “training traffic” is not one uniform pattern.
For a large, tightly coupled job, the design priorities include high sustained bandwidth, predictable collective completion, and placement that keeps frequent communication on fast paths. Storage also needs to deliver data and absorb checkpoints without starving the job. Meta specifically describes topology-aware scheduling, routing optimization, and coordinated checkpoint storage in its cluster design account.
Inference: start with how much cooperation one request needs
For independent replicas that each fit within a server, there may be little cross-server GPU communication for each request. The network emphasis shifts toward request routing, availability, model distribution, and supporting services. A full training-class compute fabric may add little value to that workload.
If one model instance spans servers, the situation changes. Tensor parallelism can require repeated communication during generation, making link latency and congestion visible in token delivery. vLLM explicitly calls for fast inter-node communication when using cross-node tensor parallelism. vLLM’s networking guidance.
Mixture-of-experts models add another possibility: token work moves among distributed experts. Expert parallelism can create all-to-all traffic, even during inference. Sparse computation does not automatically mean a light network workload. vLLM’s expert-parallel deployment documentation describes these communication paths.
And a disaggregated deployment must account for prefill-to-decode cache transfers. The client’s short text request may trigger far more internal traffic than its size suggests.
What changes in the network plan?
The following are engineering implications of those workload patterns, rather than a universal topology prescription:
| Design choice | Training considerations | Inference considerations |
|---|---|---|
| Oversubscription | Contention can delay collective operations across a job | Potentially acceptable between independent serving groups; risky on critical model or KV-transfer paths |
| Placement | Keep communicating ranks near appropriate links | Keep model shards close and consider cache locality when routing requests |
| Bandwidth | Size for the chosen parallelism and collective traffic | Size for model exchanges, cache movement, model loading, and service traffic |
| Latency | Delayed synchronization increases step time | Delays can affect first-token time and token delivery, especially at the tail |
| Isolation | Protect long-running jobs from noisy neighbors | Protect serving latency from bulk transfers and other tenants |
| Resilience | Limit lost work and checkpoint recovery time | Preserve spare serving capacity and avoid cascading overload after failures |
Oversubscription means the combined demand from downstream links can exceed the bandwidth available upstream. It is a traffic-budget decision, not a synonym for a poorly designed network. The question is whether the workload regularly needs that shared capacity at the same time.
Ethernet or InfiniBand?
The choice does not map neatly to “inference versus training.” Meta demonstrates large training on both RoCE Ethernet and InfiniBand. Distributed inference can also benefit from RDMA-capable fabrics; simple independent replicas may not require them.
RDMA enables direct memory transfers with reduced CPU involvement. RoCE carries RDMA over Ethernet, but the label “Ethernet” alone says little about congestion control, routing, NIC configuration, or application performance. For Dynamo’s multi-node disaggregated deployment, its guidance calls for RDMA over InfiniBand or RoCE for the cache-transfer path. That does not make RDMA a prerequisite for all inference. Dynamo’s transfer guidance.
Can the same cluster do both?
Yes, suitable hardware can support both workloads. Sharing it efficiently requires more than changing the application image.
My recommendation is to decide which objectives must remain protected. Training can consume large blocks of capacity for extended periods. Interactive inference needs timely admission, predictable latency, and enough spare capacity to survive disruptions. A shared deployment needs scheduling and isolation policies that enforce those priorities.
Power and cooling also follow the actual hardware and utilization. An inference rack filled with the same dense accelerators does not automatically become an easy air-cooled workload because it stops doing backpropagation.
For planning, define the model and precision, prompt and output lengths, expected concurrency, latency objectives, and whether a request crosses GPUs or servers. Then benchmark representative traffic, including bursts and reduced-capacity conditions. GPU utilization alone cannot tell you whether a service meets its users’ needs.
The most revealing question is: How many accelerators need to communicate to make progress on this job or this request? That answer helps determine the tightly connected group. Demand and resilience determine how many groups the deployment needs. Together, they explain both the cluster’s size and the network it actually requires.


Leave a Reply