Scale Up networking
Scale Up networking increases the capacity of a single system by adding more GPU resources which operate as a single logical compute unit.
Scale Up networks connect 2 to 8 GPUs per compute tray and 72 GPUs per rack.
- 800 Gbps/GPU
- GPU power 1000 W (GB300)
- 130 Tbps GPU-to-GPU throughput within the rack
- Rack power 140 kW (GB300 typical)
Scale Up networks are tightly-coupled systems optimized for latency and synchronized traffic. They rely on short-reach, under 5m, copper interconnect solutions which offer lowest latency and highest power efficiency.
Scale Out networking
Scale Out networks connect multiple scale up systems into large clusters for distributing AI workloads across nodes.
They are designed to connect thousands of GPUs. Scale Out networks are optimized for massively parallel workload distribution that enable highest GPU resource utilization.
They rely on reliable high-speed optical interconnects that deliver robust signal quality while reducing complexity and power consumption.
- GPU: 1,000s to 100,000
- Racks: 1,000+
- Energy consumption: 15 MW to 100+ MW
- Reach: 100m to 2km
Scale Across networking
Scale Across networks are designed to connect GPU clusters between data centers.
Traditional DCI connects front end CPU networks between data centers and to end-users. DCI requires Tbps of transport capacity. A key objective is optimizing scarce fiber capacity.
Power is the main constraint driving Scale Across networks. AI training capacity in 100 MW data centers requires geographically distributing backend GPU networks between data centers, creating a synchronized and distributed GPU fabric.
Scale Across can require Pbps of transport capacity between data centers. A key objective is scaling transport over many fiber pairs.