Arista's Paul Gilbert warns traditional data center design fails AI workloads
The shift from enterprise plumbing to AI networking
Building an AI data center is fundamentally different from traditional enterprise networking. In a standard data center, we design for diverse traffic moving "north-south" from users to apps. AI infrastructure flips this script. Paul Gilbert, Tech Lead at Arista Networks, explains that training Large Language Models requires a "backend network" that is completely isolated from everything else. This isn't just about security; it's about physics. When you connect NVIDIA H100 servers, you are dealing with traffic bursts at 400Gb or 800Gb per port. If one packet drops, the entire training job—costing thousands of dollars an hour—can stall or fail. We have moved from 1:10 oversubscription ratios to a strict 1:1 non-blocking architecture where every bit of bandwidth must be guaranteed.
Tools and infrastructure for 2025

To build a functional AI cluster in 2025, you need to abandon the standard rack specifications of the last decade.
- Hardware: NVIDIA or Supermicro GPU servers (HGX/DGX) with 8 GPUs per node.
- Switching: High-radix switches like the Arista 7800 Series capable of handling 800G and 1.6T throughput.
- Cabling: High-quality transceivers and Direct Attach Copper cables; at this scale, cable failure is a leading cause of downtime.
- Power: Racks capable of 100kW to 200kW, almost certainly requiring liquid cooling.
Implementing lossless ethernet and flow control
Because AI training is sensitive to latency and packet loss, we rely on RoCE v2 (RDMA over Converged Ethernet). This protocol allows for memory-to-memory transfers between GPUs without hitting the CPU, but it requires precise tuning of flow control mechanisms. You must configure Explicit Congestion Notification (Explicit Congestion Notification) to provide an end-to-end "slow down" signal to the senders before buffers overflow. If that fails, Priority Flow Control (Priority Flow Control) acts as the emergency brake to prevent packet loss. Without these protocols correctly synchronized between the NVIDIA ConnectX-7 and the switches, the network will simply melt under the synchronized pressure of thousands of GPUs bursting simultaneously.
Visibility and the AI agent
A massive blind spot in AI networking is the gap between the switch and the GPU. Traditional telemetry tells you the switch is fine, but the data scientist is calling because the job completion time just jumped from one hour to four days. We use an Arista AI Agent loaded directly onto the GPU nodes via an API. This agent allows the switch and the GPU to "talk" to each other, confirming that flow control settings match on both ends. It provides visibility into RDMA error codes that were previously hidden, allowing engineers to identify if a specific GPU or a specific cable is the bottleneck before the entire training run is ruined.
Tips and troubleshooting the backend
When things go wrong, it's rarely a software bug; it's usually a physical layer issue or a configuration mismatch. Always use BGP as your routing protocol for its simplicity and scale. If you are seeing performance degradation, check your entropy settings. Standard load balancing using five-tuple hashing often fails with GPU traffic because everything looks like one IP address, leading to oversubscribed uplinks. Use cluster load balancing that looks at actual bandwidth utilization. Finally, prepare for the Ultra Ethernet Consortium (UEC) standards arriving in 2025, which will further offload congestion management from the network to the NICs.
The outcome of proper AI design
By following this methodical approach—isolating the backend, eliminating oversubscription, and enforcing lossless flow control—you ensure that your gpus spend their time calculating, not waiting on the network. The goal is predictable job completion time. In a world where 9.6 terabytes of traffic per server is the new normal, your network is no longer just plumbing; it is the backbone of the model's intelligence.
- Arista 7800 Series
- 6%· products
- Arista AI Agent
- 6%· technologies
- Arista Networks
- 6%· companies
- BGP
- 6%· technologies
- Cisco Systems
- 6%· companies
- Other topics
- 72%

How to Build Your Own AI Data Center in 2025 — Paul Gilbert, Arista Networks
WatchAI Engineer // 23:00
We turn high signal in-person events for the top AI engineers, founders, leaders, and researchers in the world into the best free learning opportunities for millions around the world here on YouTube. Your subscribes, likes, comments, speaking, attendance, or sponsorships goes a long way toward making our biz model sustainable indefinitely. We strongly believe this industry deserves a better class of community and that we know how to do this well; we just need your support.