RoCE
Networking & InterconnectRDMA over Converged Ethernet, a technology that brings the low-latency, CPU-bypassing benefits of RDMA to standard Ethernet networks.
RoCE is one of the certified fabric options under NVIDIA Reference Architecture, offering an Ethernet-based path to the low-latency interconnect, while Eleveight AI's own cluster uses NVIDIA InfiniBand XDR.
Overview
RDMA's direct, low-latency memory access was born on InfiniBand, a specialized and high-performing but distinct networking technology. RoCE brings that same capability to Ethernet, the most widely deployed networking standard in the world. It lets organizations gain RDMA's efficiency, direct memory access with the CPU bypassed, while building on familiar Ethernet infrastructure rather than a separate fabric.
How it works
RoCE encapsulates RDMA operations within Ethernet, so that direct memory-to-memory transfers run across an Ethernet network. Achieving InfiniBand-like performance this way requires the network to be carefully configured for lossless behavior, since RDMA is sensitive to dropped packets, which is why RoCE deployments use specific flow-control and congestion-management techniques to keep the fabric clean under heavy load.
Why it matters
RoCE gives data center operators a choice in how they build a high-performance AI fabric: the dedicated route of InfiniBand, or RDMA over an Ethernet foundation many teams already know. Both are recognized paths to the low-latency, high-bandwidth interconnect that distributed training demands, and having a compliant Ethernet option broadens how a reference-architecture cluster can be assembled without sacrificing performance.
Use cases
- Low-latency AI fabrics built on Ethernet
- Distributed training interconnect
- High-throughput storage networking
- An Ethernet alternative to InfiniBand in reference-architecture clusters