Article Overview

For optimal AI performance, Dawning AI servers should use high-bandwidth, low-latency NICs with RDMA/RoCEv2 enabled, mapped per GPU via PCIe switches, and integrated into a lossless scale-out fabric.

Network Fabric Overview

Dawning AI servers rely on multiple network fabrics to support AI workloads:

  • Front-End Fabric: Connects applications, storage, and GPU cluster management. Typically requires 1–2 NICs per server .
  • Internal-AI Fabric: Integrates GPUs, CPUs, NICs, and storage via PCIe switches for efficient peer-to-peer communication .
  • Scale-Up Fabric: Handles intra-server GPU-to-GPU communication using NVLink, Infinity Fabric, or Ethernet .
  • Scale-Out Fabric: Connects multiple AI servers for distributed workloads. Requires one NIC per GPU with matching PCIe bandwidth (e.g., PCIe Gen5 x16 or 400Gbps) to ensure non-blocking, low-latency transfers .

NIC Configuration Guidelines

  1. RDMA and RoCEv2:
    • Enable RoCEv2 on Broadcom or compatible NICs to offload transport tasks from the CPU and provide direct memory access between GPUs and network .
    • Ensure NICs support UDP/IP routable RoCEv2 for cluster-wide communication.
  2. PCIe Mapping:
    • Each GPU should have a dedicated NIC connected through a PCIe switch to avoid bottlenecks .
    • Verify PCIe bandwidth matches NIC throughput (e.g., 400Gbps NICs require PCIe Gen5 x16 lanes).
  3. Server BIOS Settings:
    • For AMD-based servers (e.g., Supermicro AS-8125GS-TNMR2), disable IOMMU and ACS if required by the workload .
    • Dell XE9680 servers may require careful SR-IOV configuration; consult vendor documentation to avoid boot faults .
  4. GPU-NIC Topology:
    • Use tools like rocm-smi for AMD GPUs to verify hardware topology and NIC mapping .
    • Ensure GPUs and NICs are correctly paired for optimal RDMA performance.

Network Design Considerations

  • Lossless Ethernet: Configure Priority Flow Control (PFC) and Explicit Congestion Notification (ECN) on switches to prevent packet loss during high-throughput AI workloads .
  • Spine-Leaf Architecture: Recommended for scale-out fabrics to reduce latency and provide sufficient bandwidth for multi-node training .
  • Quality of Service (QoS): Assign DSCP values (e.g., CS3) for RoCEv2 traffic and configure class-maps and policy-maps on switches to maintain low-latency, high-priority traffic .

Monitoring and Validation

  • Use NIC and GPU management tools to monitor throughput, latency, and RDMA performance.
  • Validate network paths and congestion management to ensure non-blocking communication across the AI cluster . By following these guidelines, Dawning AI servers can achieve high-throughput, low-latency networking, essential for distributed AI training and inference workloads.

Understanding GPU Servers and Their Role in Data Centers

But as demand for infrastructure capable of supporting AI model training and inference has grown, the ability to host

Choosing a Server for Deep Learning Training

For optimal performance, the server''s configuration should include a balanced PCIe topology,

Configuring network card settings

This guide provides instructions for configuring network card settings to optimize network performance and connectivity.

How to Setup and Optimize GPU Servers for AI Integration?

Learn how to setup and optimize GPU servers for AI integration. This guide covers hardware selection, OS & drivers installation, AI

High-Performance GPU Server Hardware Topology and Cluster

This article explores the hardware topology and cluster networking of high-performance GPU servers, focusing on

Dawning Information''s Earnings Momentum: A Strategic Buy

Hygon''s expertise in advanced chip design complements Dawning''s server and storage solutions, creating a

Configuration Guidance for AI & Data Science

#41 Hardware requirements for Client AI can vary significantly dependent on the type, size and complexity of project

Guide to Building a Bare-Metal AI Server

Transforming a list of carefully selected components into a functional server requires a methodical assembly and configuration

DAWN Internet

Join DAWN, the decentralized wireless network revolutionizing Internet service. Empower communities to buy, sell, and share

Networking — NVIDIA DGX BasePOD: Deployment Guide Featuring

BMC network access is provided through the out-of-band network (marked in purple). The networking ports and their

How to Setup and Optimize GPU Servers for AI Integration

Learn how to set up and optimize GPU servers for AI. Discover best practices for configuration, tuning, and deployment.

Powering AI: The Semiconductor Ecosystem at the Foundation of

A host of power and networking semiconductors enable efficient energy delivery and seamless interconnection within AI servers and

Networking Solutions for the Era of AI | NVIDIA

Networking Purpose-Built for the Era of AI Factories AI factories now span tens of thousands of GPUs—and soon,

How to build a high-performance AI server locally

Learn how to build a high performance AI server to allow you to run large language models locally. Removing the need

Building an 8K GPU Cluster with High-Performance Ethernet

This reference design leverages DriveNets AI Fabric and the breakout capabilities of NCP5-AI leaf switches to create a

Dual network cards

In some scenarios, you may require two network cards for your setup. This article describes how to configure dual

Stand out with designations, specializations, and badges introduced in

This badge differentiates partners who have demonstrated advanced capabilities across numerous Microsoft Cloud

NETWORKING THE AI DATA CENTER

The ideal approach to solve AI networking challenges is based on open standard and interoperable Ethernet fabrics, with a focus on

High-Performance Ethernet Networking for Artificial Intelligence Systems

Before digging into the details of how to maximize the network performance, it is critical to understand the server and network

Udemy: Online Courses for Skills, Careers & AI

Learn in-demand skills with online courses, get professional certificates that advance your career, and explore courses in AI, coding,

Building an Efficient EdgeAI Server: A Guide to Dual

Learn why building a multi-GPU EdgeAI server is a smart investment. Get the benefits of

AI Server — documentation

It is necessary to configure BMC and NR1 module network interfaces to operate the NR1 AI Inference

Introduction | NVIDIA ConnectX-8 SuperNIC User Manual

With support for both InfiniBand and Ethernet networking at up to 800 gigabits per second (Gb/s), ConnectX-8

AI Server Configuration

AI Server Configuration DeGirum AI Server The DeGirum AI server software stack allows you to run AI model inferences initiated

How to Set Up and Optimize GPU Servers

Conclusion Understanding how to effectively use GPU servers can unlock new possibilities for your business, from

NVIDIA Configuration | Juniper Networks

NVIDIA® ConnectX® family of network interface cards (NICs) offer advanced hardware offload and acceleration features, and

High-Performance Ethernet Networking for Artificial Intelligence Systems

The diagram below shows a very high-level architecture and key components of a standard AI server. The highest performance

Knowledgebase

Looking for a dedicated server to deploy your AI models? Bacloud offers dedicated GPU servers tailored to your needs. Choose from

Building AI Network Infrastructure: A Step-by-Step Guide

Once you have a solid grasp of your current network setup and the specific needs of your AI applications, it''s time to

AI Data Center Networking: How GPU Clusters Are Changing Network

This guide provides a rigorous technical deep-dive into how AI workloads are reshaping data center network design:

Optimizing AI Workloads: Best Practices and Tips

Explore essential practices for optimizing AI workloads, including server configuration, software optimization, and network management.

Building an Efficient GPU Server with NVIDIA GeForce

In today''s AI-driven world, the ability to train AI models locally and perform fast inference on

DeviceNet Network Configuration

When you configure the network online, the devices on the network have parameters configured. Complete these steps to upload

dawning

When Dawning Technologies knows what initial instruments are going to be interfaced, they will send the system preconfigured.

Carde.io Admin Dashboard

Manage your gaming community with Carde.io''s Admin Dashboard, designed for seamless organization and enhanced engagement

How to manage network adapter settings on Windows 11

In this guide, I will teach you the steps to manage Ethernet and Wi-Fi network adapters on Windows 11 using the Settings app.

Designing Data Centers for AI Clusters

About this Document This document is a generic design document for building network infrastructure for high-performance AI clusters.

Quick Understanding GPU Server Network Card Configuration in AI Era

Delve into GPU server network configurations & optical communication solutions in the era of GenAI. Learn about

Best Mini PC for Ollama and Local LLMs in 2026:

Run Ollama and Llama 3 locally with confidence. Discover the best mini PCs for local AI

Related Resources

Need Precision Optical Test Instruments?

Request a free quote for OTDR, power meters, light sources, spectrum analyzers, return loss testers, VFL, or complete fiber test kits. EU‑owned manufacturer with local support in South Africa – reliable, accurate, and field‑proven equipment.