Resources /
Smarter GPU Placement Starts with Understanding the Fabric
Summary
Having available GPUs doesn’t mean they’re the right GPUs for the job. In large AI clusters, where those GPUs sit within the InfiniBand fabric can influence how efficiently workloads use the network. That makes understanding the physical topology an important part of smarter GPU placement.
Netris 4.14 exposes InfiniBand topology through the API, allowing orchestration platforms to understand where GPU servers sit physically within the network, not just whether they’re available.
This enables smarter GPU placement, more efficient use of network resources, and better operational planning as AI infrastructure scales. This enables improved GPU utilization, reduced risk in a shared network infrastructure, and a greater awareness of potential network failure domains.
The Problem
As GPU clusters expand beyond the boundary of a single rack and, therefore, require complex InfiniBand fabrics to support scale-out GPU connectivity, any two GPU servers aren’t always the same from a topology perspective. For example, two available servers could be located in different racks requiring traffic between them to traverse additional shared network infrastructure.
The challenge isn’t that InfiniBand is lossy or that an extra switch hop is inherently a problem. The issue is that an orchestration system allocating GPUs may know which servers are available without understanding where those servers sit relative to one another in the physical network.
The result is that as GPUs are continuously allocated and released, available capacity can become scattered across different racks, switches, and parts of the network similar to how a hard drive becomes fragmented from active use over time. Without topology awareness, an engineer or an orchestration platform could allocate a set of GPUs based primarily on availability, with no understanding of the network resources and failure domains they share.
The real problem is that in large shared GPU infrastructures, which GPUs you allocate together can affect the performance quality of the infrastructure, and therefore services, the end customer receives. If an orchestrator selects GPUs purely because they’re available, it can spread a workload across servers that depend on different parts of the InfiniBand fabric, consuming more shared fabric resources and potentially creating a larger operational and failure footprint than necessary, while simultaneously providing a less performant infrastructure to the operator’s end customers.
The Solution
Netris 4.14 enhances the physical placement visibility of bare-metal servers by exposing the InfiniBand topology information, including the relevant leaf, spine, and core relationships, through the Netris API. An orchestration platform can then use that knowledge as an input to make informed workload placement decisions. And operators can use the same information for assessing the scope and potential impact of network changes. Netris also shows when topology data was last updated, helping consumers understand how current the information is.
Rather than considering availability alone, the platform can incorporate network proximity into its placement logic, for example, preferring resources within the same rack or under the same leaf when possible before allocating resources that are farther apart in the fabric.
Why it Matters
For neocloud and AI infrastructure operators, this becomes very important as clusters become larger and resource availability changes continuously. GPUs are allocated, released, and allocated again. Static assumptions about where available capacity lives quickly become outdated.
In fact, the same topology information has operational value beyond initial placement. Because operators can query servers according to their physical location in the InfiniBand fabric, the data can also support blast-radius analysis and maintenance planning.
The benefits are clear. Netris 4.14 gives orchestration platforms and operators greater visibility into where GPU servers sit within the InfiniBand fabric, enabling:
- Improved GPU utilization
- Distributed computing resources rely on close physical proximity to maximize performance and network efficiency
- Reduced operational risk associated with a shared infrastructure
- This means GPU resources associated with a single workload are co-located physically adjacent or near each other in the network
- Greater awareness of shared failure domains
- Better blast-radius analysis and maintenance planning
With a current, API-accessible understanding of how available GPU servers relate to the underlying InfiniBand fabric, GPU placement can account for not just what is available, but where it is.
Join Our Community
Connect with cloud networking professionals
Slack
Join our Slack community to connect with like-minded developers.
Newsletter
Subscribe to our newsletter for the latest updates and insights.