Loading article…
Google’s new Virgo networking fabric connects up to 134,000 AI chips with 47 Pb/s bandwidth, aiming to reduce latency in massive GPU and TPU clusters.
Google has launched Virgo, a dedicated scale-out network fabric designed to connect massive clusters of GPUs and TPUs, capable of supporting 134,000 chips with 47 petabits per second of non-blocking bisection bandwidth [1]. The architecture marks a strategic shift for the company, separating high-speed accelerator-to-accelerator traffic from general-purpose data center networking to improve performance and fault tolerance in large-scale AI training and inference [1].
| At a glance | |
|---|---|
| Product | Virgo Scale-Out Fabric |
| Max Capacity | 134,000 TPU 8t chips [1] |
| Bisection Bandwidth | 47 Pb/s [1] |
| Topology | Two-layer, non-blocking [1] |
Virgo functions as an "east-west" fabric, operating independently of Google’s existing Jupiter front-end network and the internal scale-up interconnects used within individual accelerator systems [1]. By utilizing high-radix switches—which house a high number of ports—Google has collapsed the network into a flat, two-layer topology [1]. This design replaces traditional three-layer networks, a change Google claims provides up to four times the bandwidth per accelerator and up to 40% lower unloaded fabric latency compared to previous configurations [2].
The fabric is engineered to handle the "inevitable" hardware failures inherent in clusters containing hundreds of thousands of chips [1]. Virgo employs a multi-plane design where each accelerator connects across multiple independent network planes; if one plane fails, the system retains 87.5% of its aggregate bandwidth rather than suffering a total connection loss [1]. To manage congestion, the architecture incorporates high-resolution telemetry and CSIG in-band signaling, allowing the network to detect microbursts and steer traffic around bottlenecks in real-time [1].
Google is positioning Virgo as the foundation for its next generation of AI infrastructure, including A5X systems based on NVIDIA’s Vera Rubin NVL72 and the eighth-generation TPU [1]. While Virgo currently manages connectivity within and across data center buildings, the company is already developing "Scale Across" networking, which aims to extend these capabilities via long-haul RDMA (Remote Direct Memory Access) to link accelerator clusters separated by hundreds of kilometers [1].
This evolution reflects a broader trend where networking is no longer treated as generic infrastructure but as a component co-designed with specific generations of accelerator hardware [1]. By separating the back-end fabric from the front-end, Google intends to evolve its AI networking independently of general-purpose compute resources, potentially allowing for more rapid upgrades as new GPU and TPU technologies emerge [1].
The success of Virgo will likely hinge on whether its "invisible" failure model can maintain consistent performance during the massive, multi-site training runs that define the current AI arms race. As Google pushes toward architectures supporting nearly one million GPUs, the ability to manage network-level bottlenecks will be as critical to performance as the raw power of the underlying silicon [1].
Coverage is mostly measured — 219 of 231 reports stay neutral.
Every Monday — the token unlocks, Fed dates & catalysts set to move crypto and markets this week. So you’re never blindsided.
Free · 3-min read · one-click unsubscribe
AI-assisted synthesis by the TrendWatcher Editorial Desk · sourced from 3 outlets · Aug 26, 2026 · How we report
Google has integrated AI for features such as video creation, automated bidding, and reporting for AI-generated product titles.
AI agents typically perform supervised tasks and return results for human review, whereas autonomous AI systems work toward a goal for days or weeks to return finished work without human intervention.
Recent reports indicate that images have been assigned new functions within the context of Google's AI-driven search experience.