
OPTIMUSEDGE AI
The Daily Briefing on Physical AI, Orbital & Edge Infrastructure, Networking & Autonomous Agents
Built on the NVIDIA Technology Stack
| MON AI Infrastructure |
| TUE Data Science |
| WED Generative AI |
| THU Simulation & Physical AI |
| FRI AI Agent Design — Use Case 1 |
| SAT AI Agent Design — Use Case 2 |
| SUN Hottest NVIDIA & AI News |
NVIDIA
THE TECHNOLOGY STACK AT THE CORE OF EVERYTHING WE COVER
Welcome back to the OptimusEdge. One NVIDIA rack can now hold 72 GPUs and 36 CPUs, wired together as a single computer.
Not a cluster of computers. One computer.
Today we open it up and see how that's actually built.
___________________________________________________________
The Edge Upload: Today’s Insights
What "AI factory" actually means, in NVIDIA's own words
The building block every SuperPOD is made from one rack
How racks link into thousands of GPUs without falling apart
The software layer nobody talks about, but that makes the whole thing usable

NVIDIA DGX SUPER POD - Reference: NVIDIA DOCS
TECH RADAR - WHATS HAPPENING - LATEST NEWS TO LEARN FROM
Storage just got SuperPOD-certified. Pure Storage's FlashBlade//S R2 earned official certification for DGX GB200 and GB300 SuperPODs this year.
Why it matters: GPUs this expensive are only as fast as the data reaching them exactly the storage-starvation problem we covered a few issues back.
The B300 generation has its own blueprint now. NVIDIA published a dedicated reference architecture for DGX B300 SuperPODs.
Different chip generation, same underlying playbook which is exactly why understanding the architecture once pays off for years.

DGX B300 SuperPODs- Reference Architecture - Reference : NVIDIA DOC
WHAT “AI FACTORY” ACTUALLY MEANS
NVIDIA doesn't call this a data center. It calls it an AI factory.
Jensen Huang's own framing: "NVIDIA DGX AI supercomputers are the factories of the AI industrial revolution."
A factory has two inputs: electricity and data. Output: tokens.
Sounds like marketing. It's actually a useful way to think about the design every choice below exists to keep that "factory" running at full output.
How It's Actually Architected — The Full Stack
Here's the whole thing, top to bottom, in the order it's built.
Layer 1 - The GPU itself. Blackwell or Blackwell Ultra silicon, with HBM3e memory attached directly on the package.
Layer 2 - NVLink, inside the rack. GPUs within one rack talk to each other over NVLink, not a network cable. This is what turns 72 separate chips into one shared-memory machine.
Layer 3 - The rack, as one node. 72 GPUs plus 36 Grace CPUs, wired as a single NVLink domain. This whole rack is the base unit everything else scales from.
Layer 4 - The Scalable Unit (SU). A fixed bundle of racks (8 DGX GB300 systems per SU) pre-validated, not custom-designed each time.
Layer 5 - The fabric between racks. InfiniBand or Spectrum-X Ethernet at 800 Gb/s connects SUs to each other a different, slower tier than NVLink, by design.
Layer 6 - Orchestration software. Mission Control, CUDA, and Magnum IO sit on top of all of it, making thousands of physically separate GPUs schedulable as one resource.
Six layers. Each one exists because the layer below it can't scale far enough on its own NVLink can't span a data center, InfiniBand alone can't share memory, and none of it is usable without software tying it together. That's the whole architecture story in one pass the sections below go deeper on each layer.

AI Factory - Nvidia CEO Jensen Huang delivers the keynote address during the GTC 2025 conference in San Jose. Justin Sullivan/Getty Images
THE BUILDING BLOCK: ONE RACK
Start small. One DGX GB200 rack holds 36 Grace CPUs and 72 Blackwell GPUs connected as one machine via fifth-generation NVLink.
Not "connected like a network." Connected like one motherboard, just rack-sized.
The newer GB300 rack pushes this further: 72 Blackwell Ultra GPUs and 36 Grace CPUs in the same NVLink domain, with up to 13.4 TB of HBM3e memory and roughly 576 TB/s of aggregate bandwidth per rack.
That's the whole idea of DGX. One rack is already a supercomputer. Everything past this point is just more of them, wired together.
Scaling Out: The Scalable Unit
NVIDIA doesn't scale SuperPODs rack by rack. It scales them in chunks called Scalable Units (SUs).
One SU = 8 DGX GB300 systems a fixed, pre-validated building block, not a custom design each time.
Stack enough SUs and the official reference architecture scales to 128 racks and 9,216 GPUs a real, currently deployable configuration, not a future roadmap slide.
Why chunks instead of one continuous design? Predictability. A pre-validated SU means a customer isn't debugging a one-off network topology at 9,000-GPU scale.

GB200 NVLink Rack Configuration. DOCS NVIDIA
THE NETWORKING GLUE
Inside a rack, NVLink does the talking. Between racks, it's a different job entirely.
DGX SuperPOD's compute fabric runs on 800 Gb/s InfiniBand or Spectrum-X Ethernet a full generation faster than what most enterprise networks run today.
NVIDIA's Quantum-X800 InfiniBand pushes this to 1,800 GB/s of bandwidth per GPU.
BlueField-3 DPUs sit in this fabric too handling networking and security tasks so the GPU never has to slow down to do "boring" infrastructure work.
THE SOFTWARE LAYER WHICH CAN BE OFTEN OVERLOOKED
Hardware gets the headlines. This is the part that actually keeps 9,000 GPUs from becoming 9,000 separate headaches.
NVIDIA Mission Control handles orchestration across the whole SuperPOD workload scheduling, infrastructure health, and resilience, delivered as software layered on top of the hardware.
Underneath that: CUDA, NVIDIA AI Enterprise, and Magnum IO the layers that make thousands of GPUs behave like one system instead of a very expensive pile of parts.
HOW COMPANIES ACTUALLY GET ONE
Two real paths. Buy the reference architecture and build it on-prem this is where certified partners like Pure Storage plug in for the storage layer.
Or rent equivalent capacity through DGX Cloud, which we'll go deep on in an upcoming issue.
Either way, nobody is assembling this from scratch.
That's the entire point of a reference architecture the hard engineering decisions are already made before a single cable gets plugged in.
Takeaway: A DGX SuperPOD isn't a bigger version of a normal server rack it's a pre-engineered system where compute, networking, and software are designed together as one unit, from a single rack up to 9,216 GPUs. Understanding the rack is understanding the whole architecture; everything past that is repetition, not new complexity.
BEFORE YOU ORDER A SINGLE GPU 🙂
That's today's briefing. If "SuperPOD" has always sounded like a marketing word to you, hopefully it sounds like an engineering decision now. Past issues are in the archive. See you tomorrow.
INFRA TOOL OF THE DAY
NVIDIA DGX SuperPOD Reference Architecture (GB300): The actual document NVIDIA gives partners and customers to build one of these.
Dense, but it's the real blueprint worth skimming even if you never touch the hardware yourself.
QUICK EDGE HITS & REFERENCES
Full Architecture Breakdown: NVIDIA DGX SuperPOD: Next Generation Scalable Infrastructure for AI Factories the official GB300 reference architecture document
B300 Generation Details: DGX B300 SuperPOD System Architecture networking-focused breakdown of the B300-era design 💾 Storage Certification: Pure Storage + NVIDIA DGX GB200/GB300 SuperPOD Certification real per-rack specs and the AI factory framing
Original Announcement: NVIDIA Launches Blackwell-Powered DGX SuperPOD the GB200-generation launch, Jensen Huang's full quote in context
That’s it for today !☀
Edge AI is levelling up are you? Until next time, stay curious, stay building, and don’t let your machines take over. 🤖😆
Enjoyed today’s issue? Share OptimusEdge AI with your engineering team. Subscribe to OptimusEdge AI
Your Edge AI Explorer,
Sharat Sami (Let’s connect on LinkedIn)
