Every rack-scale AI platform shipping in 2026 asks you to make the same bet: pick one GPU vendor, accept the ratio of compute to memory to storage that someone else chose at the factory, re-plumb your data center for 100% direct liquid cooling, and live with that decision for the life of the asset. NVL72, Vera Rubin NVL144, Helios — all beautiful machines, all bolted shut.
The Mercury Rack takes the opposite position. It is a 48U composable platform with 72 double-wide PCIe Gen5 GPU slots, and you decide what goes in them — including changing your mind later, in software, without a forklift.
What is in the rack
Mercury is a full 48U cabinet, laid out to keep the GPU pool as large as possible while leaving room for the hosts and management that make it usable:
- 9× PRU 2500 shelves — the Photonic Resource Units that hold the accelerators and storage.
- 72× double-wide PCIe Gen5 slots across those shelves — the composable pool itself.
- 8U of host servers — an 8:1 GPU-to-CPU ratio, because pooled GPUs do not need a CPU each.
- 1U management switch and 3U reserved for expansion.
The result is up to 111 kW in a standard 19-inch rack — air-cooled, with in-rack direct liquid cooling available as an option rather than a prerequisite.
Any GPU. Any vendor. Any ratio.
The 72 slots take standard double-wide PCIe Gen5 accelerators, which means the shopping list is yours: RTX 6000, H200, MI350, MI210, R9700, Intel accelerators, and whatever ships next. NVIDIA and AMD silicon can sit in the same rack, on the same fabric, serving different workloads — and the software stack follows, with CUDA and ROCm available side by side rather than one excluding the other.
That is the part fixed monoliths structurally cannot do. A 72-GPU NVL72 is 72 B200s. A Helios rack is 72 MI455X. Mercury is 72 slots.
The fabric: PCIe Gen5 over optical
Composability is only interesting if the pooled devices perform like local ones. Mercury connects hosts and PRUs over a PCIe Gen5-over-optical fabric, giving full Gen5 bandwidth between any GPU pair in the rack — with no proprietary interconnect domain drawing a boundary around which GPUs can talk to which.
It is a vendor-neutral fabric built on the standard every accelerator already speaks, which is what makes mixing silicon possible in the first place.
Storage that composes too
Each PRU 2500 can be populated with accelerators or with NVMe — up to 40 SSDs and 1.2 PB per PRU. Those drives are composable on the same fabric, which makes in-rack KV-cache offload and GPUDirect Storage practical without a separate storage rack sitting next to the AI rack. For long-context inference, where the KV cache is the thing that actually runs out, that capacity is in the same cabinet as the GPUs consuming it.
Composer: the rack is software
Corespan Composer is the control plane that turns the pool into shapes. It assigns devices to hosts at runtime — attach eight GPUs to one host for a fine-tune, break them apart into inference workers an hour later, add NVMe to a host that needs scratch space — without cabling changes, BIOS reboots, or downtime. It sits beneath Kubernetes and Slurm rather than replacing them, so the schedulers you already run keep working.
This is the difference between a rack you buy and a rack you operate: the ratio of GPUs to hosts to storage stops being a purchasing decision made once, and becomes a scheduling decision made continuously.
Mercury vs. the Monoliths
Side by side against the 2026 rack-scale platforms, the trade is straightforward — Mercury gives up the proprietary scale-up interconnect and gains vendor choice, air cooling, mixed software stacks, and a lower entry price.
Who this is for
If you are pre-training frontier models on a single stack, the monoliths are built for you and you should buy one. Mercury is for everyone else: inference and fine-tuning fleets, agentic serving, RAG and multi-modal pipelines, storage-heavy workloads, research clusters with mixed silicon, and anyone whose data center is air-cooled today and will stay that way for a while.
It is also for teams who simply do not want their next three years of GPU purchasing decided by a fabric. If that is the position you are in, we should talk.