The AI buildout is being scored on the wrong metric. The winners of the inference era will not own the biggest boxes — they will own the fabric that composes compute, memory, and I/O around each workload, then reclaims it for the next one. That fabric is what Corespan builds.
For most of the AI industry, the answer to more demand is still a bigger box — more GPUs behind a tighter interconnect, more memory in a larger node, more power concentrated in a rack, a pod, a purpose-built hall. That reflex made sense when training dominated and the goal was to finish one job sooner.
It makes less sense for the workload now driving infrastructure decisions: serving models to users, agents, and applications at variable demand. The next trillion tokens will not come from a permanently assembled supernode. They will come from infrastructure that can reorganize around the work. That is the distinction between scale-up and scale-across — between a larger fixed domain and a pool of compute, memory, and connectivity composed into the right domain at the right time.
Where the Bigger Box Breaks
A useful design principle can harden into an expensive habit. Training large models requires high-bandwidth, low-latency communication across many accelerators, and keeping those accelerators physically close reduces friction for the job at hand. Inference is not a single synchronized training run. It is a flow of requests with different models, context lengths, latency targets, memory footprints, and service-level requirements. A rack- or pod-scale system is elegant when the components are fully occupied by the same work. It is less elegant when GPUs wait for a burst, expensive compute is stranded beside insufficient memory, or an entire domain is reserved for one customer. The fixed box is then imposing its demand profile on the application.
Scale-up is also hitting several ceilings at once. The first is physics: modern accelerators have pushed power density and cooling into a regime where power delivery, heat rejection, serviceability, and facility design are first-order constraints. Direct liquid cooling makes higher densities possible, but not concentrated power free.
The second is the practical boundary of a tightly coupled interconnect domain. NVLink-class fabrics are exceptionally effective within their intended domain, but as the domain grows, so do the costs of integration, topology management, failure handling, and maintenance. The largest coherent unit becomes the unit operators must purchase, deploy, protect, and recover.
Then there is blast radius. A firmware bug, fabric fault, or cooling event can take down a large block of monolithic capacity at once — and the usual response is more spares and reserved headroom, paid for in advance. Finally, scale-up concentrates capex before utilization is proven. Operators buy a particular shape of capacity, not only the capacity needed this quarter, and accept the risk that demand will not fill it evenly over its useful life.
None of this is an argument against dense systems. It is an argument against assuming density is the end state of the architecture. The question is no longer how large a GPU domain can we build. It is how precisely can we allocate the domain each workload actually needs.
Inference Changes the Math
Inference is where the fixed-box assumption breaks most clearly. Serving traffic is bursty. Demand can move by geography, time of day, model release, product feature, or one customer’s batch run. The request stream is also heterogeneous. A small model may need GPUs for high-throughput prefill; a larger reasoning model or long-context workload may be constrained by memory first.
Decode adds another complication. During token-by-token generation, performance is often shaped by moving and accessing model state efficiently, not simply by adding floating-point capability. “GPUs installed” is therefore a weak proxy for useful inference capacity. A GPU waiting on memory and a reserved pod carrying low steady-state load can look well provisioned while producing disappointing economics.
Monolithic infrastructure handles uncertainty by overbuilding for the peak case. That can leave expensive resources underutilized between peaks and poorly matched to the requests that actually arrive.
This is the problem Corespan was built for: evolving fixed pods into composable pools, so operators stop paying peak-case prices for average-case demand.
The Scale-Across Alternative
Scale-across starts with a different premise: the useful unit of infrastructure is a workload-defined composition, not a rack-defined box. In practice, that means separating GPU compute, vRAM capacity, and I/O into pools that can be assembled and released as demand changes. A workload that needs more memory should receive it without reserving an overprovisioned GPU node. A workload that needs a larger GPU cohort briefly should acquire it without permanently reshaping the data center.
PCIe remoting over photonics is the enabler. It extends the resource fabric while retaining a direct, standards-based relationship between hosts and devices. Latency, bandwidth, topology, and workload behavior still matter; the point is to make those tradeoffs manageable instead of freezing resource relationships at installation.
Optical switches alone are not the architecture. Optics are becoming commodity. The durable value sits above the optical layer — in knowing which devices can be composed, which paths meet a workload’s requirements, how to establish those paths safely, and how to return capacity to the pool without operator intervention.
Scale-across also rescues heterogeneous fleets. Current-generation GPUs can serve the highest-value jobs, while prior-generation GPUs continue serving workloads for which their performance, memory, and power profile remain economic. Specialized accelerators can enter without a forklift replacement. The pool becomes more capable because its differences can be orchestrated.
Economics Follows Architecture
The most important consequence of composability is not a prettier topology diagram. It is a different cost curve for tokens. Cost-per-token reflects utilization, idle capacity, power, cooling, depreciation, reserved headroom, and the cost of capacity larger than the workload requires. A fixed pod can post impressive theoretical performance while delivering poor economic output if it waits for the wrong demand.
Composable pools create more ways to extract useful work from assets already owned: allocate memory to context-heavy requests, build a temporary GPU group for a burst, or assign suitable work to a prior-generation fleet. Capacity planning remains hard, but it depends less on predicting one permanent shape of demand.
This also changes residual risk. A shift in model architecture or customer mix can leave a tightly integrated asset poorly matched to the next wave of demand. A composable pool gives the operator more ways to reconfigure around that shift. That flexibility matters as AI infrastructure moves from a race to install capacity into a discipline of operating it.
What Composable Actually Requires
Scale-across is not achieved by putting optics between servers or routers and calling them composable. It requires a real control plane — one that maintains an accurate view of resources, topology, health, policy, and ownership; composes and decomposes infrastructure predictably; enforces isolation; handles failures; and answers a simple question at any moment: what is connected to what, and why? This is what Corespan Composer does — the difference between a fabric you can demo and a system you can operate.
The control plane must also be vendor-neutral and generation-neutral. An architecture limited to one vendor or its newest devices recreates lock-in and stranded assets at a different layer. Operators need freedom to introduce new hardware, retain useful hardware, and allocate both by workload.
The optical layer is a commodity input, not the moat. The moat is PCIe remoting, policy-driven orchestration, observability, and lifecycle management — the software that turns a fabric into a fleet.
Tokens Per Dollar, Not GPUs Per Rack
The next 18 months will bring another wave of AI infrastructure spending. It will be tempting to score it by FLOPs per rack, GPUs per cluster, and megawatts commissioned. Those numbers still matter, but they do not answer the question operators and investors actually need to answer: how many useful tokens does a system produce per dollar of capital and operating cost, across a changing workload mix?
The systems that win will not be the ones with the biggest boxes. They will be the ones that can turn a diverse fleet of compute, memory, and connectivity into the right infrastructure for each job — then reclaim it for the next one. That is scale-across, and it is what Corespan is building.
Corespan Composer is available today for operators building composable AI infrastructure. To discuss scale-across for your fleet, contact the Corespan team.