Blog6 min read

The Token-Cost Trap: One Rack, Every GPU, Yours to Compose

Disaggregated inference locks the attention-to-feed-forward ratio into the silicon layout. Mercury Rack turns that ratio into a Composer policy instead.

Bill Koss - CEO and President of Corespan Systems

There is a comfortable story in AI infrastructure right now that sorts silicon into two lanes. One lane makes cheap tokens at scale. The other lane makes fast tokens for interactive workloads. Pick a lane, pick a chip, pick a customer, buy the matching rack.

The last two quarters of public filings say that story is already breaking, and the emerging industry answer — disaggregated inference — is quietly making the case for something Corespan has been building all along. The fix is not another purpose-built rack. It is a composable one. Corespan calls it the Mercury Rack.

The Lane Story is Collapsing in Public

The dominant wafer-scale inference vendor spent the June 2026 quarter turning into a cloud. Hardware sales fell to $54.1 million while cloud and other services grew to $126.0 million, roughly 70 percent of the quarter, and gross margin dropped to about 14 percent as the company began passing through data-center costs to serve a single very large customer (Cerebras 10-Q via Stocktitan). Contracted future work stood at $25.4 billion, most of it tied to one 750 MW commitment with an option to reach 2.0 GW by 2030 (Yahoo Finance).

The dominant throughput vendor spent the same window buying its way into the interactivity lane. NVIDIA's July 2026 filing carries a $2,944 million line related to a non-exclusive license agreement with Groq, on top of a $14.4 billion goodwill entry and a $2.5 billion developed-technology asset booked earlier for the same deal (NVIDIA 10-Q via SEC, NVIDIA FY26 10-K excerpt). On August 24, 2026, NVIDIA announced that Groq 3 LPX, a rack-scale interactive-inference accelerator that bolts onto Vera Rubin NVL72, had entered full production, with Nebius as first cloud customer (NVIDIA newsroom, CNBC).

Read those two disclosures next to each other and the neat lane assignment falls apart. The throughput vendor paid a full acquisition's worth of capital to own an interactivity rack. The interactivity vendor traded its margin to own a slice of hyperscale batch capacity. Both are trying to escape the lane they were sorted into, and both are paying full price for the ticket.

Disaggregation is the New Consensus — and its Hidden Liability

The industry's answer to this mismatch has a name now: disaggregated inference. Split the model's attention work from its feed-forward work, put each on the hardware it prefers, pipe them together at rack scale. NVIDIA and Groq are doing it. Cerebras and AMD announced a similar collaboration. Serious analysts describe both approaches as the direction of travel.

They are right that the gains are large. They are also honest about the catch: pipelining attention and feed-forward across dedicated hardware fixes the ratio between them at the moment the rack is built. A workload with a 1,000-token context and a workload with a 1-million-token context put wildly different demands on attention. A rack whose attention-to-feed-forward ratio is baked into the silicon layout works beautifully for the workload it was designed for, and degrades quickly for anything far outside those bounds.

That is not a wafer problem. It is not a GPU problem. It is a rack architecture problem. It is the same problem the throughput and interactivity vendors already ran into when they tried to serve each other's customers.

Models are Already Being Shaped Around the Rack

The fixed-ratio problem gets worse because model designers know exactly which rack won the last round. Sparse mixture-of-experts models are being tuned to the multiplication-unit widths of the dominant GPU. Expert counts, expert widths, and sparsity patterns are all set to what runs fastest on that specific silicon. Move the same model to a fundamentally different architecture and it slows down, sometimes drastically.

Meanwhile, hardware performance on already-shipped racks keeps improving through software — kernel optimizations, better speculative decoding, more aggressive multi-token prediction. Realistic estimates suggest 50 to 80 percent of performance is still available on existing hardware and existing models, before anyone buys another rack.

Put those two facts together and the "buy a second, purpose-built rack" answer looks even weaker. Operators are being asked to double their facility footprint to cover a workload lane whose optimal ratio will move as the models move, while sitting on racks that are not yet saturated.

Mercury Rack: One Rack, Every GPU

Mercury Rack is Corespan's answer, and it starts from a different unit of capacity.

Figure 1 — Fixed-ratio rack vs Corespan Mercury Rack. Left: a single-vendor rack with attention and feed-forward silicon pipelined in a locked 24:48 ratio. Right: a Mercury Rack with nine PRU 2500 chassis holding a mixed pool of NVIDIA, AMD, Intel and other GPUs that Composer re-splits between attention and feed-forward on demand.

A single Mercury Rack is a 48U composable system built from nine PRU 2500 resource chassis, hosts, storage, optical fabric, and Corespan Composer, with 72 mix-and-match GPU slots, 111 kW of Air-plus-DLC power, and up to 1.2 PB of storage per PRU 2500 for in rack storage. The GPU slots accept NVIDIA, AMD, Intel and other accelerators side by side. Composer binds those resources into hosts on demand, and re-binds them as the workload changes. The headline is exactly what the product is: One Rack. Every GPU. Yours to Compose.

That gives operators four things a fixed-ratio rack cannot.

Disaggregation Without Lock-in. Mercury supports the same attention-plus-feed-forward split the industry is converging on, but the ratio between the two is a Composer policy rather than a silicon layout. When context lengths grow, or when a new model changes the mix, the operator re-composes the pool. The rack does not need to be rebuilt.

Workload Fit Without a Second Rack. A Mercury Rack can run a batch profile and an interactivity profile in separate partitions, or shift capacity between them across the day. The workload-shape mismatch that forces other operators to buy a companion rack becomes a policy change.

Vendor Choice that Stays Open. Mercury is designed around heterogeneous GPU pools. When a new interactive part or a cheaper batch part changes the math, the buyer adds it to the pool instead of buying another rack. Model designers stop dictating the silicon roadmap by proxy.

An Asset that Keeps its Value. A rack whose ratios are locked to one vendor depreciates against one roadmap. A Mercury Rack that can absorb multiple generations of accelerators, from multiple vendors, across multiple workload shapes, holds its worth across the roadmaps that follow. That is what makes it financeable, and financing is what actually decides whether AI capacity gets built.

Where the Value Shows Up

The end buyer of fast inference is rarely a chip enthusiast. It is a firm whose people are expensive — quantitative trading desks, coding-assistant customers, consulting practices, executive teams — where saving a few seconds per iteration on a $500,000-a-year engineer pays for the silicon many times over. That customer does not care which lane the rack was sold in. They care that the tokens arrive fast enough at a price that fits, and that the capacity is still there when next quarter's model changes shape.

Fixed-ratio racks put a ceiling on that promise the day they are installed. Mercury Rack removes it. The GPUs inside it stay a decision. The attention-to-feed-forward split stays a decision. The workload mix stays a decision. And every one of those decisions is made in software, on a rack that was built to keep making them.

The token-cost debate keeps being told as a fight between throughput silicon and interactivity silicon. On the ground, the number that decides whether an inference business works is utilization: how much of the rack is doing useful work, and how quickly the operator can point that work at whichever accelerator is best for the moment. That is the shift Corespan is building toward, and it is why Mercury Rack is the product the market has been quietly asking for through two very expensive quarters of filings.

Corespan Systems designs photonic-native, composable GPU infrastructure. Mercury Rack is built on the PRU 2500 resource chassis, FIC 2500 fabric card, and Corespan Composer control plane. Product and partner detail is at corespan.ai.