The boundary of the computer is changing again. Not because machines are becoming smaller, and not because they are becoming bigger, but because the line around what counts as one machine has stopped being fixed by the server enclosure.
At Corespan Systems, we believe the next chapter belongs to composable infrastructure: systems large enough to support enterprise-scale agentic work, yet flexible enough to give each workload only the resources it needs.
That future has a familiar foundation. Shared capacity, coordinated resources, and disciplined utilization were central to early computing. The opportunity now is to recover those principles without recreating the rigid architecture of the past.
From a Shared Machine to a Personal Device
The mainframe era established a powerful economic principle: expensive computing resources become more useful when many people can access them. Time-sharing systems allowed multiple users to work interactively on a central machine, giving each the experience of dedicated access without requiring a separate computer for every person (Computer History Museum).
The lesson was not simply that large machines were powerful. It was that allocation mattered. A valuable resource should spend its time doing useful work, not waiting for its owner.
Our graphic traces the subsequent movement through mid-range systems, workstations, servers, PCs, and mobile devices — approximate, overlapping eras rather than a performance scale, as the caption notes. Smaller devices did not mean less capable computing.
The rise of inexpensive microprocessors made dedicated personal computing practical, while the later growth of cloud computing brought shared-resource concepts back into prominence (IBM). That history offers a useful perspective on the AI era: computing does not move in one direction forever.
AI Turns the Curve Upward
On the right side of the graphic, the emphasis changes from smaller devices to larger accelerator systems. Two GPUs become four, then eight. The question shifts from how much computing fits on a desk to how much acceleration a system can coordinate.
Eight GPUs represent a familiar server building block, not an industry-wide maximum: NVIDIA documents both eight-GPU HGX baseboards and much larger rack-scale systems. The boundary highlighted in our graphic is therefore an architectural convention to move beyond, not a claim that nobody has built a larger machine.
For Corespan, the more important question is what happens when a workload does not fit the configuration that was purchased. Imagine an organization with one application that needs more GPUs, another that needs fewer, and a third that needs additional storage rather than additional compute. Buying the same fixed combination for all three is administratively convenient. It is not necessarily the best use of capital, power, or capacity. We should not confuse a larger machine with a better-matched system. The next step is to make the system's shape adaptable.
Agentic Scale Does Not Mean Maximum-Sized Everything
Agentic computing makes that distinction especially important. Anthropic describes workflows that route easier tasks to smaller, cost-efficient models and harder tasks to more capable ones.
Consider an enterprise assistant that classifies a request, retrieves documents, extracts information, performs deeper reasoning, and checks a result. There is little reason to assume every stage should receive the same accelerator allocation. Some stages may justify substantial GPU resources; others may be better served by a smaller model or CPU-based processing.
Our position is straightforward: the hardware strategy should reflect the same discipline as the application strategy. Agentic scale should mean supporting more useful work, not automatically assigning the largest accelerator to every task.
Muse Glimmer Makes the Point Concrete
Meta introduced Muse Glimmer as a 30-billion-parameter model optimized for local agentic workflows — function calling, coding, and LLM-as-a-judge evaluation — running on a Mac or PC with a single consumer GPU. That is a useful counterpoint to the assumption that capable agents must always require the largest accelerator systems.
Meta describes approximately 4-bit quantization bringing the language model under 20 GB, which leaves headroom for the model's working memory, the perception encoder, and the speculative decoding drafter to run together inside a 24 GB or 32 GB envelope. AMD reports running Muse Glimmer on a single Radeon AI PRO R9700 using llama.cpp.
That card is familiar to us. We recently ran eight Radeon AI PRO R9700s as a single fabric in a PRU 2500, holding roughly 111 GB/s bidirectional on every GPU-to-GPU pair under full-mesh load. The same accelerator that can host a capable agent on its own also pools cleanly — which is this entire argument expressed in one piece of silicon.
Our takeaway is architectural, not a claim of a Corespan benchmark on Muse Glimmer or of a Meta deployment. If a suitable agent workload can run on one appropriately sized GPU, enterprise expansion can mean operating more such instances from a coordinated pool, rather than making every instance larger.
Imagine a pool serving coding agents, document assistants, and evaluation workers alongside larger reasoning workloads. The opportunity is to match each service to suitable resources and adjust allocations as demand changes. Model fit still requires workload testing; single-GPU feasibility is not a guarantee of production concurrency or latency.
Compose the System Around the Workload
This is the idea behind DynamicXcelerator™: disaggregate resources into pools and compose them into right-sized environments, rather than permanently binding each resource to a fixed server configuration. Our design principle is to start with the workload and make the infrastructure follow.
The graphic expresses that direction through configurations of 16 and 32 GPUs per CPU, with 64 GPUs per CPU explicitly identified as a 2027 roadmap goal. Those numbers describe the composition and scaling story, not a promise of proportional application performance.
The architectural foundation combines resource pooling, photonic connectivity, and software-controlled allocation of PCIe devices, including GPUs and NVMe storage. The purpose is to make the server enclosure less decisive in determining which resources can work together.
Corespan Composer provides the control plane, managing how disaggregated physical resources are attached to and detached from servers as requirements change. Put simply, our approach asks not only where a workload should run, but what machine should exist beneath it.
Scheduling work onto a fixed machine and changing the resources that constitute the machine are different design choices. We believe effective AI infrastructure needs both.
Storage Belongs in the Same Conversation
The graphic deliberately includes high-density SSDs alongside GPU pools. We do not view storage as an afterthought to be added once the accelerator count has been selected.
Consider two services with the same GPU requirement. One repeatedly serves a compact model; the other must work across a substantial document collection. Even if their accelerator counts match, their storage-capacity and data-access requirements may differ.
Corespan's resource-pooling approach includes NVMe devices as well as accelerators, allowing storage to participate in system composition rather than remain inseparable from one host. That gives infrastructure designers another variable to match to the application.
The principle is balance, not simply density. We would rather ask whether a configuration supplies the right combination of compute and data access than celebrate a GPU count in isolation. A well-composed system should be judged by the work it completes.
The Modern Mainframe Is a Principle, Not a Box
This is where old becomes new again. The original time-sharing insight was that a shared machine could serve many users efficiently; our interpretation for AI is that a shared resource pool should support many appropriately composed machines.
The analogy has limits, and those limits are important. We are not proposing a return to one inflexible configuration for every application. Nor should pooling be interpreted as a guarantee that every combination of devices will behave like one giant accelerator.
Our design standard is workload fit. Before choosing a configuration, ask what the application needs, how it moves data, what software it supports, and what service level it must meet. Then evaluate the system against those requirements rather than against a device-count headline.
For example, an operator should be able to consider several smaller inference environments instead of automatically assigning an entire large configuration to one service. When a larger allocation is justified, the architecture should support that decision too. The important feature is the ability to choose.
Scale the System, Right-Size the GPU
We believe the next phase of AI infrastructure will reward a different purchasing question. Instead of asking only, “How large a system can we buy?” organizations should ask, “How effectively can this infrastructure adapt to the work we need to do?”
That is the message of the curve. The return to larger systems need not mean a return to rigid systems. It can mean recovering the economics of shared computing while preserving the flexibility to compose different environments for different needs.
At Corespan Systems, our direction is clear: compose pools of right-sized GPUs for agentic workloads, rather than sizing every task for the largest accelerator. The Muse Glimmer example reinforces our conviction that a larger shared system and smaller, purpose-fit compute allocations belong in the same architecture.
Corespan. Scale the system. Right-size the GPU. Talk with our team about composing your next AI environment around the workload, not the server boundary.