By Jerry Chang, Head of Marketing, Americas & Europe, MediaTek.
AI performance is now defined at the server rack, where compute, memory, interconnect, power, and economics converge.
Until recently, performance improvements were driven at the chip level. Faster transistors, denser integration, and smarter architectures were enough to keep pace with demand.
In the AI era, that is no longer sufficient. What matters is not the performance of any single chip, but how effectively an entire system works together. And that system doesn’t stop at the board level or even the compute system. It extends to the server rack—the level at which compute, memory, interconnect, power delivery, and cooling converge into a single operational unit.
- Chips: Individual components.
- Systems: Integrated compute systems (XPUs, CPUs, scale-out fabrics).
- Racks: Fully realized infrastructure units.
The rack is where performance becomes real. This is where efficiency is determined and, increasingly, this is where value emerges. The message is clear: “Value is created at the rack, not just the chip.”
Data Center Metrics That Matter
In this new AI world, traditional metrics like raw TOPS (trillion operations per second) and FLOPS (floating-point operations per second) are no longer sufficient. A system can deliver impressive peak performance on paper while falling short in real-world deployment.
For example, a system may boast extraordinary peak TOPS based on idealized compute throughput, but large-model inference is not a benchmark—it’s a continuous, system-level workload. Tokens are generated only as fast as data can be moved between compute, memory, and interconnect. If memory bandwidth can’t keep up or interconnect latency gets in the way, expensive compute units would sit idle. Add in real-world power and thermal constraints, and the gap widens further. The result is that a system delivering impressive numbers on paper may achieve only a fraction of that performance in practice. In the AI era, peak performance is easy to quote—sustained, usable performance is what really counts.
This is because real-world AI workloads do not operate in isolation. Those workloads are constrained by power budgets, thermal limits, memory bandwidth, and communication overhead. That’s why the industry is shifting toward more meaningful metrics:
- Performance per watt.
- Performance per total cost of ownership (TCO).
Or, more concretely:
- How many useful tokens can you generate per joule?
- How many tokens can you deliver per dollar of infrastructure?
These are the metrics that determine whether an AI system is not just computationally powerful, but also economically viable.
As Dr. Rick Tsai, Vice Chairman and CEO of MediaTek, emphasized in his ISSCC 2026 keynote, Advancing Horizons for AI: Perspectives on Semiconductor Innovations, that AI-driven infrastructure spending is exploding. With investments potentially approaching $1 trillion by the end of the decade, he argued that efficiency is no longer just desirable but essential.
If you can’t optimize for performance per watt and performance per TCO, you don’t scale. If you don’t scale, you can’t compete. And if you can’t compete, you won’t last.
The Energy Wall Is Real
Another force shaping this transition is energy. AI is rapidly becoming one of the largest drivers of global data center power consumption. In some projections, a single large-scale AI data center can consume on the order of a gigawatt, which is comparable to the power usage of an entire city (the city of San Francisco consumes an average of 1 GW per day).
At the same time, global energy production is not scaling at anything like the same rate. This creates what MediaTek describes as a structural “energy wall.”
To make this more concrete, consider a data center operator with a fixed 50-megawatt power budget—a hard limit imposed by the local grid. The facility is already running near capacity, but the business opportunity is clear: deploy more AI infrastructure to meet growing demand.
Using conventional rack designs, adding new capacity quickly runs into that power ceiling. There’s simply no headroom left without expanding the facility or securing additional grid capacity—expensive and time-consuming propositions.
Now consider the impact of improving performance per watt at the rack level. With a more optimized architecture, the operator can deliver significantly more AI throughput within the same 50-megawatt envelope—effectively increasing revenue-generating capacity without increasing power consumption.
In this world, efficiency isn’t just a technical advantage—it’s a business multiplier. It’s no longer enough to build faster systems. We need to build more efficient systems, in which every watt is carefully budgeted, and every joule is put to productive use. And this level of efficiency can be achieved only through holistic design.
The Rack as a System-Level Architecture
To understand why the data center rack has become the focal point of AI compute, we need to consider what lives inside it. A modern AI rack is not simply a homogeneous block of compute. It’s a carefully orchestrated system composed of:
|
System Element |
Role in the Data Center Rack |
|
CPU |
Handles control, orchestration, and system management. |
|
DPU |
Manages data movement, networking, and offloads infrastructure tasks. |
|
XPU |
Delivers the core AI acceleration for training and inference workloads. |
|
HBM (High-Bandwidth Memory) |
Within the XPU, HBM provides the high-throughput data access required to keep AI compute engines fully utilized. |
|
Scale-Up Interconnect |
Links compute elements within a system or server rack, enabling high-bandwidth interconnect capabilities and low-latency communication between XPUs and memory. |
|
Scale-Out Interconnect |
Connects AI compute systems and racks across the data center, allowing workloads to scale efficiently across distributed infrastructure. |
|
Power and Thermal Systems |
Ensure stable operation within power and cooling limits, directly constraining achievable performance and density. |
All these elements must be co-designed and co-optimized. A bottleneck anywhere—memory bandwidth, interconnect latency, power delivery, or cooling—degrades the entire system's performance.
In other words, you don’t benefit from the performance of your best component; rather, you’re limited by the performance of your weakest constraint.
Beyond the XPU: From Infrastructure to Monetization
All of this is why MediaTek’s approach is explicitly framed as “beyond the XPU.” While the XPU remains the heart of the system, it is no longer the whole story. Real innovation lies in how compute, memory, interconnect, packaging, and system integration come together as a unified architecture—spanning from silicon to rack.
MediaTek’s ability to co-optimize across this entire stack—from custom ASICs and advanced packaging to optical interconnect and full rack-level integration—reflects a System Technology Co-Optimization (STCO) approach, in which silicon, packaging, memory, connectivity, and software are engineered as a unified system. This positions MediaTek to help its customers turn infrastructure into a strategic advantage, not just a cost center
Ultimately, this shift is about more than engineering. It’s about economics. AI is no longer a research project. It’s an operational business, and at-scale success is determined not by peak performance but by sustainable, efficient throughput.
The rack is where the transformation happens. It’s where abstract compute becomes deployable infrastructure; it’s where performance meets reality; and, increasingly, it’s where AI is monetized.
Rise of the System to Define the AI Era
If the past 50 years were defined by the rise of the chip, the next decade will be defined by the rise of the system, or more precisely, the rack.
In the AI era, the computer is no longer just something you hold in your hand or place on your desk. It spans everything from intelligent edge devices to seemingly endless rows of racks in vast data centers that consume megawatts of power and operate at a planetary scale.
In this new landscape, success will not be defined by excellence in any single component. It will be defined by the ability to co-optimize across the entire system—from silicon and packaging to interconnect, power, and rack-level integration.
This is where MediaTek stands apart. With deep expertise across the full stack and growing momentum in real-world deployments, MediaTek is uniquely positioned to translate architectural innovation into scalable, production-ready AI infrastructure.