> ## Content Index
> Fetch the complete content index at: https://www.techstories.co/llms.txt
> Use this file to discover other available public pages before exploring further.

# The real AI race isn't about deploying the most GPUs
- URL: https://www.techstories.co/the-real-ai-race-isnt-about-deploying-the-most-gpus/
- Published: 2026-09-04T11:44:37.000Z
- Updated: 2026-09-06T02:31:48.000Z
- Description: Why integration and efficiency stole the show at Asus AI Tech 2026.
- Author: Paul Mah
- Tags: AI, Data Centres

As the AI race accelerates, it's no longer about having the best model or the best data centre hardware. It's about how everything works together to maximise token output.

That was a recurring theme at Asus AI Tech 2026 in Seoul yesterday, which saw speakers from Asus, IBM, Nvidia and Samsung, among others, take to the stage.

![](https://www.techstories.co/content/images/2026/09/source-cd44bb9d74b394af.jpg)

****Photo Credit**: Asus

### Owning more of the stack

I've previously written about "useful AI", which I see as the tipping point for surging token use. Indeed, neocloud players I've spoken with over the last few months have told me of a sharp increase in demand.

For data centre operators used to working with multiple suppliers, Alber Wu of Asus argues that vertical integration, and owning more of the stack, reduces delays, cost and operational friction.

Unsurprisingly, Crusoe's Andrew Wee agrees. He thinks that owning more of the AI infrastructure stack, specifically the energy and the data centres, can accelerate deployment and reduce supply-chain risk.

### Where the losses happen

But how does one increase token yield? A good place to start is by ensuring that any losses in the process of token generation stay minimal.

On this, CH Hsieh of Asus shared the top ways these losses happen: power loss from grid to chip, performance loss from throttling, and production loss from slow deployments.

To reduce performance loss from throttling, he highlighted thermal solutions across the component, system and rack levels to sustain peak AI performance. These comprise microchannel cold plates at the component level, fully liquid-cooled servers at the system level, and coolant distribution units (CDUs) for multi-MW cooling at the rack level.

![](https://www.techstories.co/content/images/2026/09/1788522273820.jpg)

![](https://www.techstories.co/content/images/2026/09/1788522270771.jpg)

![](https://www.techstories.co/content/images/2026/09/1788522272152.jpg)

****Photo Credit**: Paul Mah

### Beyond the chatbot

While the attention is on large language models (LLMs), the reality is that AI demand is no longer confined to the likes of ChatGPT. Agentic AI and physical AI are increasingly taking centre stage, too.

One area highlighted was industrial autonomy, where AI-powered smart systems are deployed to reduce production time and improve the quality of outputs. This is probably an area worth watching in the months ahead.

### The real race

A few other themes caught my attention, and the ones I fully agree with are these. Output is increasingly judged by tokens per watt. The AI bottleneck is moving beyond the GPU. And hybrid cooling will be needed for the transition.

Ultimately, the AI infrastructure race is no longer limited to who can deploy the most GPUs, though I'm sure that still matters for the AI giants.

For the rest of us, it is more about who can convert electrons, data and capital into useful tokens quickly, efficiently and reliably.

One last thing: Asus says mass production of its Vera Rubin rack will begin "next month". Is your data centre ready?