Six-month data centres, and AI on your desk
Three takeaways from Nvidia's AI Day Singapore.
At Nvidia's AI Day Singapore today, I got a look at the AI factory and saw where AI is headed: not just in the data centre, but on our desks. Here are three things that stood out.
The six-month data centre
Any data centre observer already knows how much the field has evolved. But even I was surprised when I learned from Nvidia's Terry Yin that construction time for new greenfield facilities could be cut to just six months.
And it's not just about build times. Nvidia is also serious about squeezing every iota of AI performance out of data centres. To that end, it is studying and has come out with designs around power orchestration software, next-generation AI architectures, new data centre designs, and a shift to 800 VDC, among other things.
The focus is clear: minimise energy waste, deploy AI data centres in the shortest time possible, and squeeze out the maximum tokens per megawatt (MW) of power.


Inside Vera Rubin
I saw photos of the insides of shipping Vera Rubin servers from Asus, and it really hit me why it is the AI system to get.
The Vera Rubin platform isn't an incremental update but a complete redesign: no fans, a much simpler layout, and far more tokens. On that last point, Nvidia says Vera Rubin generates up to 10 times more tokens per MW than the GB200 NVL72.
The simpler design matters because simpler deployments mean greater reliability in data centres. Full liquid cooling is also easier to support than hybrid cooling. Taken together, I expect Vera Rubin servers to be highly sought after.



Coming to a desk near you
Finally, I'm increasingly convinced that AI processing will not stay within data centres but will gravitate towards our desks and into our offices. Let me explain.
Subscription plans offered by the likes of Anthropic and OpenAI are heavily subsidised. And though AI models are improving rapidly, there is evidence that the tokens on offer are being reduced.
As more users and businesses find value in AI, token use will only surge. For some use cases, the point will come where it is more feasible to run inference locally than through AI providers. This is where AI desktops like the DGX Spark I've been using come in.
I saw the more powerful Asus ExpertCenter Pro ET900N G3 on display today. On a benchmark, it achieved more than 11,000 tokens per second running DeepSeek V4. The secret lies with its 784GB of onboard RAM, made up of both HBM and DDR for optimal performance and cost.
Sure, it has a six-digit price tag. But for a team of, say, 10 heavy users who might otherwise pay API rates, it could well work out to be the cheaper option.
The AI factory is getting faster to build and more efficient to run. But don't assume inference will stay locked inside it. Some of that processing may soon be happening a lot closer to your desk.

