top of page


Why Compute Is Scattering Again
If you have been following AI over the past week, you have probably come across at least one of the following stories. Anthropic, which is preparing for an IPO, is reportedly weighing whether to present investors with a potential market of more than $30 trillion The "software apocalypse" argument that AI will dismantle the existing software industry, and the counterarguments to it NVIDIA unveiling an entry level edge AI module that delivers 78 TOPS of compute Capital markets,
Sep 2


Infrastructure for the AI Agent Era (Part 2): Why Consumer GPU Clouds Are the Answer
Part 1 looked at why AI agents are a 'program-shaped workload' that breaks the assumptions of existing serving systems, and at how the world's top systems conferences are approaching the problem. As the CacheBlend case showed, much of the answer depends on resources outside GPU memory, namely the DRAM and SSD that consumer PCs already have. Let's return to the cost question and continue from there, looking at why that direction converges on the picture of a "consumer GPU clou
Aug 19


How AI Infrastructure Is Reshaping the Future City
What did the city of the future look like in your childhood imagination? Maybe it was robots collecting trash in apartment complexes, self-driving vehicles gliding through traffic-free streets, or an AI assistant like Iron Man's J.A.R.V.I.S. handling your every need. In WALL-E, humans drift through life on floating chairs inside a giant spaceship, living almost entirely inside a virtual world. These visions differ wildly in their details, but they share one constant: AI is wo
Jun 11


95% of your GPU is idle
You're only using 5% of the GPUs you paid for As of April 2026, GPU utilization in enterprise Kubernetes clusters ranges between 5% and 30%. Despite costing $2 to $15 per hour depending on the hardware, most GPUs remain idle for the majority of the time. According to a Cast AI report, companies are spending up to 20× more than what they actually need for GPU compute. In the race to adopt AI, many organizations secure GPU capacity “just in case.”But simply holding onto that ca
Apr 28


How Many Tokens Per Month Before Self-Hosting Your GPU Becomes Cheaper?
If you've been running an AI service for any length of time, you've probably hit this question at some point. "Is using an API actually the cheaper option? Or would it be better to just buy a GPU and run it ourselves?" As model performance converges, cost has become the decisive battleground. Teams at every scale are starting to run the numbers on which approach is actually cheaper for their usage volume — and the answer changes significantly depending on how much you're act
Apr 14


AI Infrastructure Must Go Beyond Geography
As geopolitical conflicts intensify, the limitations of centralized AI infrastructure are becoming clear.
This article explores why distributed infrastructure is emerging as a more resilient and necessary approach.
Apr 6


Google's TurboQuant — The Era of Serving LLMs Without Expensive GPUs Is Getting Closer
Google’s TurboQuant reduces KV cache memory usage in LLM inference without sacrificing accuracy. Learn why 80GB GPUs were needed—and why mid-range GPUs may now be enough.
Mar 30
bottom of page
