top of page


Qwen3.8-27B Benchmarks: Throughput, Latency and KV Cache
Throughput, latency and KV cache headroom for Qwen3.8-27B measured on a single RTX PRO 6000, comparing BF16, FP8 and a serving baseline across context lengths and batch token settings.
3 days ago


Why Compute Is Scattering Again
If you have been following AI over the past week, you have probably come across at least one of the following stories. Anthropic, which is preparing for an IPO, is reportedly weighing whether to present investors with a potential market of more than $30 trillion The "software apocalypse" argument that AI will dismantle the existing software industry, and the counterarguments to it NVIDIA unveiling an entry level edge AI module that delivers 78 TOPS of compute Capital markets,
Sep 2


How AI Infrastructure Is Reshaping the Future City
What did the city of the future look like in your childhood imagination? Maybe it was robots collecting trash in apartment complexes, self-driving vehicles gliding through traffic-free streets, or an AI assistant like Iron Man's J.A.R.V.I.S. handling your every need. In WALL-E, humans drift through life on floating chairs inside a giant spaceship, living almost entirely inside a virtual world. These visions differ wildly in their details, but they share one constant: AI is wo
Jun 11


Two Technologies That Reduce AI Model Deployment Costs: Quantization and Prefix Caching
Hi, I'm Jinbeom Kim, a Software Developer on the AIEEV Dev Team. I studied computer science through both undergrad and graduate school, and I've been with AIEEV since the early days of the company — working on how we can operate more distributed GPU resources efficiently within Air Cloud 😊 In this post, I want to walk through two techniques we regularly evaluate when thinking about how to deploy AI models more efficiently. The first is Quantization — a method for reducing me
May 7


How Many Tokens Per Month Before Self-Hosting Your GPU Becomes Cheaper?
If you've been running an AI service for any length of time, you've probably hit this question at some point. "Is using an API actually the cheaper option? Or would it be better to just buy a GPU and run it ourselves?" As model performance converges, cost has become the decisive battleground. Teams at every scale are starting to run the numbers on which approach is actually cheaper for their usage volume — and the answer changes significantly depending on how much you're act
Apr 14


Air API is Now Live
If you've ever tried serving an open-source AI model yourself, you know the pain. Setting up GPU infrastructure takes longer than choosing the model itself. Provisioning GPUs, configuring environments, scaling with traffic... the road to running a single model is way too long. Air API eliminates that entire process. It's a serverless API service for open-source AI models. No infrastructure to build. Just an API key to get started. Key Features 💡 OpenAI-Compatible Endpoint
Apr 9
bottom of page
