top of page


Qwen3.8-27B Benchmarks: Throughput, Latency and KV Cache
Throughput, latency and KV cache headroom for Qwen3.8-27B measured on a single RTX PRO 6000, comparing BF16, FP8 and a serving baseline across context lengths and batch token settings.
3 days ago


Infrastructure for the AI Agent Era (Part 2): Why Consumer GPU Clouds Are the Answer
Part 1 looked at why AI agents are a 'program-shaped workload' that breaks the assumptions of existing serving systems, and at how the world's top systems conferences are approaching the problem. As the CacheBlend case showed, much of the answer depends on resources outside GPU memory, namely the DRAM and SSD that consumer PCs already have. Let's return to the cost question and continue from there, looking at why that direction converges on the picture of a "consumer GPU clou
Aug 19


Infrastructure for the AI Agent Era (Part 1): Agents Are Programs
Hello. I'm Yongrae Cho, and I lead AI infrastructure research and development on AIEEV's engineering team. I hold a PhD in Computer Science focused on the scalability and reliability of distributed systems, and I now work on the efficiency and reliability of LLM serving on Air Cloud, along with research and development in AI agent technology. As more services, like AI agents, call a model dozens or even hundreds of times for a single request, the inference bill grows far fast
Aug 12


How to Turn GPU Resources into an Inference API
The Distributed GPU Cloud Story and Why Ray Is at the Center of It 💡 Core Message "A GPU that never serves a request has no business value. " Air Cloud connects everything from the runtime layer to the platform layer, so hardware actually reaches users as a real service. Introduction When people talk about AI infrastructure today, the conversation usually starts with GPU scarcity. How many H100s did you lock in? Is B200 supply going to loosen up? Does your data center have e
May 29


Two Technologies That Reduce AI Model Deployment Costs: Quantization and Prefix Caching
Hi, I'm Jinbeom Kim, a Software Developer on the AIEEV Dev Team. I studied computer science through both undergrad and graduate school, and I've been with AIEEV since the early days of the company — working on how we can operate more distributed GPU resources efficiently within Air Cloud 😊 In this post, I want to walk through two techniques we regularly evaluate when thinking about how to deploy AI models more efficiently. The first is Quantization — a method for reducing me
May 7


AirCloud April Update
AirCloud's April release is built around one goal: making it faster to run AI workloads, more reliable to operate them, and more flexible to put your existing GPU resources to work. This update includes enhanced Air Container operations, the general availability of Air API, Resource Provider (RP) support, and the introduction of an intelligent scheduler. Developers can now handle container access, log monitoring, error response, and API integration more seamlessly. Enterprise
Apr 29


One Command, Done: Integrating Air API with a ClawHub Plugin
Hi,I’m CY Lee from the DevOps/SRE team. With the launch of Air API, we’ve been building out our internal infrastructure monitoring system. Along the way, we developed an OpenClaw plugin—and in this post, I’d like to walk you through what we built and why it matters. 🙂 Before We Start If you’ve used OpenClaw for a while, you’ve probably experienced something like this at least once. The moment you try to connect an external model provider, you find yourself going through the
Apr 16
bottom of page
