Blog

Deep dives into inference performance, power telemetry, the compile-and-serve layer, and tokens per watt. Written by the engineers doing the work.

Describe your image

First Fable, then NVIDIA and HuggingFace. Is it time for Owned Intelligence?

It seems that we are at an inflection point, perhaps even a tipping point. Companies are re-evaluating their AI strategy and asking an important question: is it …

Read Story

3 min read
Sep 11, 2026

GPU waste, tokens per dollar, and the impact on inference businesses

The cost of AI is in the spotlight, and every part of the AI industry has at least some focus on improving efficiency to reduce that cost.Model builders, improv …

Read Story

3 min read
Sep 2, 2026

Adaptive SM clocking for energy-efficient LLM serving

LLM serving is moving from occasional jobs to persistent service infrastructure. At that scale, GPU power is no longer a background detail. It sets thermal limi …

Read Story

9 min read
Jun 1, 2026