Download optimized containers for your model / GPU / workload

Every container runs next to vLLM in your own cloud account. Same engine, same API, same weights. The serving configuration is what changes.

Browse the container library See how a container gets measured

How a number gets on a card

Every reduction on this page is measured against stock vLLM on the GPU named on the card, at a fixed latency target, with the same weights and the same request mix. Nothing is extrapolated from one GPU class to another.

A container appears in the library before its benchmark is finished. Until the measurement is signed off, the card shows the shape of the result and not the number.

Unit of analysis: model × precision × GPU class. A container measured on A100 does not carry its number to H100.

 

Your workload is not in the library yet

We measure your production baseline on your own GPUs first.
You keep the measurement whether or not you go further.

Talk to an engineer Request a model