The problem being solved
You want to know whether your expensive GPUs are earning their keep. The obvious number — “GPU utilization” — is misleading, and understanding why is the whole point.
Layer one: is it allocated?
The first question is dull and it’s where most of the waste hides.
Is this GPU assigned to anyone at all?
Not “is it working hard” — is anybody’s job even on it. A GPU sitting in a rack with nothing scheduled to it is 100% waste, and it doesn’t show up in any performance metric, because there’s no work to measure.
This is what the Cast AI study found across 23,000 clusters. Average enterprise GPU utilization near 5% is mostly an allocation problem, not an efficiency problem.
Parking garage analogy: first ask how many spaces are empty. Not how skilfully each parked car is positioned.
Layer two: is the work efficient?
Now the GPU is allocated and running something. Is it running it well? This is where “GPU utilization” — the headline number — starts to lie.
Why the headline number lies
The standard GPU utilization metric measures the percentage of time during which at least one kernel was running.
A kernel is a unit of work sent to the GPU — “multiply these matrices.”
So the metric is answering: was the GPU doing anything at all? Not: was it doing much.
Imagine a factory with 10,000 workers. The metric asks “was at least one worker moving?” If one person is sweeping a floor and 9,999 are standing still, the factory reports 100% utilization.
That’s genuinely how it works. A job can show 95% utilization while using a tiny fraction of the chip.
SM occupancy — how much of the chip is busy
A GPU isn’t one processor. It’s thousands of small cores, grouped into units called Streaming Multiprocessors — SMs. A modern data-centre GPU has roughly 130 to 150 of them.
SM occupancy measures how many of those units are actually working, and how fully.
Back to the factory: utilization asks was anyone moving? Occupancy asks how many of the 10,000 workers were busy?
Why occupancy can be low:
The work is too small. You send a tiny batch. It only fills twelve SMs. The other hundred-plus sit idle — and the GPU reports near-100% utilization the whole time.
Each thread needs too much memory. Every SM has a fixed pool of fast local memory called registers. If each piece of work demands a lot, fewer pieces fit, and the SM runs half-empty.
The work doesn’t divide evenly. Some SMs finish early and wait for the stragglers.
The important caveat: maximum occupancy is not the goal. A well-optimized kernel can run faster at 50% occupancy than a poorly-optimized one at 90%, because occupancy measures how many workers are busy, not how much they accomplish. It’s a diagnostic, not a target.
Memory bandwidth utilization — is data arriving fast enough?
The second efficiency question, and usually the more important one.
A GPU has two resources: compute (the SMs doing arithmetic) and memory bandwidth (the rate data moves between the GPU’s memory and its cores).
Modern GPUs are extraordinarily fast at arithmetic. So fast that most real workloads are limited by memory, not by compute. The cores finish their arithmetic and sit waiting for the next batch of numbers to arrive. This is called being memory-bound, and it’s the normal state for large language model inference.
Restaurant analogy. Your chefs are world-class and can cook a dish in ten seconds. But ingredients arrive from the storeroom one tray at a time. The chefs spend most of their shift standing at the pass, waiting. Hiring faster chefs changes nothing. The storeroom is the bottleneck.
Memory bandwidth utilization tells you what percentage of the maximum data-transfer rate you’re achieving.
What the combination tells you
Reading the two together diagnoses the problem:
SM occupancy | Memory bandwidth | Diagnosis |
Low | Low | Something upstream is starving it. Storage, network, or the data loader. The GPU isn’t the problem at all. |
Low | High | Memory-bound. The chip is waiting for data. Larger batches, better data layout, or quantization. |
High | Low | Compute-bound. Doing lots of arithmetic. Usually fine and often what you want. |
High | High | Well-optimized. Rare, and worth understanding why so you can repeat it. |
The top-left row is the one that matters commercially. Low on both means the GPU is allocated, reporting high “utilization,” and accomplishing very little — because the bottleneck is somewhere else entirely. That’s the most expensive failure mode in AI infrastructure and the hardest to see.
Layer three: is the customer getting value?
Layers one and two are about the machine. Layer three is about the outcome. This is the only layer the customer cares about.
Tokens per second — for inference
A token is a chunk of text, roughly three-quarters of a word. When a language model generates a reply, it produces tokens one at a time. Tokens per second is the throughput of a running model. Two distinct versions matter:
Per user — how fast the text appears to one person. Below about 20 tokens per second feels sluggish to read.
Total across the system — how many users you can serve simultaneously on one GPU.
These trade against each other. Batching more users together raises total throughput and lowers per-user speed. That trade is a product decision, not a technical one — and it’s exactly the kind of thing a services practice advises on.
Samples per second — for training
A sample is one item of training data: an image, a document, a record. Samples per second is how fast the model is learning. It’s the direct measure of training throughput, and it’s what you watch when you change anything — batch size, precision, storage, network.
Epoch time — for training
An epoch is one complete pass through the entire training dataset. Epoch time is how long that pass takes. It matters because it converts directly into the thing the customer actually asks about: how long until my model is ready? Twelve epochs at four hours each is two days. Get epoch time to three hours and it’s a day and a half.
Why all three layers, not one
Each layer answers a different question, and a number at one layer can’t answer a question at another:
Allocation — are we wasting capacity? A finance question.
Occupancy and bandwidth — is the work efficient? An engineering question.
Tokens, samples, epoch time — is the customer getting value? A commercial question.
And here’s why they must be correlated rather than reported separately.
A customer complains that training is slow. Epoch time went from four hours to six. Why?
If allocation is fine and occupancy is low and bandwidth is low → something upstream is starving the GPUs. Look at storage.
If occupancy is low and bandwidth is high → memory-bound. Look at batch size.
If both are unchanged and epoch time still rose → the collective is slower. Look at the network.
Without all three layers on one timeline, you cannot distinguish these. You get three teams each reporting that their layer looks fine — which is the three-teams problem, arriving here in numerical form.
The version to say out loud
GPU utilization is a misleading metric — it measures whether at least one kernel was running, so a job can show ninety percent and be doing almost nothing. So I look at three layers.
Allocation — is the GPU assigned to anyone at all. That’s where the real waste is, and it’s what the five percent industry figure is actually measuring.
Then SM occupancy and memory bandwidth, for whether the work is efficient. Low on both usually means something upstream is starving it — storage or network, not compute. Low occupancy with high bandwidth means memory-bound.
Then a workload metric — tokens per second, samples per second, epoch time. That’s the only layer the customer cares about.
The gap between allocation and useful work is the number I’d want on a dashboard, and in most environments nobody has it.
Two follow-ups to have ready
“What’s a good occupancy number?”
There isn’t one, and that’s the point. Occupancy is a diagnostic, not a target — a well-optimized kernel at fifty percent occupancy can beat a badly-optimized one at ninety. I’d never set it as a goal. I’d use it to tell whether a slow job is starved, memory-bound, or compute-bound.
“How do you collect this?”
DCGM for the GPU layer — occupancy, bandwidth, clock throttle reasons, XID errors — exported into Prometheus, which the GPU Operator deploys. The workload metrics come from the framework or the serving layer. The hard part isn’t collection, it’s putting them on one timeline with fabric and storage metrics so you can correlate.

