The vLLM doctor

Know if a config change will actually help — before you ship it.

Point InferPilot at a running vLLM. It reads the live /metrics, tells you your real bottleneck, and whether a change like fp8 KV cache will help — or do nothing. It diagnoses from evidence, not GPU utilization, and says so when it can't tell.

$ uvx inferpilot doctor --url http://localhost:8000

Open source · MIT · no telemetry · no GPU needed to run the diagnosis

InferPilot diagnosing a live vLLM in the terminal

One scrape of /metrics, one honest verdict

GPU utilization sits near 100% whether you're genuinely compute-bound or just thrashing on KV cache. It can't tell these apart. Preemptions and queue depth can — so InferPilot reads those and names the regime.

⚡ KV-BOUND & PREEMPTING

fp8 KV cache → worth testing

KV cache is full and the scheduler is preempting and recomputing tokens — wasted work fp8 can relieve.

✗ COMPUTE / OTHER-BOUND

fp8 won't help — don't bother

A queue is building while KV has headroom. The GPU math is the limit. The honest "no" most tools won't give you.

◐ NEAR CAPACITY

watch the preemption counter

KV is full but not preempting yet. fp8 is pre-emptive insurance, not a measured win.

✓ HEALTHY

no KV lever warranted

KV has headroom and nothing is queuing. If latency matters, that's a separate, low-load question.

Try the diagnosis

Move the signals your vLLM server would report and watch the verdict change. Two servers can show the same ~99% GPU utilization and get opposite answers — that's the whole point.

Observed signals

What a measured window of your server reports.

99%
6

InferPilot verdict

Recomputed live — or it abstains.

…
…
KV cache
preemptions

Measured, not hypothetical

A paired study on Qwen2.5-3B (A10G, vLLM 0.29.0) using an exact recomputed-token counter added to vLLM (an open PR to vLLM core). In the preempting regime, fp8 drove recomputed tokens to zero and throughput up. Total GPU cost: $0.81.

Recomputed tokenswasted work
bf16 KV · 29,826
fp8 · 0
Throughputtokens / sec
bf16 · baseline
fp8 · +36%

Honest caveat: on a 7B model the throughput win held (+33%) but recompute did not go to zero — the mechanism is model-size dependent, and InferPilot reports that instead of overclaiming. The quality preflight also detected distribution shift, so fp8 KV isn't called "lossless."

Install

Pure Python, pydantic-based. The diagnosis path needs no GPU and no vLLM install.

Run with zero install

uvx inferpilot doctor --url …

Keep it around

uv tool install inferpilot

Classic

pip install inferpilot

Then the daily loop: doctor to glance, doctor --watch to monitor, capacity to turn a rate sweep into an SLO-capacity ceiling, a $/token, and an action plan.

copied