The vLLM doctor
Point InferPilot at a running vLLM. It reads the live /metrics, tells you your
real bottleneck, and whether a change like fp8 KV cache will help — or do nothing. It diagnoses from
evidence, not GPU utilization, and says so when it can't tell.
Open source · MIT · no telemetry · no GPU needed to run the diagnosis

/metrics, one honest verdictGPU utilization sits near 100% whether you're genuinely compute-bound or just thrashing on KV cache. It can't tell these apart. Preemptions and queue depth can — so InferPilot reads those and names the regime.
KV cache is full and the scheduler is preempting and recomputing tokens — wasted work fp8 can relieve.
A queue is building while KV has headroom. The GPU math is the limit. The honest "no" most tools won't give you.
KV is full but not preempting yet. fp8 is pre-emptive insurance, not a measured win.
KV has headroom and nothing is queuing. If latency matters, that's a separate, low-load question.
Move the signals your vLLM server would report and watch the verdict change. Two servers can show the same ~99% GPU utilization and get opposite answers — that's the whole point.
What a measured window of your server reports.
Recomputed live — or it abstains.
A paired study on Qwen2.5-3B (A10G, vLLM 0.29.0) using an exact recomputed-token counter added to vLLM (an open PR to vLLM core). In the preempting regime, fp8 drove recomputed tokens to zero and throughput up. Total GPU cost: $0.81.
Honest caveat: on a 7B model the throughput win held (+33%) but recompute did not go to zero — the mechanism is model-size dependent, and InferPilot reports that instead of overclaiming. The quality preflight also detected distribution shift, so fp8 KV isn't called "lossless."
Pure Python, pydantic-based. The diagnosis path needs no GPU and no vLLM install.
Run with zero install
uvx inferpilot doctor --url …Keep it around
uv tool install inferpilotClassic
pip install inferpilotThen the daily loop: doctor to glance,
doctor --watch to monitor, capacity to turn a rate sweep into an SLO-capacity ceiling,
a $/token, and an action plan.