GPU benchmark
How Kielo measures your hardware for routing, model fit, and earnings projection.
Why benchmark?
Raw hardware specs are not enough — the benchmark runs real inference on your GPU to measure throughput and latency. The result drives:
- Model recommendations (what fits and how fast)
- Routing priority and job matching
- Earnings projections during onboarding
- Your public capability tier visible to the network
When it runs
- Onboarding step 4 — automatic during first-time setup (onboarding flow)
- Re-benchmark — from the desktop app or web dashboard at /provider/benchmark after hardware changes or driver updates
What is measured
| Metric | Description |
|---|---|
| Tokens per second | Sustained generation throughput during the test run |
| Time to first token | Latency from prompt submission to first output token |
| Peak VRAM | Maximum GPU memory used during inference |
| GPU utilization | Compute load during the benchmark |
| Temperature / power | Recorded when hardware sensors are available |
Benchmark score
Results are combined into a normalized score from 0 to 10,000. Tokens per second is weighted most heavily; high latency and excessive VRAM pressure relative to capacity reduce the score.
User experience
- A small test model downloads on first benchmark (cached for future runs)
- Progress UI: model loading → warm-up → throughput measurement
- Live metrics display during the run
- Cancel returns you to machine detection; partial results are discarded
Hardware scan (step 3)
Before benchmarking, onboarding runs a machine detection pass that collects GPU name, VRAM, CPU, RAM, disk, OS, and a quick network speed test. Benchmark requires GPU, CPU, and RAM detection to succeed.
Tips for best results
- Close other GPU-heavy applications before benchmarking
- Plug in laptops on battery saver off
- Ensure adequate cooling — thermal throttling lowers scores
- Re-benchmark after driver or OS updates