GPU benchmark

How Kielo measures your hardware for routing, model fit, and earnings projection.

Why benchmark?

Raw hardware specs are not enough — the benchmark runs real inference on your GPU to measure throughput and latency. The result drives:

  • Model recommendations (what fits and how fast)
  • Routing priority and job matching
  • Earnings projections during onboarding
  • Your public capability tier visible to the network

When it runs

  • Onboarding step 4 — automatic during first-time setup (onboarding flow)
  • Re-benchmark — from the desktop app or web dashboard at /provider/benchmark after hardware changes or driver updates

What is measured

MetricDescription
Tokens per secondSustained generation throughput during the test run
Time to first tokenLatency from prompt submission to first output token
Peak VRAMMaximum GPU memory used during inference
GPU utilizationCompute load during the benchmark
Temperature / powerRecorded when hardware sensors are available

Benchmark score

Results are combined into a normalized score from 0 to 10,000. Tokens per second is weighted most heavily; high latency and excessive VRAM pressure relative to capacity reduce the score.

User experience

  1. A small test model downloads on first benchmark (cached for future runs)
  2. Progress UI: model loading → warm-up → throughput measurement
  3. Live metrics display during the run
  4. Cancel returns you to machine detection; partial results are discarded

Hardware scan (step 3)

Before benchmarking, onboarding runs a machine detection pass that collects GPU name, VRAM, CPU, RAM, disk, OS, and a quick network speed test. Benchmark requires GPU, CPU, and RAM detection to succeed.

Tips for best results

  • Close other GPU-heavy applications before benchmarking
  • Plug in laptops on battery saver off
  • Ensure adequate cooling — thermal throttling lowers scores
  • Re-benchmark after driver or OS updates

Next