LocalBench
How well local AI models from ollama.com code, reason and run on one GPU. Every score is measured automatically: code is run against tests, and answers are checked against known results.
Leaderboard
Overall is the average of the seven category scores, from 0 to 100. Coding averages debugging, code generation and refactoring. Click any model for its full breakdown.
By category
The 10 best models in each category for the current filter. Thin lines show the margin of error; when two models' lines overlap, the difference may just be noise.
Speed and efficiency
Measured on this machine, one request at a time. Higher tokens per second means faster answers.
Model detail
UI gallery
Real web pages the models built from the same request, shown live. These are the top scorers for the current filter.