Original benchmark dataset
Local LLM benchmarks
Tokens-per-second rows collected on owner hardware or from cited, reproduced sources. Every public row ships with a confidence badge so you know which numbers are owner-measured, independently reproduced, or source-backed.
Confidence engine
39 strong
Coverage map
0 wanted
Regression watch
version-aware
Reproduce
protocol + queue
Selected coverage slice
22 measured or reviewed / 0 estimated / 0 wanted / 6 unstarted
78.6% of 28 selected cells measured or reviewed
22 measured or reviewed cells, 0 estimated cells, 0 wanted cells, and 6 unstarted cells.
RunLocalAI measuredReproducedVendor-publishedCommunity-reviewedLow/unverified measuredEstimatedCritical wantedWantedNot started
Dataset provenance split
Public evidence split
- 1.Editorial (measured)39 (100%)Measured by RunLocalAI on owner hardware.
- 2.Reproduced0 (0%)Independently reproduced by an operator on similar hardware.
- 3.Community reproduced0 (0%)Submitted by an operator and reproduced before public evidence use.
- 4.Excluded estimates0 (0%)Formula-derived rows are excluded from public benchmark evidence. How we estimate →
Dataset provenance heuristic
How strong is the visible provenance?
- 1.Very high0 (0%)Owner-measured AND reproduced by an independent operator.
- 2.High39 (100%)Owner-measured single-source OR community independently-reproduced.
- 3.Moderate0 (0%)Community-submitted: editorially approved, reproduced once.
- 4.Low0 (0%)Estimated / unverified / single anecdote — directional only.
/benchmarks/[id] detail pages. See Confidence methodology.Benchmark aging heatmap
How fresh is the dataset, month by month?
- empty
- 1-2 rows
- 3-5 rows
- ≥6 rows
- stale (18m+)
Top hardware by median tokens / sec
Throughput leaderboard
- NVIDIA GeForce RTX 5080Editorial
- NVIDIA GeForce RTX 3080 16GB (Mobile)Editorial
Latest 39 runs
Sorted by date. Click a model or hardware name to drill into the full record.