Editorial
RunLocalAI measured

Trendyol LLM Asure 12B on NVIDIA GeForce RTX 5080

Measured 3 mo ago.

Why trust this benchmark?

Measurement

tok/s
61.5
TTFT
323 ms
VRAM used
RAM used
Power
Quant
Q4_K_M
Context
8K
Run date
2026-05-27
Source
owner
Type
measured

V36.52 rigor detail

Protocol →
Cold-start decode
61.58 tok/sTTFT 6510 ms
Steady-state median
61.52 tok/sP5 61.5 · P95 61.6
Runs captured
3
Scenario
Single-stream
Editorial notes

3-run measured Ollama API median; variance 0.1%; raw logs in public Gist

Evidence

What this row provides for independent review. Missing fields lower confidence; they are shown explicitly instead of hidden.

Source link
Open source
Evidence manifest
Open manifest
Command
Available
Raw logs
Not provided
Raw results
1 files
Log hash
73c7fe1e1310...
Operator
fred-oline
Runtime
ollama version is 0.24.0
Driver
NVIDIA 595.97
CUDA / ROCm / Metal
Not provided
OS
Windows
Run count
3
Raw stats
Available
Environment notes
Not provided

Why this confidence tier?

High confidence

Confidence is rule-based. Every factor below contributed to the tier. We never expose a single numeric score; the tier label is auditable through this explanation alone.

Factors
  • +Measured by RunLocalAI editorial
How to improve this benchmark's confidence

Cohort intelligence

How this measurement compares to the rest of the corpus. Only comparable rows (same model + hardware first, with relaxations labelled) are used. We never average across runtimes or quant formats unless explicitly told to.

Insufficient comparison data. Insufficient cohort (0 comparable measurements). Outlier detection requires ≥5.

Same model + hardware, different runtime

2 matching rows

Variance here is pure runtime / version drift. Wide spread suggests a runtime regression candidate worth investigating.

Median tok/s
80.5
Spread
79.1 82.0
  • 79.1 tok/srtx-5080ollama-local-apiunknownEditorial
  • 82.0 tok/srtx-5080ollama version is 0.24.0Q4_K_MEditorial

Same model + hardware, different quant

1 matching row

How much speed you trade for memory. Useful when the operator is deciding whether to drop to a smaller quant.

Median tok/s
79.1
Spread
79.1 79.1
  • 79.1 tok/srtx-5080ollama-local-apiunknownEditorial

Same model, different hardware

1 matching row

What this model looks like on adjacent hardware. Drives the 'should I upgrade?' question.

Median tok/s
43.4
Spread
43.4 43.4
  • 43.4 tok/srtx-3080-16gb-mobileollama version is 0.24.0Q4_K_MEditorial

Same hardware, different model

8 matching rows

What else this rig can run at the same quant bucket.

Median tok/s
134.6
Spread
79.0 443.7
CoV
65%
  • 79.0 tok/srtx-5080ollama version is 0.24.0Q4_K_MEditorial
  • 101.1 tok/srtx-5080ollama-local-apiQ4_K_MEditorial
  • 133.4 tok/srtx-5080ollama-local-apiQ4_K_MEditorial
  • 133.6 tok/srtx-5080ollama-local-apiQ4_K_MEditorial
  • 135.6 tok/srtx-5080ollama version is 0.24.0Q4_K_MEditorial
  • +3 more

Reproduce this benchmark

Got the same model + hardware combo? Run the same measurement and submit your numbers. We'll pre-fill model, hardware, quant, and context — you just add your tok/s, VRAM, runtime version. If your numbers match within ±15%, this benchmark gets a confidence lift and a reproduction badge.

Cite or export

Reference this benchmark in your work. Multiple formats; CC-BY attribution required.

Cite this benchmark or paste it into a README. Copy-to-clipboard; license is CC-BY-4.0 (attribution to RunLocalAI required).

Embed this benchmark
Paste into a Reddit thread, blog post, or README — attribution baked in.
<a href="https://www.runlocalai.co/benchmarks/345" rel="noopener">RunLocalAI: Trendyol LLM Asure 12B on NVIDIA GeForce RTX 5080 — 61.5 tok/s</a>

Direct download: .json · .md · .bib · .svg

Next recommended step

Got the same model + hardware? Run it and submit your numbers — successful reproductions lift this benchmark's confidence tier.