New Same benchmark, more providers: Ollama vs OpenCode Zen & Go. Compare on TokenDyno →

Ollama Cloud tokens per second — live benchmark

Real inference speed, measured continuously. Every row is a live Ollama Cloud model — sorted by tokens per second, benchmarked every ~10 minutes.

● live — last benchmark 13s ago
Trend
DeepSeek V4 Flash Pro 191.8 258.1 538ms 99% 40.3 9m ago
Nemotron 3 Nano 30B (non-reasoning) Pro 148.4 99.9 445ms 99% 7.4 10m ago
Nemotron 3 Nano 30B (non-reasoning) Free 147.8 112.8 445ms 96% 7.4 1m ago
DeepSeek V4 Flash 0731 Pro 147.0 135.1 5.2s 100% 49.9 13s ago
GPT-OSS 20B Free 118.8 63.0 681ms 100% 14.9 2m ago
GPT-OSS 20B Pro 116.4 116.8 445ms 100% 14.9 11m ago
DeepSeek V4 Pro Pro 116.3 124.4 672ms 100% 44.3 9m ago
GLM 5.2 Pro 114.5 60.0 703ms 100% 51.1 1m ago
Kimi K2.7 Code Pro 109.1 80.1 909ms 100% 41.9 10m ago
GPT-OSS 120B Free 105.5 94.3 484ms 92% 23.8 2m ago
Nemotron 3 Super Free 99.0 69.3 522ms 96% 25.4 59m ago
GPT-OSS 120B Pro 96.8 69.7 468ms 87% 23.8 11m ago
Nemotron 3 Super Pro 94.5 88.3 525ms 100% 25.4 10m ago
Kimi K2.6 Pro 93.2 74.1 1.1s 99% 44.2 11m ago
Gemma4 31B Pro 88.6 112.0 332ms 99% 29.4 9m ago
Gemma4 31B Free 88.2 86.3 317ms 92% 29.4 8m ago
GLM 5.1 Pro 88.2 97.5 933ms 100% 40.2 1m ago
MiniMax M3 Free 79.0 82.0 823ms 100% 44.4 2m ago
MiniMax M3 Pro 77.7 83.0 750ms 100% 44.4 10m ago
Kimi K3 Pro 69.8 88.9 1.1s 100% n=6 57.1 3h ago
Qwen3.5 397B Pro 58.2 62.9 1.0s 100% 33.7 9m ago
Mistral Large 3 675B (non-reasoning) Pro 56.9 46.5 1.6s 100% 15.9 10m ago
MiniMax M2.7 Pro 47.8 49.3 1.2s 100% 38.1 10m ago
Nemotron 3 Ultra Pro 39.4 20.1 705ms 100% 37.8 10m ago
Nemotron 3 Ultra Free 38.1 42.1 572ms 96% 37.8 57m ago

Intelligence Index scores from Artificial Analysis.

Ollama Free is sampled about hourly to avoid burning through the weekly free-tier balance.

Frequently asked questions

What is the fastest Ollama Cloud model right now?

As of the last build, DeepSeek V4 Flash is the fastest Ollama Cloud model on this leaderboard at 258.1 tokens/sec (most recent measured run). Rankings change as models are re-benchmarked roughly every 10 minutes (every 60 minutes on the Ollama Free tier) — see the live leaderboard above for the current order.

How is tokens per second measured?

TPS is generation throughput: output tokens divided by server-reported generation time, excluding time-to-first-token. For Ollama Cloud we use the server’s own reported timing (eval_count and total_duration) rather than a client-side stopwatch, so results are immune to network jitter and token buffering. Full formula and error handling are on the methodology page.

How often is data refreshed?

The worker benchmarks continuously using a round-robin priority queue. Ollama Cloud Pro is sampled about every 10 minutes; Ollama Free is sampled about every 60 minutes to preserve the weekly free-tier quota. Models Ollama bills as extra usage rather than plan usage — currently Kimi K3 — are sampled about every 4 hours, since each run is charged per token against a prepaid balance. The leaderboard also polls the API roughly every 60 seconds client-side to keep cells current without a full page reload.