New Same benchmark, more providers: Ollama vs OpenCode Zen & Go. Compare on TokenDyno →

Ollama Cloud models

Live speed benchmarks for every model available on Ollama Cloud. Numbers update every ~10 minutes (Ollama Free hourly) — sorted by latest tokens per second.

25 total models benchmarked
25 available on current plan
live data updates every ~10 min

All Ollama Cloud models by speed

Model TPS now TPS 24h avg TTFT Reliability Intelligence Index
1 DeepSeek V4 Flash 258.1 191.8 538ms 99% 40.3
2 DeepSeek V4 Flash 0731 135.1 147.0 5.2s 100% 49.9
3 DeepSeek V4 Pro 124.4 116.3 672ms 100% 44.3
4 GPT-OSS 20B 116.8 116.4 445ms 100% 14.9
5 Nemotron 3 Nano 30B (non-reasoning) 112.8 147.8 445ms 96% 7.4
6 Gemma4 31B 112.0 88.6 332ms 99% 29.4
7 Nemotron 3 Nano 30B (non-reasoning) 99.9 148.4 445ms 99% 7.4
8 GLM 5.1 97.5 88.2 933ms 100% 40.2
9 GPT-OSS 120B 94.3 105.5 484ms 92% 23.8
10 Kimi K3 88.9 69.8 1.1s 100% 57.1
11 Nemotron 3 Super 88.3 94.5 525ms 100% 25.4
12 Gemma4 31B 86.3 88.2 317ms 92% 29.4
13 MiniMax M3 83.0 77.7 750ms 100% 44.4
14 MiniMax M3 82.0 79.0 823ms 100% 44.4
15 Kimi K2.7 Code 80.1 109.1 909ms 100% 41.9
16 Kimi K2.6 74.1 93.2 1.1s 99% 44.2
17 GPT-OSS 120B 69.7 96.8 468ms 87% 23.8
18 Nemotron 3 Super 69.3 99.0 522ms 96% 25.4
19 GPT-OSS 20B 63.0 118.8 681ms 100% 14.9
20 Qwen3.5 397B 62.9 58.2 1.0s 100% 33.7
21 GLM 5.2 60.0 114.5 703ms 100% 51.1
22 MiniMax M2.7 49.3 47.8 1.2s 100% 38.1
23 Mistral Large 3 675B (non-reasoning) 46.5 56.9 1.6s 100% 15.9
24 Nemotron 3 Ultra 42.1 38.1 572ms 96% 37.8
25 Nemotron 3 Ultra 20.1 39.4 705ms 100% 37.8

About Ollama Cloud models

Ollama Cloud provides access to large language models via a hosted API. Models range from lightweight coding assistants (e.g. gpt-oss:20b, usage level 1) to flagship reasoning models (e.g. deepseek-v4-pro, usage level 4). Local model usage is always unlimited; cloud models count toward your plan's usage balance.

The benchmarks on this page come from continuous automated tests — one streaming chat completion per model per ~10 minutes, measured from outside Ollama's network. See the methodology page for the full measurement spec.

Compare models side by side

Use the compare tool to overlay speed timelines for up to 6 models. Useful for picking between similar-sized models or tracking performance changes over time.