Head-to-head The same models run on Ollama and OpenCode — who streams faster? See the race on TokenDyno →

Ollama Cloud models

Live speed benchmarks for every model available on Ollama Cloud. Numbers update every ~10 minutes (Ollama Free hourly) — sorted by latest tokens per second.

26 total models benchmarked
26 available on current plan
live data updates every ~10 min

All Ollama Cloud models by speed

Model TPS now TPS 24h avg TTFT Reliability Intelligence Index
1 GPT-OSS 120B 367.9 277.2 353ms 100% 12.3
2 Gemma4 31B 362.0 222.3 332ms 100% 15.4
3 Gemma4 31B 296.8 200.9 670ms 100% 15.4
4 MiniMax M3 282.7 169.7 436ms 100% 29.6
5 DeepSeek V4 Flash 0731 215.0 172.7 451ms 100% 34.5
6 GPT-OSS 120B 158.1 225.5 954ms 100% 12.3
7 DeepSeek V4 Pro 0813 141.1 139.8 484ms 99% 36.3
8 Nemotron 3 Nano 30B (non-reasoning) 118.4 186.6 915ms 100% 6.8
9 Kimi K3 108.5 102.6 843ms 100% 43.8
10 GPT-OSS 20B 106.1 93.2 492ms 100% 9
11 Qwen3.5 397B 104.7 87.6 867ms 99% 19.1
12 Kimi K2.7 Code 104.6 127.8 744ms 100% 26.3
13 Nemotron 3 Super 103.8 86.9 592ms 99% 13.6
14 Nemotron 3 Nano 30B (non-reasoning) 99.8 227.1 552ms 100% 6.8
15 Nemotron 3 Ultra 99.8 51.0 2.3s 88% 23.4
16 GLM 5.3 93.9 115.2 1.2s 100% 44.9
17 GLM 5.1 85.6 63.2 1.4s 99% 27.4
18 GPT-OSS 20B 83.5 94.4 592ms 100% 9
19 MiniMax M3 82.2 1.0s 0% 29.6
20 Nemotron 3 Ultra 71.1 56.3 12.9s 92% 23.4
21 Mistral Large 3 675B (non-reasoning) 69.2 68.4 579ms 100% 9.7
22 Nemotron 3 Super 68.3 89.8 563ms 100% 13.6
23 GLM 5.3 Flash 64.8 138.5 706ms 99% 41.9
24 MiniMax M2.7 30.8 55.0 5.9s 100% 23.2
25 Kimi K2.6 30.1 29.1 774ms 100% 31.3
26 GLM 5.2 27.1 79.1 712ms 100% 38.6

About Ollama Cloud models

Ollama Cloud provides access to large language models via a hosted API. Models range from lightweight coding assistants (e.g. gpt-oss:20b) to flagship reasoning models (e.g. deepseek-v4-pro), each with its own per-million-token rate on the pricing page. Local model usage is always unlimited; cloud models spend your plan's monthly usage credits.

The benchmarks on this page come from continuous automated tests — one streaming chat completion per model per ~10 minutes, measured from outside Ollama's network. See the methodology page for the full measurement spec.

Compare models side by side

Use the compare tool to overlay speed timelines for up to 6 models. Useful for picking between similar-sized models or tracking performance changes over time.