Head-to-head The same models run on Ollama and OpenCode — who streams faster? See the race on TokenDyno →

Ollama Cloud models

Live speed benchmarks for every model available on Ollama Cloud. Numbers update every ~10 minutes (Ollama Free hourly) — sorted by latest tokens per second.

27 total models benchmarked
27 available on current plan
live data updates every ~10 min

All Ollama Cloud models by speed

Model TPS now TPS 24h avg TTFT Reliability Intelligence Index
1 Nemotron 3 Nano 30B (non-reasoning) 346.0 186.1 303ms 100% 6.8
2 GPT-OSS 120B 306.5 201.0 338ms 100% 11.6
3 Gemma4 31B 288.9 248.8 342ms 100% 19
4 Gemma4 31B 270.3 200.0 343ms 100% 19
5 GPT-OSS 120B 252.3 252.5 363ms 100% 11.6
6 DeepSeek V4.1 Flash 173.1 169.6 893ms 100% 39.5
7 GPT-OSS 20B 164.9 95.8 523ms 100% 9
8 DeepSeek V4 Pro 0813 148.5 117.4 518ms 100% 36
9 DeepSeek V4 Flash 0731 141.8 131.0 1.2s 100% 34.3
10 Nemotron 3 Nano 30B (non-reasoning) 113.3 192.1 633ms 100% 6.8
11 Kimi K2.7 Code 112.1 95.8 667ms 100% 25.8
12 GLM 5.3 Flash 100.8 101.8 637ms 100% 41.8
13 GLM 5.3 98.9 115.2 746ms 100% 44.8
14 Qwen3.5 397B 98.0 92.4 935ms 100% 18.4
15 GPT-OSS 20B 95.1 91.5 429ms 100% 9
16 Kimi K3 93.8 69.7 891ms 100% 43.6
17 Nemotron 3 Super 92.9 96.9 842ms 100% 12.8
18 Nemotron 3 Ultra 91.5 41.2 586ms 99% 22.9
19 Nemotron 3 Ultra 89.3 37.0 741ms 92% 22.9
20 MiniMax M3 89.2 70.7 705ms 100% 29.2
21 GLM 5.2 89.1 79.4 1.6s 100% 33.7
22 MiniMax M3 82.2 1.0s 0% 29.2
23 Mistral Large 3 675B (non-reasoning) 76.6 70.0 1.2s 100% 9.3
24 Nemotron 3 Super 69.0 81.3 507ms 99% 12.8
25 MiniMax M2.7 54.3 60.9 797ms 99% 22.8
26 GLM 5.1 49.7 47.8 867ms 100% 26.1
27 Kimi K2.6 36.1 38.4 18.5s 100% 27

About Ollama Cloud models

Ollama Cloud provides access to large language models via a hosted API. Models range from lightweight coding assistants (e.g. gpt-oss:20b) to flagship reasoning models (e.g. deepseek-v4-pro), each with its own per-million-token rate on the pricing page. Local model usage is always unlimited; cloud models spend your plan's monthly usage credits.

The benchmarks on this page come from continuous automated tests — one streaming chat completion per model per ~10 minutes, measured from outside Ollama's network. See the methodology page for the full measurement spec.

Compare models side by side

Use the compare tool to overlay speed timelines for up to 6 models. Useful for picking between similar-sized models or tracking performance changes over time.