Ollama Cloud models
Live speed benchmarks for every model available on Ollama Cloud. Numbers update every ~10 minutes (Ollama Free hourly) — sorted by latest tokens per second.
All Ollama Cloud models by speed
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 DeepSeek V4 Flash | 258.1 | 191.8 | 538ms | 99% | 40.3 |
| 2 DeepSeek V4 Flash 0731 | 135.1 | 147.0 | 5.2s | 100% | 49.9 |
| 3 DeepSeek V4 Pro | 124.4 | 116.3 | 672ms | 100% | 44.3 |
| 4 GPT-OSS 20B | 116.8 | 116.4 | 445ms | 100% | 14.9 |
| 5 Nemotron 3 Nano 30B (non-reasoning) | 112.8 | 147.8 | 445ms | 96% | 7.4 |
| 6 Gemma4 31B | 112.0 | 88.6 | 332ms | 99% | 29.4 |
| 7 Nemotron 3 Nano 30B (non-reasoning) | 99.9 | 148.4 | 445ms | 99% | 7.4 |
| 8 GLM 5.1 | 97.5 | 88.2 | 933ms | 100% | 40.2 |
| 9 GPT-OSS 120B | 94.3 | 105.5 | 484ms | 92% | 23.8 |
| 10 Kimi K3 | 88.9 | 69.8 | 1.1s | 100% | 57.1 |
| 11 Nemotron 3 Super | 88.3 | 94.5 | 525ms | 100% | 25.4 |
| 12 Gemma4 31B | 86.3 | 88.2 | 317ms | 92% | 29.4 |
| 13 MiniMax M3 | 83.0 | 77.7 | 750ms | 100% | 44.4 |
| 14 MiniMax M3 | 82.0 | 79.0 | 823ms | 100% | 44.4 |
| 15 Kimi K2.7 Code | 80.1 | 109.1 | 909ms | 100% | 41.9 |
| 16 Kimi K2.6 | 74.1 | 93.2 | 1.1s | 99% | 44.2 |
| 17 GPT-OSS 120B | 69.7 | 96.8 | 468ms | 87% | 23.8 |
| 18 Nemotron 3 Super | 69.3 | 99.0 | 522ms | 96% | 25.4 |
| 19 GPT-OSS 20B | 63.0 | 118.8 | 681ms | 100% | 14.9 |
| 20 Qwen3.5 397B | 62.9 | 58.2 | 1.0s | 100% | 33.7 |
| 21 GLM 5.2 | 60.0 | 114.5 | 703ms | 100% | 51.1 |
| 22 MiniMax M2.7 | 49.3 | 47.8 | 1.2s | 100% | 38.1 |
| 23 Mistral Large 3 675B (non-reasoning) | 46.5 | 56.9 | 1.6s | 100% | 15.9 |
| 24 Nemotron 3 Ultra | 42.1 | 38.1 | 572ms | 96% | 37.8 |
| 25 Nemotron 3 Ultra | 20.1 | 39.4 | 705ms | 100% | 37.8 |
About Ollama Cloud models
Ollama Cloud provides access to large language models via a hosted API. Models range from lightweight coding assistants (e.g. gpt-oss:20b, usage level 1) to flagship reasoning models (e.g. deepseek-v4-pro, usage level 4). Local model usage is always unlimited; cloud models count toward your plan's usage balance.
The benchmarks on this page come from continuous automated tests — one streaming chat completion per model per ~10 minutes, measured from outside Ollama's network. See the methodology page for the full measurement spec.
Compare models side by side
Use the compare tool to overlay speed timelines for up to 6 models. Useful for picking between similar-sized models or tracking performance changes over time.