Ollama Cloud models
Live speed benchmarks for every model available on Ollama Cloud. Numbers update every ~10 minutes (Ollama Free hourly) — sorted by latest tokens per second.
All Ollama Cloud models by speed
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 GPT-OSS 120B | 367.9 | 277.2 | 353ms | 100% | 12.3 |
| 2 Gemma4 31B | 362.0 | 222.3 | 332ms | 100% | 15.4 |
| 3 Gemma4 31B | 296.8 | 200.9 | 670ms | 100% | 15.4 |
| 4 MiniMax M3 | 282.7 | 169.7 | 436ms | 100% | 29.6 |
| 5 DeepSeek V4 Flash 0731 | 215.0 | 172.7 | 451ms | 100% | 34.5 |
| 6 GPT-OSS 120B | 158.1 | 225.5 | 954ms | 100% | 12.3 |
| 7 DeepSeek V4 Pro 0813 | 141.1 | 139.8 | 484ms | 99% | 36.3 |
| 8 Nemotron 3 Nano 30B (non-reasoning) | 118.4 | 186.6 | 915ms | 100% | 6.8 |
| 9 Kimi K3 | 108.5 | 102.6 | 843ms | 100% | 43.8 |
| 10 GPT-OSS 20B | 106.1 | 93.2 | 492ms | 100% | 9 |
| 11 Qwen3.5 397B | 104.7 | 87.6 | 867ms | 99% | 19.1 |
| 12 Kimi K2.7 Code | 104.6 | 127.8 | 744ms | 100% | 26.3 |
| 13 Nemotron 3 Super | 103.8 | 86.9 | 592ms | 99% | 13.6 |
| 14 Nemotron 3 Nano 30B (non-reasoning) | 99.8 | 227.1 | 552ms | 100% | 6.8 |
| 15 Nemotron 3 Ultra | 99.8 | 51.0 | 2.3s | 88% | 23.4 |
| 16 GLM 5.3 | 93.9 | 115.2 | 1.2s | 100% | 44.9 |
| 17 GLM 5.1 | 85.6 | 63.2 | 1.4s | 99% | 27.4 |
| 18 GPT-OSS 20B | 83.5 | 94.4 | 592ms | 100% | 9 |
| 19 MiniMax M3 | 82.2 | — | 1.0s | 0% | 29.6 |
| 20 Nemotron 3 Ultra | 71.1 | 56.3 | 12.9s | 92% | 23.4 |
| 21 Mistral Large 3 675B (non-reasoning) | 69.2 | 68.4 | 579ms | 100% | 9.7 |
| 22 Nemotron 3 Super | 68.3 | 89.8 | 563ms | 100% | 13.6 |
| 23 GLM 5.3 Flash | 64.8 | 138.5 | 706ms | 99% | 41.9 |
| 24 MiniMax M2.7 | 30.8 | 55.0 | 5.9s | 100% | 23.2 |
| 25 Kimi K2.6 | 30.1 | 29.1 | 774ms | 100% | 31.3 |
| 26 GLM 5.2 | 27.1 | 79.1 | 712ms | 100% | 38.6 |
About Ollama Cloud models
Ollama Cloud provides access to large language models via a hosted API. Models range from lightweight coding assistants (e.g. gpt-oss:20b) to flagship reasoning models (e.g. deepseek-v4-pro), each with its own per-million-token rate on the pricing page. Local model usage is always unlimited; cloud models spend your plan's monthly usage credits.
The benchmarks on this page come from continuous automated tests — one streaming chat completion per model per ~10 minutes, measured from outside Ollama's network. See the methodology page for the full measurement spec.
Compare models side by side
Use the compare tool to overlay speed timelines for up to 6 models. Useful for picking between similar-sized models or tracking performance changes over time.