Ollama Cloud models
Live speed benchmarks for every model available on Ollama Cloud. Numbers update every ~10 minutes (Ollama Free hourly) — sorted by latest tokens per second.
All Ollama Cloud models by speed
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 Nemotron 3 Nano 30B (non-reasoning) | 346.0 | 186.1 | 303ms | 100% | 6.8 |
| 2 GPT-OSS 120B | 306.5 | 201.0 | 338ms | 100% | 11.6 |
| 3 Gemma4 31B | 288.9 | 248.8 | 342ms | 100% | 19 |
| 4 Gemma4 31B | 270.3 | 200.0 | 343ms | 100% | 19 |
| 5 GPT-OSS 120B | 252.3 | 252.5 | 363ms | 100% | 11.6 |
| 6 DeepSeek V4.1 Flash | 173.1 | 169.6 | 893ms | 100% | 39.5 |
| 7 GPT-OSS 20B | 164.9 | 95.8 | 523ms | 100% | 9 |
| 8 DeepSeek V4 Pro 0813 | 148.5 | 117.4 | 518ms | 100% | 36 |
| 9 DeepSeek V4 Flash 0731 | 141.8 | 131.0 | 1.2s | 100% | 34.3 |
| 10 Nemotron 3 Nano 30B (non-reasoning) | 113.3 | 192.1 | 633ms | 100% | 6.8 |
| 11 Kimi K2.7 Code | 112.1 | 95.8 | 667ms | 100% | 25.8 |
| 12 GLM 5.3 Flash | 100.8 | 101.8 | 637ms | 100% | 41.8 |
| 13 GLM 5.3 | 98.9 | 115.2 | 746ms | 100% | 44.8 |
| 14 Qwen3.5 397B | 98.0 | 92.4 | 935ms | 100% | 18.4 |
| 15 GPT-OSS 20B | 95.1 | 91.5 | 429ms | 100% | 9 |
| 16 Kimi K3 | 93.8 | 69.7 | 891ms | 100% | 43.6 |
| 17 Nemotron 3 Super | 92.9 | 96.9 | 842ms | 100% | 12.8 |
| 18 Nemotron 3 Ultra | 91.5 | 41.2 | 586ms | 99% | 22.9 |
| 19 Nemotron 3 Ultra | 89.3 | 37.0 | 741ms | 92% | 22.9 |
| 20 MiniMax M3 | 89.2 | 70.7 | 705ms | 100% | 29.2 |
| 21 GLM 5.2 | 89.1 | 79.4 | 1.6s | 100% | 33.7 |
| 22 MiniMax M3 | 82.2 | — | 1.0s | 0% | 29.2 |
| 23 Mistral Large 3 675B (non-reasoning) | 76.6 | 70.0 | 1.2s | 100% | 9.3 |
| 24 Nemotron 3 Super | 69.0 | 81.3 | 507ms | 99% | 12.8 |
| 25 MiniMax M2.7 | 54.3 | 60.9 | 797ms | 99% | 22.8 |
| 26 GLM 5.1 | 49.7 | 47.8 | 867ms | 100% | 26.1 |
| 27 Kimi K2.6 | 36.1 | 38.4 | 18.5s | 100% | 27 |
About Ollama Cloud models
Ollama Cloud provides access to large language models via a hosted API. Models range from lightweight coding assistants (e.g. gpt-oss:20b) to flagship reasoning models (e.g. deepseek-v4-pro), each with its own per-million-token rate on the pricing page. Local model usage is always unlimited; cloud models spend your plan's monthly usage credits.
The benchmarks on this page come from continuous automated tests — one streaming chat completion per model per ~10 minutes, measured from outside Ollama's network. See the methodology page for the full measurement spec.
Compare models side by side
Use the compare tool to overlay speed timelines for up to 6 models. Useful for picking between similar-sized models or tracking performance changes over time.