Ollama vs LM Studio on an M4 Pro: the same Gemma 4 file runs at the same speed
Same Mac, same model file, same 600 token limit: 11.74 tokens a second in LM Studio and 11.50 in Ollama. The difference is in what each lets you change.
Ollama vs LM Studio speed on an M4 Pro: we ran the same Gemma 4 31B file through Ollama and through LM Studio on a Mac with an M4 Pro chip and 48 GB of memory. The writing speed was within 2%: 11.74 tokens a second in LM Studio and 11.50 in Ollama. What separates them is not speed. It is what each lets you change.
Ollama vs LM Studio speed: is Ollama faster on an M4 Pro?
No, not by a margin you would notice. Each ran 15 timed runs across five tasks with the same 600 token limit, the same machine and the same model file, so the weights were identical.
| Measure | LM Studio | Ollama |
|---|---|---|
| Writing speed, middle of 15 runs | 11.74 tokens a second | 11.50 tokens a second |
| Slowest and fastest run | 11.40 and 12.04 | 11.31 and 11.62 |
| Wait for the first token, 3,800 token prompt | 46.1 seconds | 45.3 seconds |
| Same prompt sent again | 0.11 and 0.32 seconds | 0.11 and 0.13 seconds |
| Runs with no visible answer at 600 tokens | 6 of 15 | 7 of 15 |
Does Ollama use more memory than LM Studio?
It may, but our test cannot prove it. Ollama’s own listing showed 19 GB for the loaded model at both 32,768 and 65,536 tokens of context, all on the graphics chip. The Ollama process measured by the operating system was 24.9 GB after its first run at 65,536 and 33.2 GB after 15 runs. LM Studio’s process measured 20.6 to 22.6 GB, but we sampled it differently, once during a run and once when idle, so the two numbers are not a fair match. The detail is in our test of Ollama’s context and memory.
Which handles a repeated prompt better?
They behave the same. Both read a 3,800 token document in about 45 to 46 seconds the first time, and both answered in a fraction of a second when sent the same document again, because each keeps the prompt it has just read. Change the first line and both read the whole document again. This is the trap described in our guide to measuring tokens per second.
Do both give empty replies at a 600 token limit?
Yes. Gemma 4 thinks before it writes, and at 600 tokens LM Studio returned nothing in 6 runs of 15 and Ollama in 7 of 15. The cause is the same in both, explained in our study of where its tokens go.
Which should I use for Gemma 4 31B on a Mac?
On speed it is a tie, so choose on what you need to change. Ollama let us turn thinking off with one setting, and that removed every empty reply and cut the typical wait roughly in half, as set out in our guide to turning off thinking in Ollama. We did not test whether LM Studio has an equivalent. If you already have the model file, importing it into Ollama takes two commands.
How we tested Ollama against LM Studio
One Mac with an M4 Pro chip and 48 GB of memory, macOS 27.0.1. LM Studio with its llama.cpp Metal runtime 2.47.0, and Ollama 0.32.6, both loading the same 17.65 GB Q4_0 file. Five tasks, three runs each, 600 tokens allowed, temperature 0, and a different first line on every prompt. Ollama was run at a 65,536 token context to match LM Studio. Speed is each program’s own report. The method is on how we test.
Sources
Cite this page
"Ollama vs LM Studio on an M4 Pro: the same Gemma 4 file runs at the same speed". Onticpost, published 4 October 2026. https://onticpost.com/ollama-vs-lm-studio-speed-m4-pro-gemma-4/
Updated 4 October 2026. This address stays the same when the piece is updated.