Gemma 4 with thinking off: 15 runs, no empty replies and answers in half the time
With thinking off in Ollama, every run answered and the typical reply took 27.2 seconds instead of 53.8. The writing speed did not change.
We ran Gemma 4 31B in Ollama 15 times with thinking turned off, using the same tasks and the same 600 token limit that left LM Studio with 6 empty replies in 15. Every run answered. The typical reply took 27.2 seconds, against 53.8 seconds with thinking on, and the writing speed stayed at about 11.6 tokens a second.
Does turning off thinking stop empty replies in Gemma 4?
Yes, in our test. With thinking on, 7 of 15 runs returned no visible answer, because the model spent the whole 600 tokens thinking. With it off, 0 of 15 did.
How much faster is Gemma 4 with thinking off?
Roughly twice as fast to a finished answer, and the real gap is wider than it looks. With thinking on, most runs stopped at the 600 token limit, so 53.8 seconds is the cap and not the time the model needed. With thinking off, four of the five tasks finished on their own.
| Task | Time with thinking off | Tokens written | Finished on its own |
|---|---|---|---|
| Explain, about 250 words | 27.2 seconds | 295 to 302 | 3 of 3 |
| Summarise, about 200 words | 20.3 seconds | 217 to 225 | 3 of 3 |
| List ten checks | 18.9 seconds | 197 to 218 | 3 of 3 |
| Work out a train time | 36.8 seconds | 410 | 3 of 3 |
| Write a Python function | 53.6 seconds | 600 | 0 of 3 |
The code task ran to the limit even with thinking off, because the answer itself was longer than 600 tokens. Give code more room.
Does turning off thinking change tokens per second?
No. The middle writing speed was 11.61 tokens a second with thinking off and 11.50 with it on, from 15 runs each. Turning thinking off removes the wait, not the cost of each token.
Is the answer still right with thinking off?
The one we could check was. The train question gave 13:05 in all three runs, which is correct. Nothing else was graded, so we are not claiming the other answers are as good as the thinking ones. The explanations ran 234 to 238 words against the 250 asked for.
What did we not test?
We did not grade answer quality beyond the one sum and the document count. We did not repeat the thinking off runs in LM Studio, or try the middle settings some models offer between on and off. We ran one machine and one build of one model.
To set this yourself, see how to turn off thinking in Ollama. The behaviour with thinking on is in our count of where its tokens go.
How we tested
Ollama 0.32.6 on a Mac with an M4 Pro chip and 48 GB of memory, macOS 27.0.1. Gemma 4 31B QAT Q4_0, a 65,536 token context, 600 tokens allowed, temperature 0, the setting "think": false on every request, three runs of five tasks, and a different first line on every prompt. Times are the whole reply from request to finish. The method is on how we test.