How to turn off thinking in Ollama for Gemma 4, and what it changed in our test
One setting stopped the empty replies and cut the typical wait roughly in half. Two ways to set it, how to check your model supports it, and the numbers from 15 runs.
Gemma 4 thinks before it answers, and in Ollama that thinking counts against your token limit. To turn it off, set think to false. In our 15 runs with it off, every one answered, and a 250 word explanation took about 27 seconds instead of running out of tokens.
How do I turn off thinking in Ollama?
There are two routes we checked. On the command line, ollama run has a --think flag that takes true or false, so ollama run your-model --think=false starts a chat with thinking off. Through the local server, add "think": false to the body of a request to /api/chat. That is the route we used for every timed run.
How do I know if my Ollama model can think?
Run ollama show your-model and read the Capabilities list. Our Gemma 4 31B import listed completion, tools and thinking. If thinking is not listed for your model, the setting has nothing to turn off.
What does turning off thinking change in Ollama?
The wait, and whether you get an answer. The writing speed did not move.
| Measure | Thinking on | Thinking off |
|---|---|---|
| Runs with no visible answer, 600 tokens allowed | 7 of 15 | 0 of 15 |
| Middle time for a whole reply | 53.8 seconds | 27.2 seconds |
| Writing speed, middle of 15 runs | 11.50 tokens a second | 11.61 tokens a second |
| Finished the 250 word explanation inside the limit | 0 of 3 | 3 of 3 |
With thinking on, most runs stopped at the 600 token limit, so 53.8 seconds is the cap and not the real time. The full numbers are in our timed test with thinking off.
Is the answer still right with thinking off?
The one answer we could mark was. Asked when a train arrives after 180 km at 90 km/h, a 25 minute wait and 120 km at 80 km/h, it gave 13:05 in three runs of three, which is correct. We did not grade the explanations, the summary or the code, so this is not a quality test.
Does --hidethinking do the same thing?
Not necessarily. ollama run also has a --hidethinking flag, which its help describes as hiding the thinking output. We did not test whether it saves any time, so do not assume it does.
How we tested
Ollama 0.32.6 on a Mac with an M4 Pro chip and 48 GB of memory, the Gemma 4 31B QAT Q4_0 file imported from a local file, a 65,536 token context, 600 tokens allowed, temperature 0, three runs of each of five tasks. The method is on how we test. For the same model in LM Studio, see Ollama against LM Studio.
Cite this page
"How to turn off thinking in Ollama for Gemma 4, and what it changed in our test". onticpost.com, published 4 October 2026. https://onticpost.com/turn-off-thinking-ollama-gemma-4/
Updated 4 October 2026. This address stays the same when the piece is updated.