# How to turn off thinking in Ollama for Gemma 4, and what it changed in our test URL: https://onticpost.com/turn-off-thinking-ollama-gemma-4/ Published: 2026-10-04 Updated: 2026-10-04 By: Onticpost > One setting stopped the empty replies and cut the typical wait roughly in half. Two ways to set it, how to check your model supports it, and the numbers... Gemma 4 thinks before it answers, and in Ollama that thinking counts against your token limit. To turn it off, set `think` to false. In our 15 runs with it off, every one answered, and a 250 word explanation took about 27 seconds instead of running out of tokens. ## How do I turn off thinking in Ollama for Gemma 4? There are two routes we checked. On the command line, `ollama run` has a `--think` flag that takes true or false, so `ollama run your-model --think=false` starts a chat with thinking off. Through the local server, add `"think": false` to the body of a request to `/api/chat`. That is the route we used for every timed run. ## How do I know if my Ollama model can think? Run `ollama show your-model` and read the Capabilities list. Our Gemma 4 31B import listed completion, tools and thinking. If thinking is not listed for your model, the setting has nothing to turn off. ## What does turning off thinking change in Ollama? The wait, and whether you get an answer. The writing speed did not move. | Measure | Thinking on | Thinking off | | --- | --- | --- | | Runs with no visible answer, 600 tokens allowed | 7 of 15 | 0 of 15 | | Middle time for a whole reply | 53.8 seconds | 27.2 seconds | | Writing speed, middle of 15 runs | 11.50 tokens a second | 11.61 tokens a second | | Finished the 250 word explanation inside the limit | 0 of 3 | 3 of 3 | With thinking on, most runs stopped at the 600 token limit, so 53.8 seconds is the cap and not the real time. The full numbers are in [our timed test with thinking off](https://onticpost.com/gemma-4-thinking-off-empty-replies-timed/). ## Is the answer still right with thinking off? The one answer we could mark was. Asked when a train arrives after 180 km at 90 km/h, a 25 minute wait and 120 km at 80 km/h, it gave 13:05 in three runs of three, which is correct. We did not grade the explanations, the summary or the code, so this is not a quality test. ## Does `--hidethinking` do the same thing? Not necessarily. `ollama run` also has a `--hidethinking` flag, which its help describes as hiding the thinking output. We did not test whether it saves any time, so do not assume it does. ## How we tested Ollama 0.32.6 on a Mac with an M4 Pro chip and 48 GB of memory, the Gemma 4 31B QAT Q4_0 file [imported from a local file](https://onticpost.com/import-gguf-ollama-modelfile-gemma-4/), a 65,536 token context, 600 tokens allowed, temperature 0, three runs of each of five tasks. The method is on [how we test](https://onticpost.com/how-we-test/). For the same model in LM Studio, see [Ollama against LM Studio](https://onticpost.com/ollama-vs-lm-studio-speed-m4-pro-gemma-4/). ## Sources - [Google's Gemma 4 31B QAT model card](https://huggingface.co/google/gemma-4-31b-it-qat-q4_0-gguf)