Why Gemma 4 gives an empty response: we counted where its tokens go
In 18 timed runs the model spent most of its tokens on thinking nobody sees. With 600 tokens allowed, it ran out before answering in 6 runs of 15.
Give Gemma 4 a tight token limit and the reply can come back empty. We counted the tokens in 18 runs of the 31B model to see why. In every run the model spent most of its tokens on thinking the reader never sees, and when the allowance was 600 tokens it ran out before answering in 6 runs of 15.
Why is my Gemma 4 response empty?
Because the token limit ran out while the model was still thinking. Gemma 4 reasons privately before it writes, and that reasoning counts against the same limit as the answer. LM Studio reports it separately as reasoning tokens. If the limit is reached first, the visible reply is blank even though the model worked for the full time.
How many tokens does Gemma 4 spend thinking?
With 600 tokens allowed, thinking took between 408 and 597 of them in every run. The explanation task never produced a visible word.
| Task | Thinking tokens in three runs | Runs with no answer |
|---|---|---|
| Explain how a heat pump works | 597, 597, 597 | 3 of 3 |
| Write a short Python function | 597, 597, 484 | 2 of 3 |
| Summarise a trade off | 448, 449, 597 | 1 of 3 |
| List ten checks | 423, 423, 408 | 0 of 3 |
| Work out a train time | 469, 486, 486 | 0 of 3 |
Even the runs that did answer were cut short. All 15 stopped at the 600 token limit with the answer unfinished.
How many tokens should I allow for Gemma 4 31B?
We ran three of the tasks again with 3,000 tokens allowed, and let the model stop on its own.
| Task | Tokens used | Of which thinking | Time taken |
|---|---|---|---|
| Explain, about 250 words | 1,032 | 777 | 91 seconds |
| Python function | 1,354 | 610 | 119 seconds |
| Summary, about 200 words | 711 | 465 | 63 seconds |
For a reply of a few hundred words, allow at least 1,500 tokens. The model asked for between 465 and 777 tokens of thinking before answers that were themselves only 250 to 750 tokens long.
Does the thinking make Gemma 4 slower?
It does not change the rate, which held at about 11.7 tokens a second on our machine. It changes the wait. At that rate, 600 tokens of thinking is 50 seconds of silence before the first visible word.
Does the thinking help?
On the one task where we could mark the answer, yes. Asked how many pump stations a 3,806 token document mentioned, the model answered 120, which is correct, in four runs of four. It used 168 of its 176 tokens to get there.
What we did not test
We changed no setting, so we cannot say how the model behaves with its thinking turned off or limited. We tested one build, the QAT Q4_0 file described on Google’s model card, on one machine. The speed and memory figures are in our review, and the method for timing a model is in our guide.