Search
Research

Why Gemma 4 gives an empty response: we counted where its tokens go

In 18 timed runs the model spent most of its tokens on thinking nobody sees. With 600 tokens allowed, it ran out before answering in 6 runs of 15.

Written byadmin
Read time2 min
Published3 October 2026
Filed underResearch

Give Gemma 4 a tight token limit and the reply can come back empty. We counted the tokens in 18 runs of the 31B model to see why. In every run the model spent most of its tokens on thinking the reader never sees, and when the allowance was 600 tokens it ran out before answering in 6 runs of 15.

Why is my Gemma 4 response empty?

Because the token limit ran out while the model was still thinking. Gemma 4 reasons privately before it writes, and that reasoning counts against the same limit as the answer. LM Studio reports it separately as reasoning tokens. If the limit is reached first, the visible reply is blank even though the model worked for the full time.

How many tokens does Gemma 4 spend thinking?

With 600 tokens allowed, thinking took between 408 and 597 of them in every run. The explanation task never produced a visible word.

Task Thinking tokens in three runs Runs with no answer
Explain how a heat pump works 597, 597, 597 3 of 3
Write a short Python function 597, 597, 484 2 of 3
Summarise a trade off 448, 449, 597 1 of 3
List ten checks 423, 423, 408 0 of 3
Work out a train time 469, 486, 486 0 of 3

Even the runs that did answer were cut short. All 15 stopped at the 600 token limit with the answer unfinished.

How many tokens should I allow for Gemma 4 31B?

We ran three of the tasks again with 3,000 tokens allowed, and let the model stop on its own.

Task Tokens used Of which thinking Time taken
Explain, about 250 words 1,032 777 91 seconds
Python function 1,354 610 119 seconds
Summary, about 200 words 711 465 63 seconds

For a reply of a few hundred words, allow at least 1,500 tokens. The model asked for between 465 and 777 tokens of thinking before answers that were themselves only 250 to 750 tokens long.

Does the thinking make Gemma 4 slower?

It does not change the rate, which held at about 11.7 tokens a second on our machine. It changes the wait. At that rate, 600 tokens of thinking is 50 seconds of silence before the first visible word.

Does the thinking help?

On the one task where we could mark the answer, yes. Asked how many pump stations a 3,806 token document mentioned, the model answered 120, which is correct, in four runs of four. It used 168 of its 176 tokens to get there.

What we did not test

We changed no setting, so we cannot say how the model behaves with its thinking turned off or limited. We tested one build, the QAT Q4_0 file described on Google’s model card, on one machine. The speed and memory figures are in our review, and the method for timing a model is in our guide.