Search
Reviews

Gemma 4 31B QAT on an M4 Pro with 48 GB: 11.7 tokens a second, after a long think

It fits a 48 GB Mac with room to spare and writes at a steady pace. It also spends most of its tokens thinking before the reader sees a word.

Written byadmin
Read time3 min
Published3 October 2026
Filed underReviews

Verdict. Fits a 48 GB Mac with room to spare and holds a steady 11.7 tokens a second. It thinks at length before it writes, so a short answer takes about a minute and a tight token limit returns nothing.

3 of 5

Gemma 4 31B is the largest dense model in Google’s Gemma 4 family. We ran the QAT Q4_0 build in LM Studio on a Mac with an M4 Pro chip and 48 GB of memory, and timed 15 runs across five kinds of task. It wrote at a steady 11.7 tokens a second, used about 21 to 23 GB, and spent most of its tokens thinking before it wrote anything a reader could see.

How fast is Gemma 4 31B on an M4 Pro?

The middle result of 15 runs was 11.74 tokens a second. The slowest run was 11.40 and the fastest 12.04, so the speed barely moves from task to task. At that rate a 200 word answer is about 20 seconds of writing, once the model starts to write.

Measure What we found
Writing speed, middle of 15 runs 11.74 tokens a second
Slowest and fastest run 11.40 and 12.04 tokens a second
Wait for the first token, short prompt 1.43 seconds
Wait for the first token, 3,806 token prompt 46.1 seconds
Memory used by the model process 20.6 to 22.6 GB

How much memory does Gemma 4 31B QAT use on a Mac?

The model file is 17.65 GB on disk. Loaded with a context window of 65,536 tokens, the process that runs it held between 20.6 and 22.6 GB during the test. On a 48 GB Mac that leaves about 25 GB for everything else. Google’s model card gives no speed or memory figure for any machine, so there is no maker’s claim to check these against.

How long does Gemma 4 31B take to read a long prompt?

Reading is the slow part. Given a 3,806 token document, the model took 46.1 seconds before its first token appeared, which is about 82 tokens a second of reading. Sent the same document a second time, it answered in a third of a second, because LM Studio keeps the prompt it has already read. That shortcut also distorts speed tests, which is why we wrote a guide to measuring tokens per second properly.

Why does Gemma 4 take so long to answer a simple question?

It thinks first. With no setting changed, the model spent between 408 and 597 of every 600 tokens on reasoning that the reader never sees. In 6 of the 15 runs it used the whole allowance thinking and returned an empty answer. Given room, a 217 word explanation took 91 seconds from start to finish, and 777 of its 1,032 tokens were thinking. The full count is in our study of where its tokens go.

Is Gemma 4 31B worth running on a 48 GB Mac?

Yes, if the work can wait a minute. It fits with room to spare, its speed is steady, and the one answer we could mark, a count over a long document, was right every time. It is the wrong choice for quick back and forth, and anything that sets a tight token limit will get nothing back.

How we tested Gemma 4 31B

One machine: a Mac with an M4 Pro chip and 48 GB of memory, on macOS 27.0.1. LM Studio’s llama.cpp Metal runtime, version 2.47.0, with the model loaded at a 65,536 token context. Five tasks (an explanation, a piece of code, a summary, a list and a worked sum), three runs each, 600 tokens allowed per run. Every prompt began with a fresh label so that no run could reuse a prompt the model had already read. Speed and first token times are the figures LM Studio’s own server reports. Memory is the size of the model process as the operating system reported it. The measurement is recorded in the index.