Search
Index entry

Gemma 4 31B QAT Q4_0

Fits with about 25 GB to spare. Steady speed, long waits while it thinks.

3 of 5

Measured by us on 3 October 2026

11.7tokens a second
22.6 GBmemory in use
Mac, M4 Pro, 48 GBthe machine

What was measured

Fifteen runs across five tasks in LM Studio, llama.cpp Metal runtime 2.47.0, context 65,536 tokens, 600 tokens allowed per run, on macOS 27.0.1. Writing speed is the middle run. Memory is the largest size the model process reached.

Other figures

Measure Result
Slowest and fastest run 11.40 and 12.04 tokens a second
Wait for the first token, short prompt 1.43 seconds
Wait for the first token, 3,806 token prompt 46.1 seconds
Model file on disk 17.65 GB

The full account is in our review.