Gemma 4
Can a local AI agent draw an SVG illustration? Three tried, none got there
Pi, OpenCode and Hermes each thought of a good picture for an article. Asked to draw it as SVG on Gemma 4 31B, none made one we could publish.
- Research
Gemma 4 with thinking off: 15 runs, no empty replies and answers in half the time
With thinking off in Ollama, every run answered and the typical reply took 27.2 seconds instead of 53.8. The writing speed did not change.
- Research
Ollama’s default context length for Gemma 4 31B on a 48 GB Mac is 32,768, and memory climbs as you use it
Ollama chose 32,768 tokens on its own and listed 19 GB. The process itself grew from 22.0 GB to 33.2 GB over 15 runs at 65,536.
- Guides
How to import a local GGUF file into Ollama with a Modelfile, using Gemma 4 31B
Two commands run a GGUF file you already have. What Ollama showed after the import, and the second copy of the file it makes on disk.
- Guides
How to turn off thinking in Ollama for Gemma 4, and what it changed in our test
One setting stopped the empty replies and cut the typical wait roughly in half. Two ways to set it, how to check your model supports it, and the numbers from 15 runs.
- Reviews
Gemma 4 31B QAT on an M4 Pro with 48 GB: 11.7 tokens a second, after a long think
It fits a 48 GB Mac with room to spare and writes at a steady pace. It also spends most of its tokens thinking before the reader sees a word.
- Research
Why Gemma 4 gives an empty response: we counted where its tokens go
In 18 timed runs the model spent most of its tokens on thinking nobody sees. With 600 tokens allowed, it ran out before answering in 6 runs of 15.