Ollama’s default context length for Gemma 4 31B on a 48 GB Mac is 32,768, and memory climbs as you use it
Ollama chose 32,768 tokens on its own and listed 19 GB. The process itself grew from 22.0 GB to 33.2 GB over 15 runs at 65,536.
Ollama picks a context window for you, and on a Mac with 48 GB of memory it picked 32,768 tokens for Gemma 4 31B. We then raised it to 65,536 and watched the memory. Ollama listed 19 GB both times, but the process measured by the operating system grew from 22.0 GB after one run to 33.2 GB after fifteen.
What is Ollama’s default context length for Gemma 4?
On our machine, 32,768 tokens. Ollama’s startup log said it chose a vram based default, and recorded 37.4 GiB of graphics memory available on this 48 GB Mac. The model itself supports 262,144 tokens, but Ollama did not use anywhere near that. The log also lists a setting called OLLAMA_CONTEXT_LENGTH, which was 0, meaning no override.
How much memory does Ollama use with Gemma 4 31B?
It depends on which number you read. ollama ps showed 19 GB and 100% GPU at both context sizes. The operating system’s measure of the Ollama process was larger and rose as we ran more prompts.
| When we measured | Context | Ollama process size |
|---|---|---|
| After one run | 32,768 | 22.0 GB |
| After the first run | 65,536 | 24.9 GB |
| After 15 runs | 65,536 | 33.2 GB |
| After 15 runs, thinking off | 65,536 | 31.4 GB |
We cannot say from this test why the process grew. Process size on a Mac can include memory the system could reclaim, so treat 33.2 GB as an upper reading and not as a requirement. What it does show is that the 19 GB on the listing is not the whole story, and that a 48 GB Mac had room for it.
Does a bigger context window slow Gemma 4 down?
We cannot tell. One run at the default context wrote at 11.92 tokens a second, and the middle of 15 runs at 65,536 was 11.50. One run against fifteen is not a comparison, and the gap is small. In LM Studio at 65,536 the same file wrote at 11.74, so the context size is not the main thing deciding speed here.
How do I change the context length in Ollama?
We set it per request. In the options of a call to /api/chat we passed "num_ctx": 65536, and ollama ps then showed a context of 65536. The server also has the OLLAMA_CONTEXT_LENGTH setting for a default that applies to everything. We did not test that one.
What did we not test?
The default context was measured with one run, not fifteen. We did not try other context sizes, other models or a Mac with less memory, where Ollama may choose differently. The same model in LM Studio is covered in Ollama against LM Studio, and the steps to load the file are in our import guide.
How we tested
Ollama 0.32.6 on a Mac with an M4 Pro chip and 48 GB of memory, macOS 27.0.1, the 17.65 GB Gemma 4 31B QAT Q4_0 file. Process size is the largest Ollama process as reported by the operating system, converted to decimal gigabytes. The method is on how we test.