Data and index / Other
GLM 4.7 Flash Q4_K_M
on Mac, M4 Pro, 48 GB
Fits with about 27 GB to spare. Thinks through a 600 token limit in all 15 runs.
- Download this entryAs a CSV file
- How we testHow a figure gets in
- Run it yourselfThe script, step by step
- Cite this entry
- Follow Onticpost
Run history
Every run
| Date | Speed | Memory |
|---|---|---|
| 5 October 2026 | 56.2 t/s | 20.8 GB |
Change over time
How each figure moved, from the earliest run in each period to the latest.
The table fills once a second run is recorded.
Did this match your machine?
Readers who ran the script tell us whether they saw the same.
Against the maker's claim
The maker publishes no speed or memory figure for this build, so there is nothing to set our figures against.
How we measured
The set up for this entry has not been written down yet.
Run it yourself
Send us your result in the same format and it joins the index as a community entry, with your sample size.
The script needs Ollama or LM Studio running on your machine, with the model loaded. It changes nothing and sends nothing; it prints the line to send us.
About GLM 4.7 Flash Q4_K_M
What is GLM 4.7 Flash?
GLM 4.7 Flash is a 30 billion parameter mixture of experts model from Z.ai. This entry measures the Q4_K_M build, an 18.13 GB file, in LM Studio on a Mac with an M4 Pro chip and 48 GB of memory.
What was measured
Fifteen runs across five tasks in LM Studio, llama.cpp Metal runtime 2.47.0, context 65,536 tokens, 600 tokens allowed per run, on macOS 27.0.1. Writing speed is the middle run. Memory is the largest size the model process reached. In all 15 runs the whole allowance went on thinking and no answer was written.
Other figures
| Measure | Result |
|---|---|
| Slowest and fastest run | 52.53 and 58.70 tokens a second |
| Wait for the first token, short prompt | 0.23 seconds |
| Wait for the first token, 3,657 token prompt | 10.5 seconds |
| Runs with no answer inside 600 tokens | 15 of 15 |
| Model file on disk | 18.13 GB |
Measured with the same script as every other entry on this machine.
Sources
Cite this page
Hill. "GLM 4.7 Flash Q4_K_M". Onticpost, published 5 October 2026. https://onticpost.com/index/glm-4-7-flash-m4-pro-48gb/
Updated 5 October 2026. This address stays the same when the piece is updated.
Similar entries: the same model
No other machine has run this model in the index yet. Run the script above and yours can be the first.
Similar entries: this machine
| Entry | Speed | Fit |
|---|---|---|
| Bonsai 27B Q1_0 | 30.0 | |
| Qwen 3.6 35B-A3B Q4_K_M | 62.5 | |
| Qwen 3.8 27B Q4_K_M | 10.1 | |
| Gemma 4 31B QAT Q4_0 | 11.7 |