What was measured
Fifteen runs across five tasks in LM Studio, llama.cpp Metal runtime 2.47.0, context 65,536 tokens, 600 tokens allowed per run, on macOS 27.0.1. Writing speed is the middle run. Memory is the largest size the model process reached.
Other figures
| Measure | Result |
|---|---|
| Slowest and fastest run | 11.40 and 12.04 tokens a second |
| Wait for the first token, short prompt | 1.43 seconds |
| Wait for the first token, 3,806 token prompt | 46.1 seconds |
| Model file on disk | 17.65 GB |
The full account is in our review.