Issue 01
What a model does before it answers, and what it costs you.
- Dated5 October 2026
- Every issueWeekly issue
- Follow Onticpost
Issue 01 ยท 5 October 2026
Gemma 4 with thinking off: 15 runs, no empty replies and answers in half the time
With thinking off in Ollama, every run answered and the typical reply took 27.2 seconds instead of 53.8. The writing speed did not change.
- 15 runs, every one answered
- Reply in 27.2 seconds, not 53.8
- Writing speed stayed at 11.6
In this issue
Top stories
Flourish sells its Flourish One home robot as experimental, with no warranty and no returns
Qualcomm agrees to buy the company behind MoveIt in the Qualcomm PickNik deal, with no price stated
Intrinsic releases Intrinsic Core as open source under Apache 2.0, with a reference design for CNC machine tending
Boston Dynamics Atlas robots start training at a new center inside Hyundai’s Georgia plant
Agility Robotics Digit earned $1.78 million in 2025, its merger filing with Churchill Capital Corp XI shows
Arduino releases the Arduino VENTUNO Q robotics board, and its own pages disagree on its power input
Aleph Alpha Kolibri is a 78.84 GB open-weight model, and its own documents disagree on German data
IQuest releases IQuest-Q1, a 320 billion parameter open weights model, on September 28
DASLab publishes a 58.41 GB Qwen3.8 Flash Next GGUF build that keeps half the experts
Aikido announces Altar-1, GLM-5.3 cut to 168 experts and 327.99 GB
MiMo V2.6 is open weights from Xiaomi, and its 9B model runs at 24.09 tokens a second on a 48 GB Mac
DeepSeek V4.1 Flash weights are out: 552B in the announcement, 763.2B in the files
IFM releases K2 Horizon open weights, and its smallest model will not load in LM Studio or Ollama
Research
Can a local AI agent draw an SVG illustration? Three tried, none got there
Pi vs OpenCode vs Hermes on the same local model: one brief, three results
Gemma 4 with thinking off: 15 runs, no empty replies and answers in half the time
Ollama’s default context length for Gemma 4 31B on a 48 GB Mac is 32,768, and memory climbs as you use it
Why Gemma 4 gives an empty response: we counted where its tokens go
Reviews
Guides
How to run one prompt through Pi, OpenCode and Hermes from the command line
How to import a local GGUF file into Ollama with a Modelfile, using Gemma 4 31B
How to turn off thinking in Ollama for Gemma 4, and what it changed in our test
How to measure tokens per second in LM Studio without the cache fooling you