Pi vs OpenCode vs Hermes on the same local model: one brief, three results

Same Mac, same Gemma 4 31B, same four files. One agent took 22 minutes, one drew five rectangles, and one printed nothing until it was done.

This is Pi vs OpenCode vs Hermes on the same local model. We pointed all three at Gemma 4 31B on one Mac, gave them one brief, and ran each once. The model was the same, the files were the same, and the results were not.

Pi vs OpenCode vs Hermes: does the agent matter on the same local model?

It did here. The brief was to think of a picture for an article and draw it as SVG. All three reached the same kind of idea. They differed in how long they took, how much they drew and how well the drawing held together.

Agent Time Shapes Lines Script size
Pi 22 min 04 s 11 6 1387 bytes
OpenCode 15 min 13 s 5 0 727 bytes
Hermes 15 min 55 s 9 2 962 bytes

Which was fastest: Pi, OpenCode or Hermes?

OpenCode, at 15 min 13 s, with Hermes 42 seconds behind. Pi took 22 min 04 s. The quickest run was also the one that wrote the least: OpenCode’s script is 727 bytes, about half the size of Pi’s. Time here tracks how much the agent chose to do, not how fast the model writes, which was the same for all three.

Which made the best drawing?

Hermes. Its picture, a tall sheet of paper with scissors cutting across the top, was the only one that could be read at a glance. Pi drew more parts than anyone and stacked them on top of each other. OpenCode drew five rectangles. None was fit to publish, and the pictures are described in our account of the drawing test.

What went wrong in each run?

OpenCode’s first attempt to save its script failed: it called its file tool with the wrong name for the path, got an error, and saved the file correctly on the second try. In doing so it also changed the drawing, moving one strip of paper 95 units down. It then ran the script once and stopped, skipping the step that asked it to check its own work.

Hermes in one shot mode prints nothing until it has finished, so for 15 min 55 s there was no sign of life. It also has two ways to run code, and one of them, as ours is set up, waits for a person to approve it. We told it to use the other.

Pi ran without incident. In its print mode it shows only the final reply, so we cannot say which steps it took to get there.

Why would three agents get different results from one model?

Each wraps the model in its own instructions and its own tools. The model reads those before it reads the task, and they shape what it does. We ran Pi with its skills and context files switched off. OpenCode loaded its usual instructions, which we could not switch off from the command line, and Hermes ran with its own. So the three runs were not identical in everything but name, and part of the difference will come from that.

Which one should you use with a local model?

One run each does not settle it. On this task Hermes gave the best result and OpenCode the quickest, and Pi was the easiest to run cleanly. If the work matters, run your own task through all three, which takes one command each: see how to run one prompt through Pi, OpenCode and Hermes.

How we tested

A Mac with an M4 Pro chip and 48 GB of memory, on macOS 27.0.1. Gemma 4 31B QAT in LM Studio at a context window of 65,536 tokens. Pi 0.99.1, OpenCode 1.18.31 and Hermes 0.21.5, each run once on 4 October 2026 in its own folder holding the same four files. Time is the clock time from starting the command to its exit. Shapes and lines are counted from the script each agent wrote.

Sources

  1. Google's Gemma 4 31B QAT model card
  2. Pi coding agent package page
Cite this page

"Pi vs OpenCode vs Hermes on the same local model: one brief, three results". Onticpost, published 4 October 2026. https://onticpost.com/pi-vs-opencode-vs-hermes-same-local-model/

Updated 4 October 2026. This address stays the same when the piece is updated.

Share this