The Harness Is the Model
Mario Zechner (LinkedIn, X, GitHub) builds coding tools. He used Claude Code from April 2025 until it broke his workflow enough times that he wrote his own harness, pi. He came to AI Engineer Europe to argue that vendor coding-agent harnesses are misdesigned, that the resulting agents are degrading open source maintenance, and that the industry needs to slow down.
The Context Isn't Yours
Claude Code and OpenCode silently rewrite the model's context. Claude Code modifies its system prompt and tool definitions on every release. It injects system reminders into the conversation with text like "may or may not be relevant to what you're doing." Hooks are shallow and spawn a new process per invocation. There is almost no observability into what the agent is doing.
The real problem is that my context wasn't my context.
OpenCode is worse in some respects. Past a token threshold it prunes tool output -- which, in Mario's words, "lobotomizes the model." Its LSP integration injects errors into edit results after every edit, which doesn't match how humans work: "You finish your work and then you check the errors." A default CORS misconfiguration also let any site he visited reach a local OpenCode server.

The Harness Is the Model
Elaborate harnesses are not just unnecessary -- they're harmful. Models have been reinforcement-trained against coding-agent loops to the point where they already know they are coding agents. They do not need 10,000 tokens of preamble explaining the role, and the verbose tool definitions vendor harnesses ship with confuse the model rather than help it.
You don't need 10,000 tokens to tell them you're a coding agent. They know, because they are coding agents now.
Zechner points to the TerminalBench leaderboard from December 2025 as evidence. Terminus -- a harness that exposes only a tool for sending keystrokes to a tmux session and reading the output -- scores higher than each vendor's own native harness, across model families. No file tools, no sub-agents, no plan mode. Just keystrokes.

His own harness follows the same minimalism. The system prompt fits on a single slide. It exposes four tools: read, write, edit, bash. By default there is no per-bash approval dialog -- a popup, he argues, is not a real security mechanism, and inserting one confuses a model trained to expect freedom. Extensions are TypeScript files that hot-reload in a live session and distribute through npm rather than a bespoke marketplace. Sub-agents, plan mode, and MCP are deliberately absent; if a user wants them, they write them as extensions.

Clankers Are Destroying OSS
After a collaborator named Peter integrated Mario's harness into a separate project called OpenClaw, his issue tracker filled with PRs written by coding agents -- "clankers" in Zechner's terminology -- submitted by users who didn't always realize their contributions were agent-authored.
His countermeasure is mechanical. Incoming PRs auto-close with a comment asking the submitter to file an issue in their own voice, no longer than a screen of text. If a maintainer approves the issue, the contributor's GitHub account is added to an allowlist file in the repo, and subsequent PRs go through. Clankers don't re-read the comment, so the gate filters them. Mitchell turned the pattern into a tool called Vouch.

Other tactics: deprioritizing issues from contributors with a history of agent-authored noise, embedding issue and PR text into a 3D space to find clusters of similar reports, and what he calls an "OSS vacation" -- closing the tracker whenever he wants.
Delayed Pain
Agents combine errors with zero learning, no bottlenecks, and delayed pain.
The delayed pain is for you.
Humans feel pain when a codebase gets bad. They quit, they refactor, they push back. That feedback signal is what keeps complexity in check. Agents have no such bottleneck -- they will happily keep producing code at whatever volume the user allows, learning patterns from the "old garbage code" that dominates the internet and applying them locally without seeing the whole. A codebase under ten agents, Mario argues, accumulates enterprise-grade complexity within two weeks: duplication, defense-in-depth, abstractions for hypothetical futures.

Zechner is skeptical of the proposed remedies. Long context windows, he predicts, are "a hack" most people will discover the limits of this year as everyone moves to million-token windows. Agentic search patches locally and breaks things globally. The review-agent-reviewing-the-agent loop -- which he calls "the Ouroboros," after the snake eating its own tail -- catches some surface issues but not the structural ones.

You know what we call a sufficiently detailed spec? It's a program.
If you leave blanks in the spec, the model fills them with whatever it learned online.
Where Agents Earn Their Keep
He is not arguing against agents. He offers a short list of properties for tasks worth delegating: scope the work so the agent can find what it needs (modularize the codebase first); provide an evaluation function when you can; reserve agent work for non-mission-critical code, boring code, and reproducing user issues from partial information. They are also good rubber ducks when no human is awake.
Takeaway
Mario's prescription is restraint, not abstinence: say no more often, ship fewer features, polish the ones that matter, and read every line of code in anything that matters. The friction of writing it yourself is what builds the mental model that lets you maintain it later.
If you do anything important, write it by hand.
Mario Zechner spoke at AI Engineer Europe 2026. He is the creator of pi, an open-source coding agent.
Watch the full talk | GitHub | LinkedIn | X