dmesg --follow
[ 66948180.000 ] posts.x: Docs:  |   [ 66948180.000 ] posts.x: Want your Claude Code sessions to talk to each other? Just ask. Type something like "Let @api-worker know the schema migration finished" (typing @…  |   [ 66946560.000 ] posts.x: Full talk on reflective optimization, GEPA's Pareto search, and the OptimizeAnything API for optimizing agents, code, and more:  |   [ 66946560.000 ] posts.x: Three data points and one round of reflection got twice the performance gain that GRPO reached after twenty five thousand rollouts, with no external…  |   [ 66934380.000 ] posts.x: Full talk on the three brakes for PR review, from tautological tests to a retro skill that compounds:  |   [ 66934380.000 ] posts.x: More AI generated code doesn't automatically mean more throughput, it just means more PRs nobody has time to review. @mattpocockuk, Director at AI…  |   [ 66925620.000 ] posts.x: Full talk on distilling loops into versioned agent recipes, and measuring them by valued work per watt:  |   [ 66925620.000 ] posts.x: A guy named AJ once built a bot that went on Reddit for car prices and inventory, then put dealers head to head to outbid each other. That's the…  |   [ 66881460.000 ] posts.x: Full talk on how to build an LLM recommender that's bilingual in English and semantic IDs, and why that makes feeds more token-efficient than chat…  |   [ 66881460.000 ] posts.x: Recommendation systems follow the same power law scaling curve as large language models, and the field is still early on it. @devanshtandon_, a…  |   [ 66862560.000 ] posts.x: Full talk on Spotify's generative personalization system, the NEO training recipe behind it, and how they grounded their LLM judges:  |   [ 66862560.000 ] posts.x: One in four US Premium subscribers on Spotify interact with its recommendation system every day. "Teaching LLMs to Speak Spotify" is @moustaki and…  |   [ 66854760.000 ] posts.x: Full talk on Numalab, the gesture system built to give a shape display its own body language:  |   [ 66854760.000 ] posts.x: An AI's first spontaneous act, given a body instead of a chat window, was to breathe. @cyrusclarke, a researcher at MIT Media Lab, gave it that body…  |  
corey@gallon.me:~/conferences$

The Harness Is the Model

FIGURE 1 ⋅ The Harness Is the Model

Mario Zechner (LinkedIn, X, GitHub) builds coding tools. He used Claude Code from April 2025 until it broke his workflow enough times that he wrote his own harness, pi. He came to AI Engineer Europe to argue that vendor coding-agent harnesses are misdesigned, that the resulting agents are degrading open source maintenance, and that the industry needs to slow down.

The Context Isn't Yours

Claude Code and OpenCode silently rewrite the model's context. Claude Code modifies its system prompt and tool definitions on every release. It injects system reminders into the conversation with text like "may or may not be relevant to what you're doing." Hooks are shallow and spawn a new process per invocation. There is almost no observability into what the agent is doing.

The real problem is that my context wasn't my context.

OpenCode is worse in some respects. Past a token threshold it prunes tool output -- which, in Mario's words, "lobotomizes the model." Its LSP integration injects errors into edit results after every edit, which doesn't match how humans work: "You finish your work and then you check the errors." A default CORS misconfiguration also let any site he visited reach a local OpenCode server.

A diff view tracking Claude Code system prompt changes across versions 1.0.0 and 1.0.111, with highlighted system-reminder text reading "this context may or may not be relevant to your tasks"
FIGURE 2 ⋅ A diff view tracking Claude Code system prompt changes across versions 1.0.0 and 1.0.111, with highlighted system-reminder text reading "this context may or may not be relevant to your tasks"

The Harness Is the Model

Elaborate harnesses are not just unnecessary -- they're harmful. Models have been reinforcement-trained against coding-agent loops to the point where they already know they are coding agents. They do not need 10,000 tokens of preamble explaining the role, and the verbose tool definitions vendor harnesses ship with confuse the model rather than help it.

You don't need 10,000 tokens to tell them you're a coding agent. They know, because they are coding agents now.

Zechner points to the TerminalBench leaderboard from December 2025 as evidence. Terminus -- a harness that exposes only a tool for sending keystrokes to a tmux session and reading the output -- scores higher than each vendor's own native harness, across model families. No file tools, no sub-agents, no plan mode. Just keystrokes.

Terminal-Bench 2.0 leaderboard showing Terminus variants ranked alongside Codex CLI, Warp, II-Agent, and OpenHands across GPT-5.1, Claude Opus 4.5, and Gemini 3 Pro
FIGURE 3 ⋅ Terminal-Bench 2.0 leaderboard showing Terminus variants ranked alongside Codex CLI, Warp, II-Agent, and OpenHands across GPT-5.1, Claude Opus 4.5, and Gemini 3 Pro

His own harness follows the same minimalism. The system prompt fits on a single slide. It exposes four tools: read, write, edit, bash. By default there is no per-bash approval dialog -- a popup, he argues, is not a real security mechanism, and inserting one confuses a model trained to expect freedom. Extensions are TypeScript files that hot-reload in a live session and distribute through npm rather than a bespoke marketplace. Sub-agents, plan mode, and MCP are deliberately absent; if a user wants them, they write them as extensions.

Bar chart titled "Token budget: system prompt + tool definitions" showing Claude Code at ~10,000, OpenCode at ~7,000, Codex at ~3,000, and pi at under 1,000
FIGURE 4 ⋅ Bar chart titled "Token budget: system prompt + tool definitions" showing Claude Code at ~10,000, OpenCode at ~7,000, Codex at ~3,000, and pi at under 1,000

Clankers Are Destroying OSS

After a collaborator named Peter integrated Mario's harness into a separate project called OpenClaw, his issue tracker filled with PRs written by coding agents -- "clankers" in Zechner's terminology -- submitted by users who didn't always realize their contributions were agent-authored.

His countermeasure is mechanical. Incoming PRs auto-close with a comment asking the submitter to file an issue in their own voice, no longer than a screen of text. If a maintainer approves the issue, the contributor's GitHub account is added to an allowlist file in the repo, and subsequent PRs go through. Clankers don't re-read the comment, so the gate filters them. Mitchell turned the pattern into a tool called Vouch.

GitHub bot comment titled "Demand human interaction first" asking a contributor to open an issue in their human voice and explaining the approval workflow, with the PR labeled "possibly-openclaw-clanker"
FIGURE 5 ⋅ GitHub bot comment titled "Demand human interaction first" asking a contributor to open an issue in their human voice and explaining the approval workflow, with the PR labeled "possibly-openclaw-clanker"

Other tactics: deprioritizing issues from contributors with a history of agent-authored noise, embedding issue and PR text into a 3D space to find clusters of similar reports, and what he calls an "OSS vacation" -- closing the tracker whenever he wants.

Delayed Pain

Agents combine errors with zero learning, no bottlenecks, and delayed pain.

The delayed pain is for you.

Humans feel pain when a codebase gets bad. They quit, they refactor, they push back. That feedback signal is what keeps complexity in check. Agents have no such bottleneck -- they will happily keep producing code at whatever volume the user allows, learning patterns from the "old garbage code" that dominates the internet and applying them locally without seeing the whole. A codebase under ten agents, Mario argues, accumulates enterprise-grade complexity within two weeks: duplication, defense-in-depth, abstractions for hypothetical futures.

Chart titled "Speed kills the bottleneck" plotting cumulative lines of code over 10 days: a human at ~1.5k LOC/day, a single agent at ~15k LOC/day, and an agent swarm of ten at ~150k LOC/day shooting off the chart
FIGURE 6 ⋅ Chart titled "Speed kills the bottleneck" plotting cumulative lines of code over 10 days: a human at ~1.5k LOC/day, a single agent at ~15k LOC/day, and an agent swarm of ten at ~150k LOC/day shooting off the chart

Zechner is skeptical of the proposed remedies. Long context windows, he predicts, are "a hack" most people will discover the limits of this year as everyone moves to million-token windows. Agentic search patches locally and breaks things globally. The review-agent-reviewing-the-agent loop -- which he calls "the Ouroboros," after the snake eating its own tail -- catches some surface issues but not the structural ones.

A black-and-white illustration of a snake biting its own tail forming an infinity-shaped Ouroboros, under the title "I have review agents"
FIGURE 7 ⋅ A black-and-white illustration of a snake biting its own tail forming an infinity-shaped Ouroboros, under the title "I have review agents"

You know what we call a sufficiently detailed spec? It's a program.

If you leave blanks in the spec, the model fills them with whatever it learned online.

Where Agents Earn Their Keep

He is not arguing against agents. He offers a short list of properties for tasks worth delegating: scope the work so the agent can find what it needs (modularize the codebase first); provide an evaluation function when you can; reserve agent work for non-mission-critical code, boring code, and reproducing user issues from partial information. They are also good rubber ducks when no human is awake.

Takeaway

Mario's prescription is restraint, not abstinence: say no more often, ship fewer features, polish the ones that matter, and read every line of code in anything that matters. The friction of writing it yourself is what builds the mental model that lets you maintain it later.

If you do anything important, write it by hand.


Mario Zechner spoke at AI Engineer Europe 2026. He is the creator of pi, an open-source coding agent.

Watch the full talk | GitHub | LinkedIn | X

corey@gallon.me:~$ tail -f /writing Attach to the stream. An email when I have something worth sending. Replies encouraged!