Another Agent Harness, Really?
Yes, really.
What began as a joke turned into a serious side project. It taught me a lot about how coding agents work and left me wondering how much of a moat software itself still provides.
The loop inside the harness
Harness · session, tools, permissions
- ReadRepository + task
- ActTools + permissions
- VerifyResult + task completion
Findings return to context
After seeing what felt like the 101st agentic coding harness launch on X, I figured this market had to become saturated at some point. Claude Code and Codex already exist. They come from companies with enormous engineering budgets. Not just software companies, either, but the frontier AI labs building the models these products depend on.
Surely they were in a pretty good position to cover this.
Yet other teams kept building their own harnesses, addressing workflows and preferences the larger products didn't quite accommodate. My contribution to this observation was, apparently, to make the problem worse.
A joke with a specification
I asked GPT 5.6 Pro to review the leading open-source coding harnesses on GitHub and design one that covered their most important features. My constraints were simple: zero dependencies, entirely Python.
That is a paraphrase, but it captures the level of ceremony involved. I wasn't starting a company or responding to a carefully researched gap in the market. I wanted to see what would happen.
A couple of hours later, ChatGPT came back with a result good enough that I wanted to keep working on it. The project became Borealis Coder.
It now has an interactive terminal interface, support for different model providers, persistent sessions, code-editing tools, permission controls, checkpoints, and verification. The original zero-dependency constraint didn't survive unchanged: the runtime now uses prompt-toolkit alongside the Python standard library.
I'm mentioning that because “I prompted a complete coding agent into existence” would be a convenient story. It would also leave out the subsequent work. The first result made the project worth pursuing. It didn't make the project finished.
What the harness actually does
A coding harness is the software around the model: it supplies tools and context, manages the session, and controls what actions the agent can take. From a distance, the job looks simple. Give a model some tools, let it inspect a repository, and keep calling it until the task is done.
Almost every interesting problem is hidden inside that last sentence.
What does the model need to read before it can make a useful change? How much of that information should remain in context? Which actions need approval? How do you recover when an edit goes wrong? How do you tell whether the task is done rather than whether the model has decided to stop?
Two products can use the same model and behave differently because they give it different information, tools, permissions, and feedback. Building Borealis made those decisions much more concrete for me. It also gave me a particularly annoying example.
Efficient, without losing anything, please
At one point, Borealis kept compacting its context every couple of turns. Compaction is supposed to make room for more work by replacing accumulated conversation and tool output with a shorter account of what matters.
When I looked into it, the context was only shrinking by about 20%. A couple of turns later, the agent was compacting again. It was repeatedly interrupting the work to make just enough room to need another cleanup almost immediately.
Just enough room to do it again
- Before compaction
- Context accumulated so far
- After compaction
- About 20% less context; the rest remains
Work adds context againAnother compaction
Part of the problem, I realized, was in what I had asked the AI working on the codebase to optimize for. I wanted high efficiency and the best possible output quality. Both sounded reasonable. I hadn't said enough about what should happen when they pulled in different directions.
Keeping more information can help an agent avoid losing something important. But if you preserve so much that you're constantly interrupting the work to compact again, that caution has a cost. Compress more aggressively and you risk throwing away something the agent will need later.
“Make it efficient without compromising quality” didn't settle that trade-off. It left the AI to settle it for me.
Things like this kept coming up. I could ask for a feature and get an implementation, but I still had to use it, notice where the behavior was wrong, and work out which of my requirements needed to be more specific.
For compaction, a better brief would say what must survive: the current task, decisions already made, unresolved failures, and references to relevant files. It would distinguish that information from material the agent can look up again. It would also define how much room compaction should create for continued work. Then I'd need to check whether the agent could still finish the task with the reduced context.
That is an example of a more useful specification, not a claim that there is one compression ratio that works for every task. A smaller context is no victory if the agent forgets what it was doing. Preserving detail isn't free if the result is constant interruption.
The lesson for me was about how I prompt the agent building the software: decide which goals take priority, where compromises are acceptable, and how I'll recognize a bad compromise when I see one.
The uncomfortable part
Despite that work, I was surprised by how little it took to get a serious starting point.
The architectural ideas were public. Existing projects had already explored many of the important decisions. A capable model could examine that material and help turn it into another implementation.
I don't have a benchmark showing that Borealis matches Claude Code or Codex. A feature list wouldn't establish that, either. Having something called “context management” tells you very little about how well it works during a difficult task. I had just given myself a demonstration of that.
Even so, the experience made me question an assumption: that the effort required to build a piece of software necessarily gives its creator much protection.
Something can take substantial work to build the first time and still become relatively easy for someone else to reproduce. Once the behavior is visible and the relevant ideas are available, the next implementation may require much less effort.
If your advantage depends mostly on competitors not having enough engineering time to build the same features, that seems like an increasingly uncomfortable place to be.
Is there a moat left?
I wouldn't conclude from a side project that software businesses no longer have moats.
Getting someone to try a product is a different problem from implementing it. Giving them a reason to stay is another. A tool can become difficult to replace because people trust it, because it fits their work unusually well, or because moving away would disrupt processes they depend on. My experiment didn't reproduce any of that.
The compaction problem also complicates my own argument. Getting a plausible implementation was surprisingly easy. Making the decisions that turned it into something I wanted to use took more involvement. I'm not sure that judgment amounts to a durable moat, but the code certainly hadn't made it unnecessary.
And the obvious problem with finding a niche that larger products ignore is that another small team can notice the same niche. They have access to these tools too.
I don't have a satisfying answer to where that leaves software development. I do think “we built it, and building it was hard” needs more scrutiny as a business argument.
As for why I built another coding harness: you probably don't need it. Most people will probably never use it.
I kept working on it because it taught me things I wouldn't have learned by watching another launch on X. Somewhere along the way, I became the person whose launch had made me roll my eyes in the first place.