← Back to Blog
Aug 15, 2026 •11 min read Engineering

The Bottleneck Wasn’t Writing Code. It Was Knowing When the Agent Was Done.

What coding agents taught me about infrastructure discovery, compute costs, and knowing when an agent has enough evidence to stop.

Cover image

Jensen Huang said he would be “deeply alarmed” if a $500,000 engineer used less than $250,000 worth of AI tokens in a year. I understand the argument: if compute makes an expensive engineer much more productive, refusing to use it can be its own kind of waste.

And it’s not like he or his business has any sort of vested interest in people consuming fairly expensive chunks of compute, right?

Of course not. Moving on.

My issue is with turning consumption into the goal. A routine task can quietly burn $200 of compute if you let an expensive agent wander. I would rather give repetitive investigation to a cheaper model and save the heavier reasoning for the work that actually needs it. Requiring engineers to spend six figures on tokens, regardless of the task, sounds like a billing strategy. Are we sure everyone knows what we are optimizing for?

At the other end of the spectrum are the one-shot demos: type a sentence, walk away, come back to a working application. They are impressive. They also hide a lot of work after the prompt. In one OpenAI experiment, Codex spent about 25 hours and 13 million tokens building a design tool, testing it and repairing failures. One human prompt started the job; it did not make the job one operation.

After using coding agents on a real brownfield system, I find a different question more interesting: how do I know when the agent is actually done?

The agents are good now

I also do not buy the “AI can’t code” argument anymore. Claude Code and Codex can trace unfamiliar code, inspect logs and history, implement meaningful changes and stay with problems longer than the coding assistants I was using a couple of years ago. I use them for that work.

The catch is that I can generate changes faster than I can responsibly understand them. I can start four agents in parallel and fairly quickly have four things that need my attention. I can ask one to fix a bug and have it find five other plausible improvements along the way.

Give a weak agent too much autonomy and it gets stuck. Give a strong agent too much autonomy and it can successfully change six things you never asked it to change.

Writing code is no longer the only bottleneck. Understanding the system, setting the boundary and verifying the result take real time.

Discovery before implementation

One of the biggest wins has been using agents to understand a codebase before asking them to change it. When you inherit an existing system, your first job is usually to figure out what is going on. Where does a request enter? Who writes this table? Why does that Lambda exist? Why does the repository suggest one deployment path while production appears to use another?

At Rafters, I inherited a working sports platform spanning React/Vite, Node/Express, PostgreSQL and several AWS services. The pieces worked, but the relationships between source code, deployed infrastructure and production data flows were not always obvious. The repository alone did not tell the whole story.

I used Claude Code and Codex as scoped investigators. Instead of asking for an explanation of the entire platform, I asked for one ingestion path, one deployment comparison or one answer that separated what the source proved from what we were inferring. I checked the conclusions against the actual system.

That got me to a useful map of the infrastructure and its important production relationships in roughly four days. For comparable manual discovery in a brownfield system, I would normally budget closer to three to four weeks before feeling comfortable making consequential changes. That is my estimate, not a controlled benchmark. Still, it is the kind of productivity improvement I care about: time to understanding.

What we knew, and what we only thought we knew

An agent can give a very coherent explanation before anyone has proved it correct. I started being more explicit about the status of each conclusion. “The source says this” is different from “this seems likely from the source.” Both are different from “we checked production and confirmed it.”

That distinction sounds fussy until you are about to touch production data.

How you end up playing whack-a-mole

One automatic job refreshed the platform’s list of free agents: players who were not currently on a team. The data is seasonal, but a player keeps the same ID across seasons. That relationship matters when you rebuild a list: a row from an earlier season can still collide with a row you are trying to add today.

The refresh began failing shortly after a performance optimization went out. PostgreSQL reported a unique-key violation. In plain English, the job tried to insert a player ID that the database was already storing in a field that was supposed to be unique.

The tempting explanation was immediate: new optimization, new error, therefore the optimization broke the job. I have made that leap before. It is an efficient way to fix the wrong thing, then discover the next symptom, then fix that too. Eventually the original defect is still there and you have edited half the neighborhood.

So I gave Claude a narrower job: inspect the source, schema and downstream consumers. No production queries, no AWS changes, no deployment and no rerun of the refresh. Its first useful finding was that the optimized query did not appear able to create a duplicate that the old query could not also create.

Then it found a mismatch in the data rules. The refresh deleted only rows for the current NFL season, while the database required player IDs to be unique across every season. If an older row for that player remained, inserting the current one could fail. That was a plausible explanation, not proof that it had caused this failure.

We stopped there and widened the boundary one step. Claude could read the existing Elastic Beanstalk logs, but still could not change the database or deploy anything. Those logs showed the same player ID failing multiple times on the old code, before the optimization was deployed.

Now the timing was no longer our evidence. The failure predated the change we had suspected. The next question became: what should the data model allow across seasons, and what is the smallest fix that preserves the intended behavior? That is a better starting point than tweaking a query because it happened to be nearby when the alarm went off.

From prompts to context

This is why I care more about context engineering than finding a clever prompt. The agent sees source code, repository guidance, logs, test output, prior messages and whatever tools I let it use. The quality of the answer depends on which of those things it can inspect and which conclusions it is allowed to act on.

More context is not automatically better context. I have had huge sessions that knew half the history of civilization and fresh sessions with a focused handoff that were far more useful. For the refresh issue, source inspection was enough to form a hypothesis. The logs came later, when we had a precise question for them.

That is how I ended up at harness engineering

I wanted the environment to remember those boundaries instead of making every task depend on a perfect opening prompt. That is the idea behind the development harness I am formalizing: a repeatable way to give an agent a defined job, the right context, a place to work and a way to prove what it did.

The structure I am building toward starts with a short task brief: the goal, the boundaries, the rules that must stay true and what would count as done. Versioned repository notes and runbooks would give the agent a map of the relevant architecture, so it does not have to rediscover the same facts every time. Recording the source revision makes the starting point clear. When independent work calls for it, a separate branch and Git worktree give the task its own copy of the code so parallel changes do not trip over one another.

The permissions should match the task. Reading source and logs is one level of access. Editing code is another. Production access is another again. At each step, I want a checkpoint: what did the agent observe, what changed and what evidence supports the result? Tests, diffs and logs can be checked outside the agent’s own summary. A human reviews consequential changes before they are integrated or deployed.

The bounded investigations, explicit stop conditions and human-controlled production decisions are already how I work. The repeatable briefs, versioned guidance and checkpoints are part of the more formal Rafters harness I am building toward. I do not want a neat diagram to imply I have a finished platform where I have a developing process.

Task harness diagram showing task briefs, separate workspaces, scoped permissions, tests, independent validation and human review
The task model I am formalizing: clear scope, a known starting point, appropriate access and evidence a reviewer can check.

The principle is simple: an agent saying “PASS” is a claim, not acceptance. I need to be able to reproduce the reason it passed. OpenAI describes a similar shift in its harness engineering work: as agents produce code faster, more engineering effort goes into the environment, guidance and feedback that make that work reliable.

A worktree is not a force field

A Git worktree gives a task its own directory and branch. That is useful when two agents are working independently. It does not, by itself, prevent a process from reading other files, using credentials or reaching network resources. If a task needs a security boundary, that comes from the sandbox and the permissions around it. The box in the diagram is a workspace, not magic.

Make the workflow remember

People talk about agents “learning” from each run, but the practical version I care about is less mysterious. If an agent repeatedly misses the same architectural rule, put that rule in repository guidance. If a regression escapes twice, add the test that should have caught it. If reviewers keep making the same correction, make the workflow surface it earlier.

The model weights did not change. The engineering system did. That is enough for me.

Parallel work fits the same logic. I will run separate agents on genuinely independent tasks, then review their outputs. I will not split one tightly coupled problem into eight branches just because eight looks futuristic. Once the agents can produce changes faster than I can inspect them, I become the bottleneck. Concurrent-agent count is not a productivity metric.

There is real overhead

Briefs take time. Worktrees need management. Verification costs compute. Permission boundaries can interrupt a reasonable command. I am not putting a seven-stage harness around a three-line typo; that would be an impressive way to automate myself into a bureaucracy.

Risk determines how much machinery a task deserves. More tokens, more tool calls and more agents can reduce risk on a difficult migration. They can also mean an agent reread the same three files for the fourth time because its context turned to soup.

The metric I care about is closer to correct, reviewable, maintainable changes per engineer-hour. Discovery counts. Avoided rework counts. A task whose correct result is “we do not have enough evidence yet; stop” counts too.

That last one does not make for a particularly exciting demo. It makes for pretty decent engineering.

The work moves up a level

I do not think engineers stop understanding or writing code. The more implementation capacity I can call on, the more important it is that I understand what I am asking the system to do. What are we changing? What must stay the same? What evidence would convince me that the result is correct? How much authority does this task actually need?

The code still matters. I just spend more of my time deciding what the code is allowed to change, and occasionally asking everyone to stop touching the thing for a minute.

AI makes it easier to get to an implementation. The surrounding engineering determines whether that implementation was worth getting to.