The smart zone is gone before work starts
Start-up reading eats more than half of the smart zone — the stretch of context where a model still reasons well. The task gets what's left.
For coding agents that work while you're away
One command installs rules, a docs kit, checks and session hooks into your repository. Your agents stop losing context, stop writing into each other's zones, and stop overpaying just to start a session.
The problem
Several agents, many sessions, and no human holding the whole project in their head at every moment — that's where it falls apart. What we keep seeing on live projects:
Start-up reading eats more than half of the smart zone — the stretch of context where a model still reasons well. The task gets what's left.
Agentic development moves so fast that docs written for humans go stale almost as soon as they're written. A person can usually work out which copy is right. For an agent, a stale doc is fatal: it believes what it reads.
Left to keep docs on their own, agents do it clumsily: closed issues and useless notes pile up, and every session drags them along for the life of the project.
Splitting a big project into smaller subprojects should help. Instead, agents get confused and work outside their own zones — and the chaos only grows.
After auto-compaction, docs the agent had read leave a trace of “I remember reading it” — and not a single value.
Every turn resends the whole context, and the context grows every turn. So the input
bill grows with the square of the session: b·n + g·n²/2. Twice the turns —
roughly four times the bill.
Eager to please, agents try to report what they've spent — but they have no tools for it and don't know where to look. So they make the wrong calls and misread what they find.
Who it's for
If you sit next to your agent, you don't need this. If your agents work while you're away, they don't work without it.
How it works
A zone is the folder an agent was started from; its docs live right there. It writes only to its own kit — a problem it finds next door becomes an entry in the neighbour's issues file. Agreements go stale; locations don't.
Map, open issues, decisions, archive, contract with neighbours — each with its own rule and its own token budget.
A closed issue moves to the archive whole and at once; links to it keep working. Session start cost depends on the number of open issues, not on the age of the project. There is no session log on purpose: 92% of its entries duplicated git.
Budgets, format, links — verified by machine at the end of every session. An agent can't testify about itself, so accounting is built on transcripts, not on the agent's reports.
What's inside
CLAUDE.md, root + zone: zone ownership, what to read and when, working
rules, what to update after work.
map, issues, decisions, archive,
howto — each with a token budget.
Deploys the kit into any project, filling in zone names.
Inject the zone map at start, log reads, daily summary, context-window meter with thresholds, a “where we left off” snapshot for the next session.
The agent asks for zones and prefixes and rolls out the standard — including migrating existing docs.
Turns a request into a task: items with a checkable “done when”, questions for the human.
Every rule with its reason and its measurement.
Machine check of the docs, run by a hook at session end. Ships checking the rules budget; our internal build runs thirteen checks.
The agent works through task items on its own, with safeguards and a gate that asks the human.
Archive + installer; update the harness on a live project without touching your own files.
A metrics catalogue and one collector over transcripts: who spends what, who makes mistakes.
How we differ
There are popular rule collections with huge GitHub followings. All of them are rulebooks and packaging, and none of them targets agents working unattended.
They sell rules. We measure whether they work.
Honest numbers
Some collections ship telemetry on by default — and still list the effect as “not published”. Ours come from transcripts on disk (Claude Code, September 2026), never from what the agent says about itself.
The cost law
input = b·n + g·n²/2
Cumulative input is base × turns plus the area under a growing context. Context grows linearly per turn (R² 0.96–1.00); the formula matched ten runs within 1.5% and held across 50 billion tokens of real sessions.
What the law gives
T* = 5·√(b/g)
The optimal moment to start a fresh session: 21–43 turns, 60–98k tokens of context. Restarting there, the law predicts up to 72% less input on long runs.
+68k
tokens per extra 1k tokens in the root
CLAUDE.md, over a 708-turn session — for every agent. 16% of that session's
bill was resending unchanged text.
500+
counted runs, across three models.
48×
input billed vs. final context on one 708-turn session.
Outside context
ETH Zurich, February 2026: a context file generated by an agent lowers success by 0.5–2% and costs 20% more; one written by a human adds +4% for up to +19%. A tool named in the context file gets called 50× more often. Our reading, stated openly: in “assistant” mode, docs barely pay off. We build for “tool” mode — the one that study didn't see.
Why it's a business
The genre leader gives away 25 skills and earns on courses. We tested “collect rules, sell a bundle” — and dropped it.
A subscription to maintained rules was dropped too: it competes with a free catalogue and can't prove its value.
The open spot: show users their own number — what they spend, where they overpay, whether the rules help in their project. Telemetry from the wild can't do that — there's no control group. An A/B inside the repository can. Technically hard — which is why the spot is empty.
Narrower than “everyone with a coding agent”, and worse served: agents working unattended, in parallel, across several zones.
FAQ
There is, and we agree — for “fix one bug in a mature repo with a human beside you”. It didn't measure the second session, the neighbour you can break, or work with nobody watching. That's exactly what the harness is for.
The rules are. The difference is the machinery: one-command install, budgets, a machine check, hooks that hand the agent what it needs and count what it spends.
The rules layer stays within a budget, and a script enforces it. The harness is an investment: a bare agent is always cheaper on a single task; the harness pays off over the long run.
For now, yes.
Write to us — we'll answer and can open the code for your review.
hello@claude-harness.techSmart zone — the part of the context window where a model still reasons well. Past it the model doesn't run out of room; it gets worse: follows instructions less reliably, loses details, makes more mistakes. Practitioners put the edge at roughly 40% of the window, about 80–100k tokens on today's models. The term is coding-agent practitioners' slang; the edge is a rule of thumb, not a measured constant.