Ask a coding agent for a feature and you get the feature plus forty tests. They all pass. You stop reading them somewhere around test six.
I don’t blame the agent. I told it to leave the change well tested, and the cheapest way to look well tested is volume. A mock returns a list, so the test checks that the list came back. A switch has three cases, so the test restates all three. Every green run makes the next refactor more expensive, because now forty tests have an opinion about how the code is shaped.
Peter Steinberger saw the same thing at a much bigger scale:
OpenClaw deleted around 400k LOC of its own tests without much change in code coverage. Modern models just love writing tests for every tiny change, even if they aren't useful. This skill helped. https://t.co/a1jSBpfpns
— Peter Steinberger 🦞 (@steipete) September 24, 2026
His tweet sent me back to a book I had meant to reread for years, Orta Therox’s Pragmatic Testing, and I wanted the same cleanup for a Swift app.
The screenshot at the top is the result in Oratio, my iOS reader app. The suite dropped from 1,224 tests to 846, and fourteen thousand lines of test code went away. This post is about the bar that decided which tests stayed, and the numbers that checked whether the agent deleted something it shouldn’t have. It did, once. The numbers caught it.
If you only want the skill, it is a gist. Drop it into your agent’s skills folder.
What junk looks like
All three of these pass. An agent wrote each one in good faith.
The first restates a switch. The library screen has a count that is zero while loading, zero on failure, and the number of documents once they arrive:
#expect(Documents.loading.count == 0)
#expect(Documents.failed.count == 0)
#expect(Documents.loaded([]).count == 0)
#expect(Documents.loaded(summaries).count == 3)That is the production code read back line by line. It says nothing about whether the screen shows the right thing.
The second computes its expected value with the same branching as production:
let expected: VoiceExperienceSelectionEligibility = if version == "2.0" {
.installationRequired(currentStatus: .unavailable("Requires Oratio 2.0 or later."))
} else if status != .installed {
.installationRequired(currentStatus: status ?? .notInstalled)
} else {
.available
}The production rule, copied into the test and expanded into twenty rows. If the precedence is wrong, it is wrong in both places, and the test agrees with the bug.
The third is my favourite. A test called chunkControlsSkipChapterHeadings drove the reader’s next and previous buttons and checked that playback stepped over headings. The skipping happened inside the hand-written mock playback manager, which carried its own copy of the pipeline’s algorithm:
// MockPlaybackManager, in the test target
if candidate.type == .body {
remainingBodySteps -= 1
}The test proved that the mock skips headings. Delete the check from the real pipeline and it still passes.
They rarely fail when the product breaks. They always fail when you move code around.
A book from 2016
Orta wrote Pragmatic Testing as a short book about testing iOS apps. It comes from the Objective-C era, and several chapters were never finished. The tooling has aged. The ideas haven’t.
If you try to test every
ifin your app, you’re going to have a bad time.
He frames coverage as a fight between “perfect” and “enough”, and sides with enough. You ship apps, not codebases. He also paints with two brushes. A wide brush, like a snapshot of a whole screen, covers setup and wiring in one pass. A fine brush goes to the edges where the logic lives.
My agents painted every pixel with the fine brush, then painted it again one layer up.
I part ways with the book in one place. Orta avoids mocking code he owns. I kept the generated mocks and changed how tests use them.
The bar
I started from Peter’s skill, but it is written for a TypeScript monorepo and half of it is about his repository’s scripts. So I had the agent read Orta’s book with me, idea by idea. Some ideas we kept, some we reworked for Swift and for agents, and a few we dropped. The skill that came out hangs on three questions. A test that can’t answer all three doesn’t get written, and an existing one gets deleted.
- Which bug does it catch? Name a plausible one-line mistake. Not “someone deletes the function”, but a mistake a refactor or a wrong fix could actually introduce.
- What ships broken if I delete it? If the answer is “nothing, another test covers that”, it goes.
- Does it survive a refactor that keeps the behaviour? A test that breaks when you move code around without changing what the user sees is testing the shape of the code.
Expected values come from the rule, written as literals. The second junk example could only ever agree with the code, because it computed its answers with the code’s own branching. The rewrite has five rows, each a fact about the product with a short comment saying why it holds. One of them says outright that an incompatible app version wins even when the voice is installed.
One owner per contract. Each rule gets tested once, at the layer that owns it. The screen above it gets one test that it wires the result through. Inputs that only change the data become rows of one parameterized test.
Stub freely, verify sparingly. Mocks can return whatever the test needs. Checking that a call happened is worth it when the call is an effect the user would notice. “A failed delete must never remove the book’s files” is a good verify. “The use case asked the repository for the list” is not, because the list could not have come back without the call.
A budget. A behaviour change usually lands one to three tests. More needs a reason, and races, persistence and money count. No test is also a valid answer for a one-line forward, as long as the agent names the existing test that would catch a mistake there.
Bug fixes still start with a failing test. After the fix the agent keeps it, folds it into an existing table as one more row, or drops it because a stronger test now covers the bug. I don’t ask for red, green, refactor on every change. I ask the agent to break the production line once after writing the test and watch it go red.
The skill is a gist if you want to drop it into your own repository. The examples are Swift. The rules are not.
The run
The agent worked through the suite one module at a time, one commit per batch. My part was the doubles policy. Every protocol gets a generated mock, and a hand-written double has to say in a comment what the mock can’t do. That removed about forty hand-written doubles, and with them most of the mocks that carried copies of production logic.
| Before | After | |
|---|---|---|
| Tests | 1,224 | 846 |
| Test code | 45k lines | 34k lines |
| Hand-written doubles | ~54 | ~14 |
| Production changes | +452 / −360 |
Did we lose anything?
In Refactor by Numbers the rule was that the agent never gets to be its own referee. The agent that deleted the tests doesn’t get to say they were junk. So I measured coverage per function, before and after, on three modules:
| Module | Tests before | Tests after | Functions that lost coverage |
|---|---|---|---|
| KokoroTTS | 78 | 20 | 0 |
| ReaderDB | 81 | 63 | 1, by 10% |
| Reader | 452 | 321 | 22 |
KokoroTTS lost three quarters of its tests and not a single line of coverage. Those 58 tests ran code that other tests already ran.
Reader needed reading. Twelve of the 22 were one-line forwards, like a button handler that calls the coordinator. One was the count from the first example, the expected “no test needed” outcome. That left four functions where a real rule lost its test.
What the numbers caught
The worst was the welcome tour. When the app has no recorded audio for a scene, it fakes the narration by revealing words on a timer. That path went from 65% coverage to zero. The agent had deleted the only test that reached it, a bad one that slept for 0.7 seconds and checked that at least one word had appeared. A weak test that is the only one isn’t junk. It needs rewriting, and that is now a rule in the skill.
I also ran twelve mutants against both versions of the suite. Ten were killed on both sides and the same two survived, so the mutants missed the welcome tour entirely. The auditor only puts mutants in code some test still runs, and that path had no tests left. Coverage caught it. That is why the skill now checks coverage first and puts mutants only where coverage dropped.
The fix was three new tests and five assertions folded into existing ones, with no production change. The welcome tour test now waits for the scene to finish instead of sleeping, and asserts that all three words appeared. Every new assertion was watched failing under its mutant, including the two old survivors.
| Before | After prune | After fix | |
|---|---|---|---|
| Mutants killed | 10/12 | 10/12 | 12/12 |
| Welcome tour coverage | 65% | 0% | 88% |
| Functions with a real coverage loss | 4 | 0 |
What these numbers do not say
Twelve mutants is a sample, and 12/12 fits it, because I wrote the fixes after seeing which ones survived. A fresh set is the honest next check.
I measured three modules, not the whole app.
Fewer tests also fail per mutant. One used to break four tests and now breaks two. That is the point of one owner per contract, but some rules now rest on a single test.
Delete freely, measure before you merge
The test pyramid still stands. Agents should still write unit tests, and most should be small and fast. What changed is the bar. A test names the bug it catches, uses the rule’s own numbers, and lives in one place.
And when an agent deletes tests, it doesn’t grade its own work. Measure coverage per function before and after. Every drop needs a reason or a new test. Then break a few lines and see what still goes red.