I benchmarked tdd-guard. It didn't write better code

I benchmarked tdd-guard. It didn't write better code

Sep 28, 20264 min read

tdd-guard blocks code edits until a failing test exists. 3 projects, 3 Claude Code configurations, a follow-up round, 54 runs. It didn't write better code, but it did buy consistency.

Repo · tdd-guard plugin

I benchmarked tdd-guard, a Claude Code plugin that blocks code edits until a failing test exists for them. The short version: it costs 2 to 4x more, it can make the agent skip requirements that no test asks for, and its real payoff only shows up once you start extending an existing codebase.

This is the long version of the video breakdown.

Subscribe

Regular posts on AI evals, agent reliability, and what the measurements actually show. No spam.

Testing setup

I used three small projects, each a different shape:

  • rate limiter (state-heavy)
  • CSV parser (transformational)
  • retry helper for async operations (control-flow heavy)

Each one ran under three Claude Code configurations:

  • A, no instruction: bare Claude Code.
  • B, CLAUDE.md only: the tdd-guard rules pasted into CLAUDE.md as plain instructions, with nothing enforcing them.
  • C, tdd-guard: the full plugin with the same rules as B, wired in as a hook that can deny a Write or Edit.

B and C use the identical rule set. The only difference is whether the rules are enforced or just requested.

Each project also got a follow-up prompt that extended the original spec. A single-shot run can't tell you much about discipline, so I wanted to see whether the plugin held up through a second round of changes.

harness diagram

That's 3 specs, 3 configurations, 2 rounds, and 3 runs of each: 54 runs total.

Judging

To grade quality, a script took the source and tests from two runs of the same project and put both into a prompt alongside a rubric, the original spec, and the follow-up spec. The judge picked a winner on four dimensions and explained each call.

judging pipeline

LLM judges tend to favour whichever candidate they see first, so I randomised the order on every call.

Each configuration had 3 runs, so comparing two configurations means pitting all 3 runs of one against all 3 of the other: 9 pairs. Doing that for A vs B, A vs C, and B vs C gives 27 pairs per project. Across three projects, with three judge calls per pair to get a majority, that's 243 judging rounds.

judging pairs

Quality results

Test quality. A won clearly. The judge liked its longer suites and better edge-case handling. C edged out B, but the judge's reasoning usually came down to C writing one extra test, so it was close.

Design quality. A won again, for the same reasons. B and C tied.

Spec adherence. This measured how closely each implementation matched the spec. A won because B and C both missed an edge case, and both skipped a piece of functionality the spec explicitly asked for. The TDD rules kept them anchored to the tests they'd already written, and once those passed, they stopped. I checked the test files to confirm: neither B nor C ever wrote a test for the missing case.

Restraint. Here the judge penalised the exact thing it had rewarded in the first two categories and called it scope creep. This is where tdd-guard's approach pays off: leaner implementations that do what was asked and nothing more.

judge win rates

So A wins on test and design quality because the judge likes bigger suites and extra handling, and B and C win on restraint. The spec adherence result is the one worth remembering: test-first discipline can stop an agent from building something the spec asks for, if no test happens to demand it. B and C also tracked each other closely on every dimension, so on quality alone it's hard to say what the plugin gives you over pasting the rules into CLAUDE.md.

Test count

A wrote more tests across the board, which lines up with the test quality result. The rate limiter was the exception, where tdd-guard nearly matched it.

test count divergence

Round 2 tests added

The follow-up round is where the plugin earns its place. When extending the codebase, C added new tests in every run, with the least variance of the three. A was the opposite: some runs added no new tests at all.

tests added round 2

That's the plugin doing exactly what it claims, forcing the same development cycle every time.

Cost

In round 1, tdd-guard cost 2.4x more than plain Claude Code on the CSV parser, 3.8x on the rate limiter, and 2.2x on the retry helper. In round 2 that dropped to 2.4x, 2.3x, and 1.9x. Extending an existing codebase is cheaper than bootstrapping a blank one, so the gap narrows.

cost divergence

That premium is what you pay for test-first discipline.

Verdict

tdd-guard costs 2 to 4x more on a fresh codebase and about 2x when extending one. Its strict rule of "failing test first, always" means anything without a test can slip through, including things the spec asked for. Where it earns its cost is ongoing work: the hook forces the same cycle every time, so tests keep growing with the code instead of relying on the agent to remember.


Repo: tdd-guard-bench · Video: youtu.be/EjtkHa_FH0s · tdd-guard plugin: nizos/tdd-guard