BlogEngineering

TDD with Claude Code: The Workflow That Gets What You Actually Asked For

The Claude Code TDD workflow I use: plan first, small stories, acceptance criteria as tests, Ready and Done gates, and verification evidence on every story.

In this article

Ask an agent to "build a feature" and it will almost always hand you something. It compiles. It runs. A couple of the obvious checks pass. And buried in the diff is a decision you never made, an assumption the agent quietly filled in because your instruction left a little room for it.

That gap, between what you asked for and what you actually got, is the thing I think about most these days. I have been building software for a long time, from the years when I shipped hundreds of small apps by hand to today, when agentic tools like Claude Code and OpenAI Codex do a growing share of the typing. So the question stopped being "can the agent write the code?" a while ago. It can. The question now is whether the code it wrote is the code I meant.

This post is the Claude Code TDD workflow I use to close that gap. TDD stands for Test-Driven Development: you define the expected behavior with tests first, and only then write the code that makes them pass. With an agent doing the typing, that order becomes a control system: plan first, break the work into small stories, turn acceptance criteria into tests, give the agent a Definition of Ready and a Definition of Done, and ask for verification evidence at every step. I used it on a small workflow tool I built recently, and it is also how the little Angular todo app in this guide comes together. None of it is exotic, and all of it is the difference between reviewing work and hoping for the best.

Why does Claude Code write tests that always pass?

Left to itself, an agent will often write the implementation first and then produce tests that describe whatever the code already does. The tests go green on the first run, the summary sounds confident, and nothing was actually verified. The test suite is a mirror, and a mirror agrees with everything it sees.

That is one failure mode of several, and they all have the same root: the agent optimizes for "done" unless you define "correct." The ones I run into most:

  • Tautological tests. The test asserts what the implementation happens to do, so it can never fail. Countermeasure: tests come from acceptance criteria written before the code, never from the code itself.
  • Tests written after the code. The same problem in a different order. Countermeasure: make "write the tests first, confirm they fail" an explicit step the agent must report on.
  • Weakened assertions. A test goes red, and instead of fixing the code the agent quietly loosens the expectation until it passes. Countermeasure: the acceptance criteria are the contract; a changed test is a changed requirement and needs your sign-off.
  • Scope creep. The agent implements the story and, while it's in there, rewires something you never asked about. Countermeasure: an expected-files list, which gets its own section below.

The industry numbers say this is the moment to care. The 2025 DORA report from Google found that 90% of software professionals now use AI at work, and that AI acts as an amplifier: it magnifies the strengths of a good delivery setup and just as happily exposes the weak spots.1 GitLab's 2026 AI Accountability research, surveying 1,528 developers and technology buyers, found that 85% agreed the bottleneck has shifted from writing code to reviewing and validating it, and 92% reported governance challenges with AI-generated code.2

That last number matches my experience exactly. For a lot of us the hard part has moved to knowing whether the generated code is correct, safe, maintainable, and actually aligned with what was requested. Years ago I built an e-invoicing integration for the Greek tax authority, and work like that teaches you early that "it looked fine" is not a sentence you get to use. Judging what an agent hands back turns out to be a skill you can name and practice, and I came back to it later in AI fluency and the 4D framework, where it goes by discernment.

TDD helps with every failure mode on that list because it gives the agent something far more useful than a broad instruction. It gives it a contract.

How do you get Claude Code to write tests first?

You ask for it before any implementation exists, in plain words: "Define the expected behavior with tests before implementation." Then you hold the order, tests, then code, story by story. The phrasing matters less than the sequence, and the sequence starts at the plan.

The most common way this goes wrong is starting implementation too early. You open Claude Code or OpenAI Codex and type:

Create a todo app with Angular.

The agent will happily produce something. It scaffolds components, services, templates, and tests, reaches for local storage, adds routing, picks a state pattern, and makes a dozen quiet assumptions about dependencies, structure, and styling along the way. The output might even be decent. The problem is that far too many decisions got made before the work was ever shaped.

A better first instruction looks like this:

Planning mode only. Do not edit files yet.

I want to create a todo app with Angular.

First, break the deliverable into small composable stories.

For each story, define the goal, acceptance criteria, Definition of Ready, Definition of Done, expected files to create or modify, dependencies, required tests, risks, assumptions, and implementation order.

Stop after the plan.

This changes the whole shape of the work. The agent now plans first instead of guessing its way into code, and you get something worth reviewing before a single line of production code exists. That review is where the real value sits, because it is far cheaper to correct a plan than to untangle a large diff after the fact.

On the tests themselves, one wording tip. Traditional TDD language says "write a failing test," which lands strangely on anyone who isn't steeped in the practice, because it sounds like a request to write something broken on purpose. Clearer:

Write the tests first. Run them and confirm they fail because the behavior is not implemented yet.

The failure is the point. It proves the test is checking something real, which is exactly the proof a tautological test can never give you. From there the loop is simple to hold in your head:

  1. Write the tests first.
  2. Confirm they fail for the expected reason.
  3. Implement the smallest change.
  4. Run the tests again.
  5. Refactor only when the tests are green.
  6. Report the evidence.

The first story is the scaffold

Story 0 sets up the project scaffold (architecture, folder structure, tooling, build and test) before any feature work begins
Story 0: create and verify the foundation before building features.

Before the agent builds any todo behavior, the first story should be the scaffold. If the Angular project doesn't exist yet, the agent should first create the application structure, confirm the testing setup, establish the folder conventions, and prove the app can build and test successfully. That is Story 0, and its job is to create the foundation every later story will lean on.

Story 0: Create the Angular application scaffold. The acceptance criteria could be:

  • Given the repository does not contain an Angular app, when the scaffold task runs, then a new Angular application is created.
  • Given the Angular app is created, then the project can be installed, built, and tested successfully.
  • Given the project uses a current Angular version, then standalone components are used unless the existing convention says otherwise.
  • Given the scaffold is complete, then a dedicated todo feature folder exists or is clearly planned, and no todo business logic is implemented yet.

Keep this story intentionally simple. That simplicity is scope control. Without Story 0, the agent tends to build the foundation and the first feature in the same breath, which means a bigger diff, more assumptions baked in, and a harder review at the end. With Story 0, the foundation gets verified first, and only then does the agent move into behavior.

Break the app into small stories

Once the scaffold is in place, break the app itself into small stories. A useful order for the todo app:

  • Story 0: Create the Angular application scaffold.
  • Story 1: Create the todo domain model and state service.
  • Story 2: Add a form to create a todo.
  • Story 3: Display the todo list.
  • Story 4: Toggle a todo between active and completed.
  • Story 5: Delete a todo.
  • Story 6: Filter todos by all, active, and completed.
  • Story 7: Persist todos in local storage.
  • Story 8: Add empty, validation, and error-safe UI states.

I spent years running restaurant kitchens before I ever wrote software, and the thing that keeps a kitchen alive on a busy night is that nobody tries to cook the whole order at once. Tickets come in, and you work them one plate at a time, in an order that makes sense, so the pass never turns into chaos. Breaking work down for an agent is the same instinct. The scaffold gives the app a foundation, the domain model gives it a data shape, the state service gives it controlled behavior, the UI stories consume that behavior, and persistence comes last because it is a genuinely separate concern.

This decomposition is what makes agents reliable. The agent never has to solve the whole app at once. It solves one small, testable behavior at a time, which cuts the ambiguity dramatically. It also cuts the review effort, because looking at one focused change is far easier than reviewing a large AI-generated diff that mixes setup, state, UI, validation, persistence, and styling all in one pass.

Acceptance criteria as tests: the contract the agent can't misread

Vague intent becomes clear acceptance criteria, then tests, then passing results, with scope trimmed to core behavior
From ambiguity to clarity: intent becomes criteria, criteria become tests, tests prove behavior.

Acceptance criteria written in Given/When/Then form are the single most useful thing you can hand an AI coding agent, because each criterion converts directly into a test the agent must satisfy. Take Story 2, adding a form to create a todo. A weak requirement would be:

The user can add todos.

That sentence looks simple, and that is the trap, because it leaves almost everything open:

  • What happens when the input is empty?
  • Should spaces be trimmed?
  • Should the input clear after submit?
  • Is a new todo active or completed by default?
  • Should duplicate todos be allowed?
  • Should pressing Enter submit the form?

The agent can make reasonable assumptions about all of these, but assumptions aren't requirements. Better acceptance criteria would be:

  • Given the todo input is empty, when the user submits the form, then no todo is created.
  • Given the todo input contains only spaces, when the user submits the form, then no todo is created.
  • Given the todo input contains text, when the user submits the form, then a new todo is added to the list.
  • Given the todo text has leading or trailing spaces, when the todo is created, then the stored text is trimmed.
  • Given the todo is added successfully, then the input field is cleared.
  • Given a todo is newly created, then it is active by default.

Now the behavior is specific, and the agent can turn these criteria into tests before it writes a single line of the implementation. That, more than anything, is the core of TDD with an agent: you're asking it to prove a behavior. Before implementing addTodo, the agent could write a test like this:

it('does not add a todo when the text is empty', () => {
  const initialCount = store.todos().length

  store.addTodo('')

  expect(store.todos().length).toBe(initialCount)
})

At first this fails, either because addTodo doesn't exist yet or because the current implementation happily adds empty todos. Then the agent writes the smallest amount of code that makes it pass:

addTodo(text: string): void {
  const trimmed = text.trim()

  if (!trimmed) {
    return
  }

  this.todos.update((todos) => [
    ...todos,
    {
      id: crypto.randomUUID(),
      text: trimmed,
      completed: false,
    },
  ])
}

Each behavior becomes testable, each test becomes feedback, and each implementation step becomes smaller than the last. The agent never has to guess what "correct" means, because the tests already define it.

A Definition of Ready and a Definition of Done for AI coding agents

A Definition of Ready tells the agent when it may start, and a Definition of Done tells it when it may stop. In human teams these often get treated as ceremony. With an AI agent they become the two most practical gates you have, because an agent will otherwise start on incomplete context and stop at "some code came out."

Ready, for the add-todo form story, is a checklist worth writing out in full:

  • The scaffold, the todo model, and the state service exist.
  • The add-todo behavior and its validation rules are defined.
  • The location of the input form is known.
  • The testing approach is agreed.
  • Persistence is explicitly out of scope for this story.

That last line matters more than it looks. Leave it out, and the agent may decide to wire up local storage while it's building the form. It sounds helpful, but it quietly mixes responsibilities, and mixed responsibilities mean larger changes that are harder to review and harder to test.

Done is stricter, and it is also a checklist:

  • The form is implemented, and empty or whitespace-only submissions create nothing.
  • Valid submissions create active todos, and the text is trimmed before storage.
  • The input clears after a successful submit.
  • Unit tests cover the logic; component tests cover the interaction.
  • The existing tests still pass.
  • No unrelated files changed, and no new dependencies were added.
  • The agent reports the changed files and the test results.

The last item is the one I refuse to skip. Done means the evidence is attached: the commands that ran, the tests that passed, the files that changed, and any risks or assumptions still open. A confident paragraph without that evidence is a summary of intentions, and intentions are exactly what TDD exists to replace. The tests are one layer of that evidence. The compiler, the linter, and the type checks are the ladder underneath them, and I went through that tooling side in Do UI frameworks still matter when AI writes the code?.

Expected files: keeping the agent in scope

Before the agent implements anything, ask it to name the files it plans to create or touch, and just as usefully, the ones it must leave alone. For the todo app:

  • src/app/todos/todo.model.ts
  • src/app/todos/todo-store.service.ts (and its .spec.ts)
  • src/app/todos/add-todo-form.component.ts (with template and .spec.ts)
  • src/app/todos/todo-list.component.ts (with template and .spec.ts)
  • src/app/todos/todo-storage.service.ts (and its .spec.ts)

And off-limits: src/app/payment/*, src/app/auth/*, src/environments/*, and package.json unless a dependency is explicitly approved.

Now there is a scope boundary. If the agent is implementing "delete todo" and suddenly reaches into app-wide routing, authentication, or package dependencies, the change can be challenged straight away. File scope is a form of architectural control, and agentic development needs it precisely because agents move across files so quickly.

The same discipline applies to dependencies. The agent should name them before it writes any code, and for a todo app the honest answer is short: Angular standalone components, Angular forms, a dedicated state service, and browser local storage in the persistence story only. No backend, no NgRx, no database, no UI component library, no new runtime dependency. Agents reach for tools before they've proven they need them, so making the dependency plan explicit keeps the implementation honest and simple.

Reusable TDD prompts for Claude Code (copy-paste)

These are the three prompts I actually reach for, in the order I use them. First the plan:

Planning mode only. Do not edit files.

Inspect this Angular project and create a delivery plan for a todo app.

Break the deliverable into small composable stories. The first story must be Story 0: Create or verify the Angular application scaffold.

For each story, provide: goal, acceptance criteria, Definition of Ready, Definition of Done, expected files to create or modify, files that should not be touched, dependencies, required tests, risks and assumptions, and suggested implementation order.

The todo app should support adding todos, displaying todos, toggling completion, deleting todos, filtering by all, active, and completed, and persisting todos to local storage.

Do not generate production code yet. Stop after the plan.

Then, once the plan is reviewed, the scaffold on its own:

Implement Story 0 only. Create or verify the Angular scaffold. Confirm the app builds and tests successfully. Create or identify the todo feature folder, but do not implement todo business logic yet. Do not add unnecessary dependencies. Report changed files, commands run, and verification results.

And then each feature story in the same shape:

Implement Story 1 using TDD. First write tests based on the approved acceptance criteria. Run them and confirm they fail because the behavior is not implemented yet. Then implement the smallest code needed to pass. Do not implement UI, persistence, filtering, or deletion in this story. Run the relevant test suite, linting, and type checks. Provide changed files and verification evidence.

The rhythm never changes. One story, one behavior, one test set, one implementation, one verification report. Then you repeat it.

Frequently asked questions

Does this workflow work with OpenAI Codex?

Yes, unchanged. Codex is built as a software engineering agent that works in isolated environments, edits code, runs checks, validates its work, and shows the changed-file diffs,3 and Claude Code is built as an agentic coding tool that understands a codebase, edits files, runs commands, and works through tasks from natural language.4 The workflow lives in the prompts and the gates, not in any one tool, so everything above applies to both, and to whatever agent arrives next.

Is TDD still worth it when the AI writes the tests?

More than before, and the reason is who the tests are for. When I wrote all my own code, tests protected me from my future self. With an agent, the tests written from your acceptance criteria are the only part of the loop that carries your intent into the code without passing through the agent's assumptions. They also make the review tractable: instead of squinting at a diff and hoping, you check that the tests state the right behavior and that the evidence shows them passing.

What about vibe coding?

Vibe coding, prompting your way forward and accepting what comes back, is honestly great for throwaway exploration, prototypes, and finding out what you actually want. The moment the code has to survive real users, teammates, and future changes, it needs a contract, and that is the whole difference. This workflow is what I switch to when a piece of software stops being an experiment.

The takeaway

A chain of control links requirement, stories, tests, code, and verification, producing traceable verification evidence
The chain of control: requirement to stories to tests to code to verification evidence.

The value of all this is the chain of control that runs through it. The requirement becomes stories, and the stories become acceptance criteria. The acceptance criteria become tests, and the tests drive the implementation. Ready controls when the agent can start, and Done controls when it can stop. Expected files control the scope, the dependency plan controls creeping complexity, and the verification evidence controls false confidence.

None of this guarantees perfection, because no workflow does. What it does is reduce the ambiguity, make the work easy to inspect, and give both the agent and me a shared, concrete definition of correctness. A humble todo app shows the whole pattern, and it's the same instinct that keeps a kitchen calm on a busy night: work one ticket at a time, and check every plate before it leaves the pass.

That is the role of TDD in agentic development: a way to turn vague intent into controlled delivery, and to say what "correct" means before the code exists. Agents already generate code faster than any of us can comfortably review it, and that difference is the whole game.

Footnotes

  1. Google Cloud, "Announcing the 2025 DORA Report", 2025, covering the report "State of AI-assisted Software Development" (dora.dev). Based on responses from nearly 5,000 technology professionals worldwide. ↩

  2. GitLab, "GitLab Research Reveals Organizations Are Generating AI Code Faster Than They Can Control It", 2026. States the 92% and 85% findings, from a survey of 1,528 developers and technology buyers across six countries, conducted by The Harris Poll. ↩

  3. OpenAI, "Introducing Codex", 2025. Describes Codex as a cloud-based software engineering agent that runs each task in its own sandbox and proposes pull requests for review. ↩

  4. Anthropic, "Claude Code overview": "an agentic coding tool that reads your codebase, edits files, runs commands, and integrates with your development tools." ↩

Share this articleLinkedInX