AI agents
Let agents write the code. Keep the architecture.
Claude Code, Codex, Cursor and Copilot write code faster than anyone can review it. CAITS gives them a map to read, a shape to copy, and a harness that rejects what doesn't fit, with errors precise enough to fix without you.
Why agents drift, and what holds them
An agent has seen countless codebases, so it knows every way to write yours. Without rules it can't skip, it picks a new one each session. Here is what goes wrong, and what CAITS puts in the way.
Every session invents its own structure.
One shape, checked by 52 rules. The agent copies the shape of an existing context, and anything else fails the check.
Exploring the codebase eats the context window.
Paths that never surprise, and an AGENTS.md of 28 lines: the agent knows where to look before it looks.
The tests pass, and check nothing.
Every mutant must be killed. A test without a real assertion lets a mutant live, and the check fails.
A failing check gets fixed by loosening the check.
AGENTS.md says it: fix the code, not the rule, and don't touch quality/ without asking. A change there shows in the diff you review.
"Done" means it compiles.
Done means npm run check passes: types, lint, architecture, coverage, mutants, end-to-end. Nothing less.
What your agent reads first
init writes an AGENTS.md at the root of every new project: the commands, where the code goes, and the rules, in 28 lines. Many coding agents read it on their own. If yours reads another file, point that file to AGENTS.md. For Claude Code, which reads CLAUDE.md, one line does it:
echo '@AGENTS.md' >> CLAUDE.mdAGENTS.md
# AGENTS.md
A CAITS project (Clean Architecture In TypeScript): a Marko app with a backend and a frontend. A harness checks every change.
## Commands
- npm run check: all the checks at once: types, lint, architecture rules, unit tests (100% coverage), mutation tests (100% score), end-to-end tests. The work is done only when it passes. It prints only the output of the checks that fail.
- npm run tdd: the unit tests, then at each save the tests of the changed files. In an AI agent, Vitest runs them once and stops: run it again after a change.
- npm run lint:architecture: the architecture rules alone, fast.
- npx caits mutation-summary --survivors: the mutants that survived the last mutation run.
## Where the code goes
- The business lives in bounded contexts: src/backend/<context>/ and src/frontend/<context>/, each with domain/, application/, infra/, presentation/ and a di.ts. To add one, copy the shape of an existing one.
- Domain: entities, value objects, errors, domain events, and ports (domain/port/<port>/), each with its contract (<Port>.contract.ts).
- Application: one folder per use case, command/<use-case>/ or query/<use-case>/, with its DTOs, its handler and its test.
- Infra: the adapters of the ports, in infra/<port>/. The test of each adapter runs the contract of its port.
- Presentation: the controllers in the backend; the views and the store in the frontend.
- Dependencies point to the domain. A context imports its own files with #context/, never with ../.
- src/routes/ holds the pages and the HTTP handlers: a page places views, a handler gives the request to a controller.
## Rules
- Write the test first. The coverage and the mutation score stay at 100%.
- A check that fails names the file and the rule: fix the code, not the rule.
- Do not change quality/ (the rules, the thresholds, the configuration of the tools) without asking.
Documentation: https://caits.dev/docs
The loop that keeps it honest
The agent writes, runs npm run check, reads what failed, and fixes it, until everything is green. A check that passes takes one line; a check that fails gives the file and the rule. Each attempt costs a few lines of context, not a wall of logs.
When the agent breaks the architecture
Here, the use case imported an adapter instead of its port. Real CAITS output:
npx caits architecture
src/backend/messages/application/command/add-message/add-message.handler.ts → src/backend/messages/infra/message-repository/file-message-repository.js: application must not depend on infra (layer-direction)
1 architecture violation.When its tests don't really check
For this run, one test was removed from the example app, and Stryker ran on Message.ts only. The coverage stayed at 100%, but no test sent a message of exactly the maximum length any more: Stryker turned > into >=, and every test still passed.
npx caits mutation-summary --survivors
Mutation score: 91.7% (11 killed of 12 valid mutants)
92% src/backend/messages/domain/entity/Message.ts (killed 11, survived 1, no coverage 0)
L16 [survived] EqualityOperator: if ([...trimmed].length > Message.MAX_LENGTH) throw new InvalidMessageError(`A message has → [...trimmed].length >= Message.MAX_LENGTHBoth outputs say what to fix, and where. The agent needs nothing else: no screenshot, no guess, no question for you.
Prompts that work
Short prompts are enough: the shape and the rules are in the project, not in the prompt. Each one ends the same way, because npm run check is the definition of done. They build on the example app, whose context is messages.
Add a context
Add orders next to messages, in the backend and the frontend, with the use case place-order. Follow the shape of messages. Run npm run check until it passes.
Add a use case
In the backend context orders, add the command cancel-order, test first. The domain refuses to cancel an order that was shipped, with an error of its own. Run npm run check until it passes.
Add a port
The context orders needs to send an email when an order is placed. Add the port Mailer to its domain, with its contract, an in-memory adapter for the tests, and an SMTP adapter for production. Run npm run check until it passes.
Kill the survivors
Run npx caits mutation-summary --survivors. For each mutant that survived, write the test that kills it. Change the production code only if the mutant shows code that does nothing. Run npm run check until it passes.
Review less. Trust more.
When the check is green, the questions that eat a review are already answered: the types hold, the files are where they belong, every line is covered and every mutant killed, and the app works in a browser. What's left for you is what only you can judge.
The intent
The names of the tests read like a specification: "the domain refuses an empty text, and the message keeps its text". Read them first.
The domain
The business rules live in small classes, free of any framework. That's where a wrong idea would hide.
The diff of quality/
It should be empty. If the agent touched a rule or a threshold, you see it at once.
npm run check also writes the documentation of the tests, reports/test-docs/index.html: what the app does, one line for each test. A test named “Given …, When …, Then …” shows as three steps.
Works with your agent
The harness is plain npm scripts, and the rules are a Markdown file. Any agent that can edit files and run a command can use them, such as:
- Claude Code
- Codex
- Cursor
- GitHub Copilot
- Gemini CLI
- Aider
- and the next one
Start a project, open it in your agent, and give it the first prompt above.
npx @ingenioz-it/caits init my-app --example