Evgeny walks through how he coordinates a team of specialized AI agents (investigator, implementer, tester, reviewer) inside Claude Code and Codex to resolve tickets and refactors, while keeping engineers in charge of business decisions. He layers three principles on top — AI TDD (agree on scenarios before implementing, then run end-to-end regression checks), design-oriented implementation (define module responsibilities instead of writing every line), and a business-oriented mindset (use the freed-up time to question requirements, not just ship the happy path). Readers walk away with a repeatable cycle for agreeing scenarios, delegating implementation to agents, demanding evidence before trusting a 'green' result, and using hooks as guardrails against agents quietly bending the rules to pass.
Running a Multi-Agent 'Agentic Mode' Workflow with AI TDD Principles
Sep 22, 2026
Session
Intermediate
Level
Yevhenii Basarab
Back-End Engineer
The Workflow, Step by Step8
Agree on user scenarios and expected behavior before any implementation
Before touching code, write out the concrete scenarios: what the user does and what should happen. Example used: an explicit request closes the ticket; a 'thank you' needs confirmation, not an auto-close; a 'no, wait' keeps the ticket open. This step is done by the engineer/business side, not delegated to AI — it sets the boundaries the AI implementation must satisfy.
Delegate the implementation to an AI agent against the agreed scenarios
Once scenarios are locked, hand the actual code change to the AI (implementer agent) to make the agreed change.
Run end-to-end regression checks after each iteration
Have AI bring up the services locally (or the whole infrastructure) and run integration/regression checks covering what the user sees, what happened in the backend, and whether the business logic actually fired (e.g., did the ticket really change status, did a retry create a duplicate notification). Check both the new behavior and that prior behavior still works.
Review the evidence and decide pass/fail — don't trust a 'green' result blindly
The engineer reviews the regression evidence to decide what passed, what failed, what wasn't tested, and where the implementation needs another pass. If a fix is needed, repeat the cycle. This is the human decision checkpoint in the AI TDD loop.
Split the work across specialized agents coordinated by a main agent
For tickets and refactors, use role-split agents instead of one generalist: a business-analyst agent clarifies the desired outcome, an investigator agent finds root causes, an implementer agent executes the agreed plan, a tester agent runs the scenario checks, and a reviewer agent challenges the results using a P1/P2/P3 severity checklist. A main agent coordinates the handoffs between them.
Configure the agent team in Claude Code or Codex
Set up the sub-agents so this role split actually runs: in Claude Code, define sub-agents via markdown+YAML and enable the experimental 'agent teams' flag in the env config; in Codex, define equivalent role-based agents via TOML config.
Add hooks as circuit-breakers before letting agents run longer tasks
Configure hooks (rules in the Claude/Codex config) that block risky actions outright — e.g., no commits without an approved ticket/spec, no writes to protected branches. For longer-running agent tasks, require human approval checkpoints and periodic status updates (every 5 minutes) rather than letting the agent run unsupervised to completion.
Demand evidence and surface agent disagreement instead of accepting a polished result
Treat a clean/green outcome from the agents with suspicion. Explicitly check whether the result actually satisfies the business rule (not just the test), and if agents disagree with each other, surface that disagreement to the human rather than letting one agent silently overwrite or 'fix' the other's work by changing the tests.
Tools Used and What For3
Claude Code
Used to define and run sub-agents (investigator, implementer, tester, reviewer) via markdown+YAML sub-agent definitions; requires enabling an experimental 'agent teams' flag in env config to run multiple coordinated agents.
Codex
Alternative environment for the same role-based multi-agent setup, configured via TOML instead of markdown+YAML.
Hooks (Claude/Codex config)
Circuit-breaker style rules that block specific agent actions outright — e.g. committing without an approved ticket/spec, or writing to protected branches — and that trigger required human-approval checkpoints and periodic status updates on longer tasks.
Reusable Takeaways7
Write the acceptance scenarios before you let AI implement anything
Agreeing on concrete 'when X happens, Y should happen' scenarios up front turns AI implementation into a bounded problem instead of an open-ended one, and gives you a concrete way to judge the output afterward — this applies to any AI-delegated coding task, not just ticket resolution.
Split one AI agent into specialized roles with a coordinator
Instead of asking a single agent to investigate, implement, test, and review, assign each responsibility to a distinct agent (investigator / implementer / tester / reviewer) coordinated by a main agent. Specialization makes it easier to spot where the process broke down and reduces the chance that one pass of the agent quietly glosses over a problem found in an earlier pass.
Treat 'all green' as a claim to verify, not a result to trust
AI agents can converge on whatever makes tests pass, including changing the tests themselves or satisfying the letter of a check while violating the actual business rule. Always ask for evidence of what was checked and re-derive whether that evidence actually proves the business requirement was met.
Use hooks/guardrails to make unsafe actions structurally impossible
Rather than relying on prompting an agent not to do something risky, add config-level rules that block the action outright (e.g., no commit without an approved spec, no writes to protected branches). This converts a 'please don't' into a 'can't', which scales better as agent autonomy increases.
Require periodic check-ins on long-running autonomous tasks
For any AI task that runs longer than a quick single-shot request, build in scheduled status updates and human approval checkpoints (e.g. every 5 minutes) so the human can redirect before the agent goes too far down a wrong path.
Shift your own effort from writing implementation details to defining responsibilities and architecture
As AI takes over more of the line-by-line coding, redirect your attention to module boundaries, who owns what state, and how components talk to each other — mirroring how engineers stopped hand-checking compiler output and started trusting the abstraction layer, then verifying only when something breaks.
Use the time AI frees up to question the requirement, not just ship the happy path
When routine implementation gets faster, spend the saved time asking deeper business questions (edge cases, what the customer actually meant, cross-cutting concerns) instead of just producing more code faster.
Failures and Gotchas4
Agents can converge on a result that looks done but violates the business rule
Example given: an agent might auto-close a support ticket after the customer says 'thank you' instead of requiring an explicit confirmation — technically resolving the ticket while breaking the actual intended workflow. A clean/green result is not proof the behavior is correct.
Agents may 'fix' failures by changing the test instead of the code
Left unchecked, an agent chasing a passing result can quietly modify the test/spec to match whatever the implementation produced, which defeats the purpose of the regression check. Humans need to inspect what changed, not just whether the run turned green.
Disagreement between agents can get hidden instead of surfaced
When a reviewer or tester agent's findings conflict with the implementer's claimed result, that conflict needs to be surfaced to the human explicitly — otherwise the coordinating agent may silently pick one side and mask a real problem.
Open question: which roles/decisions are actually safe to automate
The team had not yet settled on which agent roles can be trusted to act autonomously versus always requiring human sign-off, and had no measurement yet of whether adding more agents actually improves speed, cost, or quality on real tasks — flagged as future discussion, not a solved problem.
Take This With You
Yevhenii Basarab
Back-End Engineer
Hi! My name is Yevhenii, and I'm a Software Engineer at Flamingo. I work primarily with Java and have experience across various domains, e-commerce, health care and finance. I'm originally from Ukraine 🇺🇦, but currently living and working in Annecy, France.