← All posts

Agentic SDLC in practice: where AI agents speed up a software team, and where they don’t

Most teams already use AI in software development, usually as autocomplete in the editor. Agentic SDLC is a different thing. The agent does not suggest the next line. It takes a task, plans it, changes the code, runs the tests, reads the errors and tries again, and hands a finished pull request to a person who decides whether it goes in.

I work this way every day with Claude Code and Codex, and I use agent frameworks such as OpenClaw and Hermes for work outside the code. Below is what I have learned about where this pays off, where it does not, and how to start without putting a whole team at risk.

What “agentic” actually changes

In a classic software lifecycle every stage waits for a person: someone writes the ticket, someone codes, someone writes tests, someone reviews. In an agentic lifecycle agents do the first pass of each stage, and people move to the gates between stages. The job of the team shifts from typing code to defining what “done” means and checking that it was met.

  • Analysis: an agent turns a rough request into acceptance criteria and lists open questions.
  • Implementation: an agent writes the change in a branch and keeps going until the tests pass.
  • Testing: an agent adds missing tests and reproduces reported bugs as failing tests first.
  • Review: an agent does the first review pass, a person does the final one.
  • Release: an agent prepares release notes and migration steps, a person presses the button.

Where agents are already good

  • Well-defined changes in code that has tests: new endpoints, fields, validations, small features.
  • Tests, especially for existing code that nobody wanted to cover.
  • Refactors and migrations that repeat the same pattern across many files.
  • Documentation, changelogs and onboarding notes kept in sync with the code.
  • Triage: grouping bug reports, finding duplicates, pointing to the likely module.

The common factor is simple: there is a fast, objective way to check the result. If tests or types can say “this is wrong”, the agent can fix it on its own.

Where they still fail

  • Unclear requirements. An agent will build exactly what was written, including the parts that were wrong.
  • Architecture decisions and trade-offs that depend on business context nobody wrote down.
  • Security-sensitive code: authentication, payments, permissions. Here the agent can draft, but a person must own it.
  • Legacy code without tests. Without a way to verify, the agent is guessing, and so are you.
  • Anything where “looks right” is the only check.

Guardrails that make it safe

  • An executable definition of done. Tests, types and linters are the agent’s real manager.
  • Small pull requests. One task, one branch, one reviewable change.
  • A sandbox. Agents work in a branch and a disposable environment, never with production credentials.
  • Human gates. Merge and release stay with people.
  • Traceability. Every agent change links to its task and keeps its log, so you can see why something happened.
  • Cost control. Route by difficulty: a frontier model for the hard reasoning, a smaller open model such as Qwen or GLM for bulk work like tests and docs.

A 30-day pilot on one team

  • Week 1: pick one repository with decent tests and one type of task. Measure today’s baseline: lead time from ticket to merge, review time, bugs found after release.
  • Week 2: set up the agent workflow, the sandbox and the gates. Start with tests and small changes only.
  • Week 3: run it on real tickets. Keep a simple log of what the agent got right, what needed fixing and what it could not do.
  • Week 4: compare with the baseline and decide what to scale, what to change and what to drop.

Measure the same three things before and after: lead time, review load and escaped defects. If lead time drops but escaped defects rise, you have only moved the work to your users.

The short version

Agentic SDLC is not about replacing developers. It is about moving people from typing to deciding, and letting agents do the repetitive first pass under clear rules. Teams that define “done” well get the most from it. Teams that cannot say what “done” means get faster at producing the wrong thing.

If you want to test it on your team, a 20-minute call is enough to pick the first process and the right pilot. Book a call.

Next step

Want this working in your team?

20 minutes is enough to pick the first process and a sensible pilot.