My complete AI agent workflow, from problem to shipped software

Hashan Wickramasinghe17 min read
A simplified flow of the AI agent workflow in flat vector style: problem, think, brief, build, check, gate, ship, all resting on a bar labeled the library

Seepient, the open source AI assistant I build and run a business on, failed one evening while I was nowhere near a desk. A model ended its turn without producing anything, the assistant came back with an empty answer, and every layer around it treated that empty answer as a success. Nothing crashed. No error reached a log. The workday was long over, which is exactly when a one-person company does its real building, and it gave the hole all the quiet it needed.

That is the solopreneur math, and it never balances. You are the CEO who wins the work, the engineer who builds it, and the QA department that is supposed to catch exactly this kind of failure. When an AI tool writes the code, the QA department quietly becomes the machine that wrote it.

The numbers say this is now everyone's math, not just mine. Stack Overflow's 2025 survey found 84 percent of developers use AI tools or plan to. Veracode's 2026 report found AI code passes security checks about 56 percent of the time, a rate that has not moved across four years of model releases. And METR's 2025 trial found experienced developers were 19 percent slower with AI tools while believing they were about 20 percent faster. That last one is the dangerous number, because the person who wants the speedup is the same person keeping the score.

The opening failure is real, and it comes from Seepient, the open source AI assistant I build and run my own business on. Seepient is where every idea in this workflow earns its keep: the plan reviews, the brief with right answers written first, the four checkers, the gate. This article walks through the full machine I now run, from the moment a problem walks in to the moment the software ships. It is the working summary of my book, The One-Person Engineering Team, which is free to read on this site. I want you to read the book, so I am giving you the whole machine here, and the book holds the wiring diagrams.

Why "just review it" stopped working

The obvious fix is to review the code after the AI writes it. I tried. Two problems killed it.

First, a second AI agreeing with the first is agreement, not correctness. Research on AI judges, including Panickssery, Bowman and Feng, 2024, shows models recognize and favor their own generations. Ask the machine that wrote the code whether the code is good and you get a confident yes.

Second, a human reviewing AI code at volume is rubber-stamping. Plausible code is hard to review. It looks clean, it compiles, its tests pass, and the flaw sits in the path nobody exercised, the way Seepient thanked its user with an empty answer on a night nobody was watching.

So verification cannot be a step at the end. It has to be a system, and the system stands on one rule.

The one law underneath everything

Nobody grades their own exam. Whoever did the work never signs off on it.

Audit firms have run on this idea, called separation of duties, for a century: the person holding the pen never also holds the eraser. With a human team, a checker for every doer means salaries doubled. With AI workers, the second hire costs the same as the first, which is roughly nothing, and onboards in thirty seconds. There is no budget argument left.

Here is the whole machine on one map. The rest of the article walks it stage by stage.

Stage 1: think first, no code allowed

A problem walks in. A client email, an idea over coffee, a bug report. In a one-person company the next move is always the same temptation: start typing. The editor is open, the model is fast, and typing feels like work.

Two failures wait on either side of that temptation. Type immediately and you build the first idea rather than the best one, because the first idea is the only one that got a hearing. Plan forever and the document gets refined instead of released. The workflow takes the middle path: an afternoon of paper before any code.

The brainstorm. I open a fresh conversation with a good model and run it like a whiteboard. The rule is that no code and no decisions come out of it. I interrogate the problem, consider at least one alternative, and let the model argue back. This replaces the colleague I do not have; a solo founder arguing with herself is having a fixed fight, because both sides share one head and one incentive to be done by lunch. I end every brainstorm by asking the model to state the problem back to me in one paragraph, plus the two alternatives we rejected and why. If that paragraph is wrong, the thinking is not done. I keep the transcript.

The plan and the to-do list. Next, a planning step turns the brainstorm into a plan: what gets built, how, and where it fits in what I already have. Then the plan breaks into an ordered to-do list, and the ordering is the test. If I cannot write the tasks in order, the plan was still vague, because an order that hesitates is hiding an assumption. Better to surface that on paper, where it costs a sentence, than in code, where it costs a week.

The paper gets graded too. Before anything else, fresh AI reviewers attack the plan itself. Same law as everywhere else: the conversation that wrote the plan never reviews it. Two rounds maximum, then I stop and rethink, because a third round usually means the premise is wrong. A wrong plan caught on paper costs minutes. Caught in production, it costs the Seepient failure all over again, this time with a paying user on the other end.

The brief: what done means. This is the document most people skip and the one that does the most work. A prompt says what you want tried. A brief defines what done means, with the right answers written down before anyone builds. It is a numbered promise list, and each promise carries an independent right answer, an oracle I can check the result against later. "The total must come to 1,204.30" is a promise, because I computed it by hand from the ledger and wrote the number down first. "The total should be right" is a vibe.

A slice of a real brief looks like this:

IDThe promiseThe right answer, written first
A1Each active customer gets one statement PDF3 customers in the test ledger, so 3 PDFs
A2The total due is correct1,204.30, computed by hand from the ledger
A3An amount with a letter in it is rejected loudlyRow 14 of the test ledger, amount "12O" (letter O), must be refused with a named error
A4When the payment provider is down, nothing is inventedSimulated outage must show "statement unavailable", never a guessed number

Notice A3 and A4. The brief invites the ugly inputs in on purpose: the ambiguous case, the invalid case, the failed case. The builder builds exactly what I describe, and every path I leave out becomes a silent surprise for a customer.

Writing the brief means doing the thinking the prompt was going to skip. That thinking is my job. Prompting let me pretend otherwise for about a month.

Full detail: chapter 5 and chapter 6 of the book.

Stage 2: the factory floor

Paper done, the code gets built, and the build is where most AI workflows quietly die, because everybody reviews the desk while the builder is still working on it. I know, because I ran it that way first. The builder edits six files, I kick off a review, the builder keeps working, and every approval that comes back describes a version of the change that no longer exists. I call that review theater: the form of a review with none of the protection.

The Builder. One role, and the only one allowed to touch the code. It reads the library first, the standing documents I will describe below, then follows the to-do list and nothing else.

The seal. Before any review starts, the pipeline photocopies the exact work and locks the copy away. Every checker judges only the copy. If the original moves while reviews run, the machine notices at the end and throws the whole review out. Every verdict is bound to one frozen snapshot, so an approval can never describe a change that no longer exists. Seepient's own release runs through this exact gate; when a fix went in for the silent empty answers, the proof was bound to one frozen copy that the gate had judged.

Machines check first. Code tidiness, type checks, and the test suites run automatically, and the referee software records every result itself. Agents never report machine results, because a robot writing "tests passed" is not evidence.

Four checkers at once. Each one is a brand-new conversation that knows nothing about the Builder, and each judges only the frozen copy. They run in parallel, in different roles with written job descriptions.

  • The Code Reader reads the frozen code for quality and correctness, and how far one change can reach into the rest of the system. It never reads the Builder's story, so it cannot inherit the Builder's blind spots.
  • The Stand-In Customer uses the real product like a first-time user and never sees the code. It grades against the brief. Beautiful code aimed at the wrong feature gets caught here.
  • The Auditor reads across the whole estate, hunting silent ripple damage: the copy of a function in another project that now quietly disagrees, the consumer of an interface nobody told about the change.
  • The Burglar is paid to break in. Its job description is one line: find flaws, leaks, holes, mistakes, and oversights. It succeeds exactly where polite review forgives.

Five reviews in sequence cost a morning. Five in parallel cost a coffee. And depth follows risk: a copy change wakes two reviewers, a change that touches money or customer data wakes the whole fleet.

Sort the complaints. Four reviewers produce signal and noise, and the noise reads like signal. A merge step deduplicates the reports, so a defect found twice becomes one finding with two witnesses, and enforces the rule that decides everything: a complaint only counts if someone can show it. A generic "looks fine overall" cancels nothing. If the Burglar demonstrated a hole, nobody's clean impression elsewhere un-demonstrates it.

The Architect looks at the whole map. One more fresh conversation reads the sealed change against the entire system. Do the connections still line up? Does the change still serve what the product is for? The fleet's defects are local. The expensive ones rarely are. A feature can pass every local check and still break the product's spine.

The gate that can't be bribed. Then everything meets the referee, and this is the part of the whole workflow I care about most. The referee is plain code, not a chat. It takes the machine results it recorded itself, the sorted findings, the Architect's notes, and the promise list from the brief, and asks one question per promise: does proof exist for this?

The verdict is computed, not judged. AI can fail a change; it can never wave a broken test through, because waving things through is not a power the code grants. Three outcomes:

What counts as proof is strict, because the gate only accepts evidence it can re-run: a test run the referee itself recorded, a real click-through of the product, data saved and then read back, a number worked out by hand first. What never counts: a robot writing "tests passed", a screenshot nobody captured, the Builder's confidence, a second robot agreeing.

Full detail: chapters 3, 7, 8, and 9 of the book.

Stage 3: out the door

Green is not shipped yet. Three more stations stand between a pass and a customer.

The receipt. A pass earns a signed record of every check that ran and what each one found, sealed to the exact frozen copy it judged. Change one letter of the code afterwards and the receipt describes something it never judged, so it is void. When a client asks how I know the release is safe, I show them the file. When I ask myself the same question at 11pm, the answer is also the file.

The finisher. A release role takes over: bumps the version, writes the changelog, updates the docs so they stop lying about the feature. Its own edits get checked like anyone else's, because nobody grades their own exam, including the role whose whole job is finishing.

The second gate. When the code arrives at the hosting pipeline, the pipeline re-verifies that the code arriving is exactly the code the receipt describes, then runs the full suites again. Missing proof means stopped, never a quiet pass.

The first minutes in the wild. After the deploy, a smoke check watches the live system for a fixed window, looking for the obvious to break. Finding a break in the first minutes costs an apology. Finding it next quarter costs a customer.

Then it is done, in the only sense I trust now: shipped, and I can prove it.

The library everything sits on

None of this works on a blank desk, which I learned the hard way. Early plans kept inventing a second way of doing something the system already did well, because nothing told the agents where code goes and which parts may talk to which.

So four standing documents sit under the whole workflow: what the product is for, how it should look, how the pieces are wired, and the house rules. Every worker reads its reading list before starting, in every stage, including the reviewers. The documents are the desk the factory stands on, and they are also why the plan step cannot wander far: the architecture document already settled where things belong.

Full detail: chapter 4.

What this buys a one-person business

The honest ledger, because I keep one now.

What it replaces: the QA department a solo business could never afford. What it does not replace: deciding what the business needs. The brief made that explicit. The machine can verify that the total is 1,204.30; only you can decide the statement was worth sending.

What it costs: the first guards take an afternoon. The middle of the ladder, the seal, the parallel checkers, the merge step, takes a weekend. The full referee took me weeks, and the book includes a build-versus-buy path if you would rather put those questions to a vendor than write the code.

What it earns: you can finally promise a client speed with receipts. Not "trust me, the AI was fast", but here is the record of every check that ran against the exact code you are running. Some clients will never care. The ones who have lived through an assistant failing in silence, on a night nobody was watching, care a great deal.

Full detail: chapter 10 for the build ladder, chapter 13 for the business case.

Questions people ask

How do you quality-check AI-generated code? Separate the writer from the checkers. The AI that wrote the code never grades it. Fresh AI reviewers in different roles judge a frozen copy of the change, machine checks run first and get recorded by the tooling itself, and a deterministic program, not another AI, computes the final verdict from recorded evidence.

Can AI agents review each other's work reliably? Only with structure. On its own, a model reviewing another model's work tends toward agreement, and models favor their own outputs. Reliability comes from the surrounding system: separate conversations per role, written job descriptions, a frozen snapshot so nobody reviews a moving target, and a rule that a complaint only counts when someone can demonstrate it.

What is a deterministic referee in AI development? Plain software, not a chat, that computes the ship-or-block decision. It counts proof: for each promise in the brief, does valid recorded evidence exist? Because the verdict comes from code, a convincing AI cannot wave a broken test through. Judgment proposes; arithmetic disposes.

What is a good workflow for shipping software with AI agents? Think first with no code allowed: brainstorm, plan, ordered tasks, reviewers on the plan itself, then a brief defining done with right answers written in advance. Build with one role allowed to touch code, freeze an exact copy, run machine checks plus several independent AI reviewers in parallel, merge and verify their findings, then let a deterministic gate decide. Ship with a receipt, re-check in the hosting pipeline, and watch the first minutes live.

Can a solo founder really run QA like a real team? Yes, because the economics changed. Separation of duties used to mean a checker's salary for every doer. AI workers onboard in thirty seconds and cost roughly nothing, so a one-person business can now afford more checkers than builders. The constraint is not budget anymore; it is whether you design the system at all.

Read the book

This article is the map. The territory is The One-Person Engineering Team, free to read on this site, thirteen chapters, about two and a half hours, and there is a PDF download if you read on paper. The manuscript is still being polished and the web edition updates as it does.

Start where your wound is. If you want the case in full numbers, start at chapter 1. If you are standing a new project up this month, start at chapter 4, the library. If the hole has already shipped, start at chapter 3 and hire the team that catches the next one.

Filed under: StrategyAI AgentsQualityWorkflow

Share this post