How to Review What AI Agents Build: A PM's Acceptance Checklist
Building got cheap; accepting the result did not. The four-rung Acceptance Ladder helps PMs review agent-built changes against the decision, the brief, the edge cases, and the outcome.
Share
AI coding agents have made one part of product work cheap: producing a working diff. The harder question is what happens after generation: who decides whether the change is correct, within scope, and safe to ship? Anthropic's 2026 Agentic Coding Trends Report describes the shift from writing code toward orchestrating and supervising agents, while McKinsey's recent analysis argues that faster development does not automatically produce better product outcomes.
Here's our answer. The PM job is moving from writing the ticket to deciding whether the result is acceptable. Most teams haven't built a process for that second job. They review agent output the way they reviewed a human pull request: skim the diff, click around the preview, say "looks good." That works when a teammate carries context from the planning meeting and customer conversations. It fails when an agent carries only what the brief encoded.
This post gives you a concrete method: the Acceptance Ladder, four rungs a PM climbs before an agent-built change counts as done. It's the same sequence we follow when we hand work to coding agents while building Specky.
Why "looks good" is the wrong bar for agent output
A human engineer's pull request carries context you can't see in the diff: they sat in the planning meeting, they remember the customer call, they know which edge case Support complained about. An agent carries only what was written down. If the spec was vague, the agent fills the gap with the most statistically plausible behavior, and plausible isn't the same as correct.
That produces a specific failure pattern. The change compiles, the preview renders, the happy path works, and the PM approves it. Two weeks later someone notices it quietly dropped a constraint nobody wrote down, or solved an adjacent problem instead of the one in the brief. Nothing was wrong with the code. The acceptance step simply never asked the right questions.
Three things make agent output harder to judge than human output:
S
Specky Team
Writing about AI-native product development at Specky.
Volume. Cheap generation means more candidate changes per week, so review time becomes the bottleneck.
Confidence. Agents describe their own work fluently. A tidy summary isn't evidence.
Missing intent trail. Without a link from the change back to the decision it implements, you can't tell "built as specified" from "built as inferred."
The Acceptance Ladder is designed around those three problems.
The Acceptance Ladder
Each rung asks one question. You only climb to the next rung when the current one passes. If a change stalls on rung one, you don't spend time on rung four.
Rung 1: Traceability — which decision is this implementing?
Before reading any code, answer one question: what decision or evidence does this change trace back to?
A change should point to a specific item: a decision record, a customer insight, an opportunity, an experiment. If you can't name it in one sentence, stop. Either the work was never justified, or the justification lives in someone's head and the agent never saw it.
In Specky, this is what the product graph is for. Decisions, insights, and tickets are linked nodes, so an agent run can pull the surrounding context before it plans anything. The reviewer's first check is simply whether that link exists and whether the cited evidence says what the change claims it says.
Pass condition: one decision or insight, named, with its evidence reachable in two clicks.
Rung 2: Contract — did it build what the brief said, and only that?
Now compare the change to the written brief, line by line. Not "does it feel right," but a literal diff between acceptance criteria and observed behavior.
Do two passes:
Coverage pass. For every acceptance criterion in the brief, find the behavior that satisfies it. Mark each as met, partially met, or missing.
Surplus pass. List behaviors the change introduces that no criterion asked for. Agents often add helpful extras: a new setting, an extra column, a fallback path. Each one is unreviewed scope. Keep it or cut it deliberately.
The surplus pass is the one teams skip, and it's where agent work differs most from human work. A person tends to under-build against a brief. An agent tends to over-build.
Pass condition: every criterion marked met, every surplus item explicitly kept or removed.
Rung 3: Edges — what happens when the input is wrong, empty, or hostile?
Agents are strong on the happy path because that's what prompts describe. Rung three goes where prompts rarely do. Write down five situations and try each:
Empty state: no data yet.
Error state: the upstream call fails.
Boundary: the largest and smallest realistic input.
Permission: a user who shouldn't see this tries to.
Repeat: the action runs twice in a row.
You don't need exhaustive QA. You need the five situations a real customer will hit in the first week. If the brief didn't specify behavior for them, that's a brief defect, and the fix goes into the brief, not just the code.
Pass condition: all five situations behave in a way you'd defend to a customer.
Rung 4: Evidence of outcome — how will we know it worked?
A change can pass the first three rungs and still be useless. The last rung asks what you'll measure and when you'll look.
Write the check before shipping: the metric or signal, the expected direction, and a date. "We'll see how it goes" isn't a check. This is also where hype gets filtered out: a feature that ships fast but has no outcome check is a guess with a deploy date.
Pass condition: a named signal, an expected direction, and a calendar date to read it.
A worked example
Suppose your team asked an agent to add a "bulk archive" action to a feedback inbox.
Rung 1. The brief links to an insight: support spends time clearing stale items, and three customer conversations mention it. The evidence is reachable. Pass.
Rung 2. The coverage pass finds all four criteria met. The surplus pass finds the agent also added an "undo" toast and a new "archived" filter tab. The undo is worth keeping and gets a criterion retroactively. The filter tab changes navigation and wasn't asked for, so it's cut and filed as a separate opportunity. Pass after one edit.
Rung 3. Empty state shows a blank panel with no explanation. Fails. The fix goes into the brief as a missing criterion, then back to the agent.
Rung 4. The signal is median time to clear the inbox, expected to drop, read in two weeks.
The review took one fail and two edits, caught before release. A "looks good" review would have shipped the blank panel and the unrequested navigation tab together.
Make the ladder automatic where you can
The ladder works best when part of it isn't manual. Two practical moves:
Validate the plan before the agent writes code. Specky exposes this through its MCP tools: an agent calls validate_agent_plan against your engineering standards before it starts, and validate_agent_change before it reports success. That moves rung two earlier, so the plan gets checked against the rules while it's still cheap to change.
Keep the evidence attached. When the decision, the brief, the agent run, and the outcome check are all nodes in one graph, rung one stops being a hunt. The reviewer opens the change and the reason is already beside it.
Running the ladder in a small team
You don't need a review board. For a solo founder or a three-to-ten person team, the ladder is a five-minute habit per change, and the order matters more than the ceremony.
Start by writing acceptance criteria as short, checkable sentences in the brief itself, such as "archived items disappear from the inbox list" rather than "archiving works well." Sentences like that are what make the coverage pass mechanical instead of a matter of taste. Next, keep a standing list of your five edge situations so rung three isn't reinvented each time. Finally, make the outcome check part of the definition of done: if the change has no signal and no date, it isn't ready to merge.
When a change fails a rung, resist the urge to patch it in the review thread. Send it back with the missing criterion added to the brief. The brief is the durable artifact; the thread disappears. Over a month, your briefs get sharper and your failure rate on rungs two and three drops, which is the real payoff.
What to measure on your own review process
Treat the ladder as a process you can inspect:
Rung-fail distribution. If most failures land on rung two, your briefs are the problem. If they land on rung three, your prompts skip edge cases.
Surplus rate. How often does an agent add unrequested scope? A rising rate means your briefs leave too much room.
Outcome-check completion. How many shipped changes got their rung-four read on the date you set? This is the number that shows whether you're learning or only shipping.
You don't need a dashboard for this. A table with one row per reviewed change is enough for the first month.
The takeaway
Faster building didn't remove the hard part of product work. It moved the scarce skill to the review seat. Climb the four rungs in order: trace the decision, check the contract, probe the edges, define the outcome. A change that clears all four is genuinely done; one that doesn't is just a diff that compiles.
Specky is the AI product workspace that shows its work — every agent-built change stays linked to the decision, evidence, and outcome check behind it. → specky.space