
Most engineering leaders can make a decent case that their foundations are solid, since features ship on time and incidents rarely get out of hand. That case leaves out the checking developers do by hand to make up for thin tests and loose requirements, and nobody counts that checking because it never shows up on a dashboard.
Cassie Shum, VP of Ecosystems, Product Engineering at RelationalAI, argues that AI acceleration has pulled that hidden work out of the loop, and most teams haven't caught on yet. "They show up faster because you removed the shock absorber and didn't notice," she says of the gaps now reaching production. "For years, the developer was the shock absorber."
For a leader, that means the most dangerous gaps are often the best covered, where developers compensated so smoothly that nobody saw a reason to fix anything.
Agents guess where developers used judgment
When validation was weak, the developer checked things by hand, and when acceptance criteria were vague, the developer worked out what the ticket really meant. “That compensating work was invisible, so a lot of teams genuinely believe their foundations are solid,” Shum explains. “What they’ve actually got is a human silently doing the verification that the tests don’t.”
Hand the same work to an agent and that cover disappears. “Take the human out of that loop, and the gap is exposed immediately,” she points out. “Vague acceptance criteria don’t get filled with judgment anymore, they get filled with a guess, and the agent guesses confidently.” Thin test coverage causes the same trouble, because “missing tests used to be backstopped by someone poking at the feature by hand before shipping,” as Shum observes. “That exploratory validation isn’t in the agent’s training, so it writes the code, declares it done and working, and moves on. There’s nothing between that and production except the review you’re now drowning in.”
Shum warns that old weaknesses now reach customers much sooner. “So the same weaknesses that were always there surface faster and harder, as incidents rather than as a slightly slow week.” She sees an upside in the same exposure. “Legacy debt, thin tests, poor docs all get exposed, but agents are also very good at fixing exactly those things once you point them at it.” And she’s clear that the problem predates the tooling. “None of this is really about AI,” she reasons. “It’s about whether you can move at speed safely, which is a twenty-year-old question that AI just made urgent.”
Start by making the invisible work visible. Ask each team lead to write down the checks their developers do by hand before a change ships, like clicking through a flow or taking a second look at an edge case the ticket never mentioned. Each item on that list points to a missing test or an unwritten acceptance criterion, so fix those before you hand more of that part of the codebase to agents.
Onboarding an agent exposes the same gaps
Most leaders wouldn’t think to look for this exposure in onboarding, yet Shum finds it there too. Slow onboarding “was rarely about the person,” Shum says. New engineers worked through wikis of uncertain age and waited on permissions while they pieced together a mental map of a system nobody had written down in a usable form. “A couple of weeks to your first real task was considered fine, because a feature took a week to build anyway,” she recalls. “The onboarding cost was in proportion to everything else.”
“When a feature that took a week now takes half a day, two weeks to first commit is suddenly enormous,” Shum points out. The part that surprised her was how much the two onboarding problems overlap. “The work you do to make an agent productive on a codebase is almost exactly the work that makes a human productive fast,” she shares. “Onboarding an agent and onboarding a person are more alike than people expect, the agent is just faster and more ruthless about exposing what’s missing.”
Her metric changed as a result. “So the metric I care about shifts from time to first commit to time to first commit for an agent,” she explains. “It tells you whether your knowledge is actually discoverable or just theoretically documented.” Teams that make their context structured and queryable in plain language let a new engineer find out who owns a service and why it was built that way in a single pass, without two weeks of asking around. “Weeks instead of quarters isn’t magic,” she adds. “It’s the payoff from finally writing down the things the team always knew but never made legible.”
The number worth adding to your onboarding metrics is time to first commit for an agent. Pick a repository, give an agent a starter task with no extra help, and time how long it takes to produce a working change. Every fix the agent forces on you, a missing README, say, or an owner nobody documented, shortens the path for your next new hire too.
Measure readiness on one small task
Timing a starter task gives a first read, and the fuller version Shum uses with customers follows the work all the way to release. Pick a small, deterministic task from the backlog, something a developer would have finished in under a day, and hand it to an agent with nothing beyond what’s already in the task and the dev environment. “No hand-feeding,” as she puts it. Then carry it to ready for release against your real quality bar rather than stopping at a generated pull request, since “a PR tells you nothing about whether the software works.”
At the end, count two things. Human interventions are the moments a developer had to step in to keep the task on track. Agentic friction covers every time the agent thrashed, repeated the same command five times, couldn’t find a file, or couldn’t isolate a test. “That count is your baseline,” Shum notes.
Shum argues that the check ends up measuring “the friction the team has been quietly absorbing for years.” “Everyone knew it was there,” she continues. “It never got prioritized because a developer always found a way around it.” “The agent doesn’t have that judgment, so it either fails or, worse, succeeds at something subtly wrong, and suddenly the pressure to fix the real thing is finally there,” she underscores.
One run gives you a baseline, and repeated runs tell you which way the team is heading. “The leverage-versus-debt read comes from running it repeatedly,” Shum maintains. “If the intervention count trends down and you’re paying off the friction faster than you produce it, you’re building leverage. If the count stays high, or the agents keep getting bottled up on the same debt, you’re accumulating, and the early productivity bump you felt is about to be erased.”
Repeat the check once a month on a comparable task in each team and log every intervention by cause, so you can fix the most common one before the next round. Put the tally next to your delivery metrics, where the trend will tell leaders far more than any count of merged agent PRs.
Review turned into the bottleneck
The check tends to expose one bottleneck faster than any other, and most teams reach for the wrong fix. “The forcing function is that review flipped from cheap to the bottleneck,” Shum explains. “When coding took three days and review took one, review was a rounding error.” “When coding takes half a day, that same review is now longer than building the feature,” she notes of the new ratio. With more changes flowing through, teams either stall at review or, under pressure to show AI is working, lower the bar and let unvalidated changes reach customers.
The usual response is to have AI suggest review comments, which Shum calls “a band-aid.” “The real move is to stop trying to catch quality at review and shift the work left, so you arrive at review with far higher confidence,” she says. “That means acceptance criteria stop being prose and become something verifiable, expressed as clear goals and concrete examples that an agent and a gate can both check against,” she adds. “It means breaking work into small tasks, each its own PR, small enough to reason about.” She rounds it out with a harness of “verification gates, standards, checks that run before a human ever opens the diff.” Qodo’s Itamar Friedman made a similar case for treating governance as infrastructure, with standards written down where machines can enforce them.
Try it on your next few AI-assisted tickets by writing the acceptance criteria as concrete examples with expected outcomes, the kind a test can check, and by capping PR size so one reviewer can reason about each change in a single sitting. Add gates one at a time, beginning with the check reviewers most often do by hand, and let real review load decide what comes next.
Human review shifts to intent
With the gates carrying the mechanical checks, the reviewer’s job looks different. “Once that’s in place, review changes character,” Shum observes. Reviewers no longer need to read every line to confirm it works, “because the gates already answered the mechanical questions.” That leaves them with the judgment only a person can make, about intent, architecture, and the problem the change was supposed to solve. “That’s the supervisory position, humans on the loop maintaining the harness rather than in the loop rubber-stamping every change,” she reasons. EY’s Dipanjan Sengupta arrived at the same human-on-the-loop model from the governance side, for deciding what agents are allowed to decide.
The change has an up-front cost that Shum doesn’t play down. “It takes discipline, and honestly, it may mean slowing down at first to build the gates,” she admits. “Without it, you’ve just made writing code cheaper and moved the traffic jam one step downstream.”
Write down what reviewers own once the gates pass, and keep it to intent and architecture. Tell the team to expect a slower sprint or two while the gates go in, so the dip reads as part of the plan.
Put architectural intent where agents look
Architecture is the hardest of those to check, since working code can break it without failing a test. “Architectural intent is the why behind the structure, and it’s exactly what ‘working code’ doesn’t capture,” Shum argues. She gives RelationalAI’s own rules as an example, noting that “in our own system, there’s one backend that owns the tables, no new UI goes in the legacy app, services are API-first.” “An agent can generate code that compiles, passes tests, and still drives straight through one of those boundaries, because working and intended are not the same thing,” she warns. “The compiler doesn’t know the boundary exists. The agent doesn’t either, unless you tell it in a way it will actually use.”
Writing those rules into a wiki doesn’t help much. “Intent buried in a wiki is intent the agent will never read and, to be fair, developers stopped reading as well,” Shum cautions. “So the answer is to put it where it gets picked up at the moment of need.” That starts with encoding conventions “as lint and gates so violations fail loudly.” Standards and context go out through skills and MCP, so the agent reads the relevant rule while it works on that part of the system rather than trying to hold one giant document in its head and ignoring most of it. The intent also gets assembled into the agent’s context at pickup “rather than hoping it goes hunting.” “Every one of those moves makes the intent legible to a human too, because a boundary enforced by a gate is clearer to a new engineer than a paragraph nobody opens,” she adds.
Pick the two or three architectural boundaries a new engineer is most likely to break and turn each into a lint rule or CI gate that fails with a message explaining the rule. Then move the related guidance out of the wiki and into the skill or context file that loads when someone, human or agent, picks up work in that area.
Code stays the source of truth
Shum parts ways with one popular school of thought on where intent should live. “This is also where I disagree with the pure spec-driven crowd,” she contends. “I don’t believe the spec is the source of truth that regenerates the code. That might work in greenfield with well-understood requirements, but in reality it doesn’t work in enterprises, where you discover the right answer through implementation and feedback loops with enterprise knowledge.”
Specs, acceptance criteria, gates, and context files earn their place because they help agents and engineers work on the real system with less friction, and in Shum’s view that system is always the code. A team can start on two moves this month. Try the readiness check on one small task and share the tally with your leads, then turn the manual checks your developers already make before shipping into tests and gates, beginning with the most common one.
“Code is still the source of truth,” Shum says.


