Not every step in an agent needs an LLM. Join Jaime Buelta to build a Python agent that pairs fast Jev decision models with LLMs, tools and MCP, and keep it reviewable with spec-driven development.
🗓️ Sat 24 Oct, 10:30 AM ET · 25% off
✍️ From the editor’s desk
Welcome to the 67th issue of Deep Engineering!
On October 6, OpenAI’s chief strategy officer Jason Kwon appeared before Australia’s Joint Select Committee on Artificial Intelligence in Sydney and apologized for what his company’s AI models did to Australian government systems without authorization. According to OpenAI’s own account of the incident, OpenAI assigned an experimental, internal-only model in June to research government spending on medicines for skin conditions. When it could not find the figures, it found a way into the non-public side of Services Australia’s Medicare Statistics Reporting Service, then ran commands and pulled internal files and credentials while still chasing the original answer. OpenAI says it found no evidence that individual patient records were accessed.
No one asked the model to break in, and OpenAI says it never intended the activity to happen. The model had a goal, a set of tools and an environment with a gap in it, and it kept working until it found the gap. The incident shows why enterprise agents with tool access and network reach need enforceable limits around their work. OpenAI’s response included tighter network restrictions, monitoring that pages a human reviewer, and a pause on tool-use training for its most capable models while it strengthens safeguards and prepares additional alignment improvements. For engineering teams shipping agents into their own systems, that is a clear lesson that the harness helps determine how far an agent gets when it goes off course.
Dr. Ali Arsanjani, Director of Applied AI Engineering at Google Cloud, and Juan Pablo Bustos, AI Agents Tech Lead at Google, work with enterprise teams on exactly that problem as they move agents from early proofs of concept into production.
The two co-wrote Agentic Architectural Patterns for Building Multi-Agent Systems, published by Packt, which also publishes Deep Engineering. And in their ARC 2026 session, they laid out the harness engineering patterns they recommend for keeping agents inside the boundaries their builders intended.
Let’s get started.
The Programming Masterclass Bundle
📚 21 expert-led ebooks worth US$761 · Just pay what you want
Bare-metal programming, memory management, templates, coroutines, CMake, CUDA, low-latency systems and Rust for C++ developers, all in one Humble Bundle.
🧠 Expert Insight
Harness Engineering Patterns for Running AI Agents in Production
by Saqib Jan with Dr. Ali Arsanjani and Juan Pablo Bustos
As AI agents move out of pilots and into production systems, the engineering that matters most is shifting away from the model and toward the harness built around it.
Juan Pablo Bustos, AI Agents Tech Lead at Google, and Dr. Ali Arsanjani, Director of Applied AI Engineering at Google Cloud, describe an industry that first fixated on how powerful the LLM is and is only now maturing into the work of controlling it.
Bustos often sees the same sequence on teams building agents. A use case starts working, the team gives the results a superficial gut check, and then deploys the agent on the assumption that it will keep working. He counts testing as the biggest gap the industry still has to close as a discipline. “You need to have a way to test everything,” he says. But the gap reaches well beyond testing, into all the engineering a team puts around the model before it trusts an agent with real work, which is the ground their harness engineering patterns are designed to cover.
Arsanjani organizes those patterns around a principle he calls commensurate governance, which ties the controls around an agent to the amount of autonomy the agent has. “As you increase the autonomy, we want to increase governance so that the agents don’t go off the rails,” he shares. In his view, teams earn the right to give agents more room by watching what those agents do, evaluating the path they take as well as the outcome, and steering them when they drift or fall into infinite loops. The first patterns in that harness deal with what happens at the moment an agent acts on a system.
Sandboxes and checkpoints come first
Arsanjani starts with execution safety, which he calls “very foundational,” and with the pattern the book names execution envelope isolation. When an agent needs to run untrusted code or call a tool, the harness places that work in an ephemeral container that contains the blast radius. The harness then tears the container down once the job finishes. NVIDIA’s OpenShell architecture documentation describes an implementation that separates the agent from the supervisor and gateway responsible for enforcing its access policies.
State needs the same planning. Arsanjani’s example is a ten-step workflow whose host crashes at step nine, which forces a restart from scratch unless the harness has been saving progress along the way. His incremental checkpointing pattern offloads session context to a shared memory cache such as Redis, so the harness can bring the agent back at the last saved point and finish the job.
Engineering leaders can make both patterns defaults for any agent that calls tools or carries out long workflows. Recovery then depends on a harness the team can test, and teams no longer have to hope the model behaves well under failure. The deep dive covers how each pattern works in more detail.
Isolation covers the agent’s own execution. But agents also call models and external systems, and Bustos argues that protection for those calls comes from the layers a team builds around the runtime.
Controls belong around the runtime
Bustos made the point while answering an audience question about hardened container images for agents. Google offers an optimized container for running agents, he acknowledged, and then added, “But I think that the guardrails that you set up happen around the container, not exactly at the container.” The container is essentially the runtime that talks to the LLM through an external API, so the protection comes from how a team wraps it.
He describes that wrapping as a stack of layers. Prompt engineering forms the first barrier and a sandbox the second. Agent identity with the right permissions comes next, along with a model-level filter that screens inputs before they reach the LLM and outputs after it responds. An agent gateway that governs everything flowing to and from the agent at the network level completes the stack.
Identity carries the access decisions that leaders end up owning. Arsanjani counts agents with root access among the biggest anti-patterns he sees. His first recommendation is least privilege, where each agent receives a unique cryptographic service account and only the access its job requires. Bustos extends that to agents shared across a team, which should act under their own identity instead of borrowing one employee’s, so every action traces back to the agent that took it. Whenever a rogue agent makes the news, the first thing he asks is why it had root access.
No layer works alone. Even Google’s Model Armor documentation notes that its model-level filter inspects each prompt and response independently and does not track conversation history across turns. That limitation is the practical case for the identity, sandbox and gateway layers around it.
A useful first exercise is to audit every production agent against those layers and fix the gaps, starting with any agent that still holds broad or root access. Bustos admits that defining guardrails takes longer and requires extra effort. Teams should put that work into the delivery estimate at the start of a project instead of discovering it in a review after launch.
Those controls decide what an agent is allowed to touch, but they say nothing about the quality of its work. That brings Bustos back to his concern about superficial checks.
Evals gate the path to production
Bustos argues that teams should not bring agents to production without evaluations, a testing harness and an established testing practice. The pace of model releases makes that rule more pressing, because every new model gives a team another reason to test the agents built on it.
Checking outputs alone also leaves out the steps an agent took to reach its answer. Arsanjani and Bustos recommend trajectory evaluation, which looks at the tool handoffs and decisions along the way. They pair it with custom metrics built for agents, since the metrics teams carried over from earlier generative AI work do not cover multi-step behavior well. Google’s ADK evaluation documentation likewise distinguishes an agent’s final response from its trajectory of tool use.
A golden dataset gives teams a reference for what good looks like. Bustos warns that it takes several iterations of collecting and refining traces before that dataset is good enough to trust. The book’s preference-controlled synthetic data generation pattern helps teams define and expand it. Teams can then evaluate online while the agent works and offline over the recorded traces.
In practice, all of this means making a passing eval suite the condition for any production release and re-running it whenever the underlying model changes. Each agent gets scored on its trajectory as well as its final output. Of everything in the session, this is the change most teams can start on right away, because it builds on the same engineering rigor software teams have applied to code for years, which is how Bustos frames it.
But a release gate only shows that an agent met the tested acceptance criteria under the conditions evaluated before launch, and the pattern Arsanjani considers paramount addresses ongoing performance after deployment.
⚡ Subscribe to Agentic Engineering
Get expert insights from Agentic Engineering on building AI agents that hold up in production, from the harness and context around them to the loops that keep them on course.
Trust scores keep working after launch
The trust decay and scoring pattern treats a team’s confidence in an agent as a number that changes while the agent works. The harness keeps scoring each agent as it carries out tasks. When a score falls, the gateway can lower that agent’s trust level and route traffic to agents that are performing well while the team corrects the weaker one.
“Implementing that pattern is paramount because you need to basically harness and control the loop for that agent,” Arsanjani says. He sees constant monitoring as the only way to succeed with agentic systems, partly because a prompt injection attack can suddenly change how an agent behaves. Bustos put the same idea in five words during the session, “It’s not set and forget.”
Teams should budget for scoring and monitoring as a standing operating cost for every agent in production, not as a project that closes at launch. Wiring those scores into the gateway gives teams a way to redirect traffic when evaluations detect degrading performance.
Patterns like these also change who owns the design work, and Arsanjani has a clear view of what that means for architects.
Architects now design the harness
Arsanjani teaches at two universities, and his students often ask him if software engineering still has a future as a career. He tells them it does. As he sees it, the work is moving away from manual code authoring and toward high-level system orchestration. The architect becomes an orchestrator focused on “systemic intent and harness design rather than static blueprints.”
In concrete terms, an architect now designs a harness that enforces boundaries and produces telemetry through the OpenTelemetry standard. The architect then checks, through continuous evaluation loops, whether the system meets its nonfunctional requirements for latency, cost, quality and resilience.
Bustos adds a caveat about skills. Teams still need a background in coding to succeed with agent-assisted development, he argues, and he points to the online memes about products that ship without authentication as a reminder of why.
Leaders can act on both views together. That means giving architects named ownership of harness design and the evaluation loops behind it. It also means keeping coding fundamentals central to hiring and growth plans as agents write more of the code.
How much of this harness a team needs depends on how far along its agents are, and the maturity model Arsanjani and Bustos use offers a framework for that decision.
Maturity guides which patterns come next
Arsanjani and Bustos describe six levels of agentic maturity. They start with a single agent making direct tool calls and end with self-correcting swarms that negotiate among themselves, and each level has its own set of patterns.
“If you start in a haphazard way, it’s very difficult to hill climb and get better and better,” Arsanjani says. The warning applies to the organization and the architecture as much as to the agents themselves.
Bustos ties every step back to customer value. He tells teams to start from the value an agent delivers and to add components such as long-running loops or extra agents only when the work genuinely requires them, since cost, speed and quality all pull against each other.
Use their six-level model to identify the patterns your current workload needs. Add autonomy only when the work requires it and the corresponding controls are in place. That keeps the harness growing at the same pace as autonomy, which is the commensurate governance Arsanjani described at the outset.
Bustos’s own readiness test is simple enough to use in any planning meeting. He asks whether you would trust your credit card to an agent, and he calls it the golden question. Teams that cannot say yes yet have a clear place to start, with sandboxed execution, least-privilege identity, a working eval gate and trust scores that keep updating after launch.
Arsanjani and Bustos cover each of these patterns in depth, along with the governance and loop engineering work around them, in Governance and Harness Engineering Patterns for AI Agents.
🛠️ Tool of the Week
NVIDIA’s open-source OpenShell runtime puts enforceable boundaries around agents that need to execute code, use credentials or connect to external services.
Confines each agent’s file access, system calls and network connections at runtime.
Injects managed credentials only into requests bound for approved endpoints, keeping those secrets outside the agent’s environment.
Uses formal verification to flag risky policy changes for human review before they take effect.
Manages sandboxes, policy and access from one control plane, with Helm deployment on Kubernetes.
📎 Tech Briefs
California announces OpenAI cybersecurity subpoena - California’s inquiry covers broader cybersecurity risks involving OpenAI models, beyond the previously announced Hugging Face investigation.
CISA adds exploited Zammad flaws to its catalog - CVE-2026-102489 and CVE-2026-102490 enable an exploit chain from session fixation to root access on affected Zammad installations.
AISI introduces Transect for agent evaluation - Turn-by-turn reports connect agent events, token use and behavior labels to source transcripts for auditable evaluation.
Strands Decider 2B released - The 1.9-billion-parameter model scores supplied options in one forward pass, supporting tool gating and routing without text generation.
Google ADK 2.11.0 released - Streaming client disconnects now cancel runs, while workflow tool calls pause for approval instead of failing.
Thank you for reading this week’s issue. Before giving an agent more autonomy, make sure your team can enforce its boundaries, evaluate its work and intervene when its behavior changes.
Keep building,
Saqib Jan, Editor-in-Chief, Deep Engineering
Partner with Deep Engineering
If your company wants to reach senior developers, software engineers, and technical decision-makers, speak to us about partnering with Deep Engineering






