This deep dive is based on the ARC 2026 session led by Packt authors Dr. Ali Arsanjani and Juan Pablo Bustos, and has been edited into a first-person account in their own words.
Over the past year, as we compiled the patterns for our book, we kept arriving at the same realization. The industry has focused almost entirely on the engine itself, the LLM, and on how capable that engine is becoming. That focus made sense early on.
But the field is maturing now, and maturity shows you that a capable engine needs a set of disciplines around it before anyone can depend on it. Juan likes to compare it to a rocket ship, which nobody can hold without a launcher. This piece covers those disciplines and the patterns behind each one, starting with the maturity model that tells you which patterns you need next.
Agentic systems mature in six levels
We think about agentic systems across six levels of maturity.
Level one covers basic agentic systems, a single agent making direct function calls or tool calls.
Level two is the dynamic single agent, where one LLM decides at runtime which tool to call.
Level three adds introspection through the ReAct pattern or the reflection pattern, along with retries.
The first three levels all rely on one LLM, working through function calls or, at best, repeated rounds of introspection and reflection.
Level four is the first true multi-agent system. It is hierarchical, with one orchestrator directing several collaborative agents from the top down, and it brings governance and auditability into the design.
Level five is advanced multi-agent coordination, where a meta agent resolves conflicts and drives consensus. The complex systems that banks and healthcare providers build today need this level.
Level six is the self-correcting, self-learning, adaptive multi-agent system. Imagine a swarm of agents that reaches consensus without a central coordinator, with agents negotiating directly with each other.
Patterns are what move a system from one level to the next. Juan finds that the most important conversation with customers is about defining that journey, starting with how a system begins to deliver value. As teams adopt these patterns, they learn what observability on an agent actually means and why it matters. They then use what each level teaches them to reach the next one. The ultimate outcome is the goal, and the path to it always starts with defining the journey.
Three disciplines tame the engine
Juan frames the work of taming the engine as three core disciplines.
Governance is the seat belts.
Harness engineering is the chassis.
Loop engineering is the navigation loop that keeps the vehicle on course.
Governance has to grow with autonomy. As agents become more autonomous, the governance around them has to increase in scope and capability so they do not go off the rails. The only way to place more trust in what agents do is to observe them and to evaluate both the path they take to execution and their final outcomes. That is also how you learn to steer them when they drift or fall into infinite loops. We call this commensurate governance, meaning governance that keeps pace with every increase in autonomy.
Juan sees the opposite habit often. A team defines a use case, builds an agent, releases it into the field and expects it to work every time. Our governance patterns start from the other direction, with the ability to see what the agent actually did.
Governance starts with seeing what agents did
Governance for agents comes down to two things, knowing what each agent did and controlling what it is allowed to do. Instruction fidelity auditing covers the first, tracking whether the original intent survives as it passes between agents and leaving a trace you can inspect afterward. Agent identity and least privilege cover the second, giving every agent its own identity and only the access its job needs, which keeps the blast radius small when an agent goes wrong.
Instruction fidelity auditing
One of the governance patterns in our book is instruction fidelity auditing. In a multi-agent system, you give an instruction at the top. As that instruction passes from one agent to the next, it grows more muddled in the context, especially as the context window grows.
Instruction fidelity is the ability of the original intent to travel from the orchestrator down through each set of agents and come back as the action you asked for.
Auditing is the ability to see what each agent actually did along the way.
In practice, every agent has to leave a log trail behind it, and the industry standard for that trail is OpenTelemetry. At Google, where we both work, our Agent Development Kit (ADK) has native support for emitting OpenTelemetry traces from your agents, which gives you that visibility from the start.
Agent identity and least privilege
Agent authentication and authorization become essential as you move up the maturity levels. One of the biggest anti-patterns we see is agents granted root access. Every agent should receive only the least privilege it needs to get its job done. Assigning each agent a unique cryptographic service account ID creates a strict, auditable boundary around what it can touch.
Agent identity also opens up use cases that did not exist before. Most agents today interact with you directly and act on external systems on your behalf. Now imagine an agent with its own access to systems that are not necessarily yours, working as a team member that several people use. Its actions should not be tied to any one person’s identity. They should be tied to the agent.
Our patterns for agent authentication and authorization treat the agent as an entity in its own right, with its own service account, and at Google we have a product called Agent Identity that does exactly this. The most important question to start asking is which emerging use cases your identity model has to accommodate, because that is where the value of agents comes from.
Identity, authentication and authorization together limit the blast radius. Everyone has followed the news about some agent that went rogue, and the question that matters is why that agent had root access at all. We have both set up OpenClaw on our personal, non-corporate laptops. It was instructive to experience firsthand how you limit an agent’s blast radius and still keep it useful.
Harness engineering builds the chassis
With the seat belts in place, the next job is building a harness that constrains what agents do. The chassis is about execution safety, and everything else depends on it.
Execution envelope isolation
The execution envelope isolation pattern, which we also call the sandbox pattern, says that agents need a secure, isolated boundary to operate within. When an agent executes untrusted code or calls a tool, the harness has to execute that work inside a container that limits the blast radius. The container is ephemeral, and the harness dismantles it once the execution finishes.
Filtering at the model and at the network
Isolation handles local execution, but agents also need to interact with external systems. Juan thinks about that at two levels.
The model level. You filter inputs before they reach the LLM and filter outputs after it responds. On Google Cloud, our product for this is Model Armor.
The network level. A gateway gives you one point where you govern everything that flows from the agent and to the agent. Our product for this is Agent Gateway.
The pattern matters more than any product. With both levels in place, you cover access to the model and access to external systems.
Incremental checkpointing
A solid chassis also has to handle state, which maps to our incremental checkpointing pattern. Say you have a ten-step workflow and the host crashes on the ninth step, just before the workflow concludes. You should not have to restart everything from scratch.
Incremental checkpointing offloads the session context to a shared memory cache, which could be Redis or anything else that fits your stack. The harness can then resurrect the agent so it resumes exactly from the last saved point.
Loop engineering learns from real executions
Once governance and the harness work together, you need a way to drive continuous improvement. That brings us to a core architectural discipline people now call loop engineering. It is the practice of designing feedback loops that let an agentic system learn from its real-world executions.
Loop engineering gives you two things.
Control over outputs. This matters because models sometimes report unfinished work as finished.
Self-improvement. This is one of the most valuable patterns we see emerging from these practices.
Gut checks are not evaluations
The biggest gap the industry has to close as a discipline is testing. It is very common for a use case to start working and for the team to settle for what Juan calls a vibe test of the results, a superficial gut check that the agent seems to work. Agents should not reach production without evals, without a testing harness and without a practice around both.
The new discipline of software engineering, and of agent engineering with it, depends on rapid iteration. New models arrive constantly, with announcements yesterday, last week and the week before, so you need a way to test everything.
Custom evaluation metrics and trajectory evaluation
Much of that comes down to what our book calls custom evaluation metrics, which means deciding which metrics you will use to evaluate your agents. Traditional metrics are not enough, as we already saw with the metrics used to evaluate earlier generative AI systems.
Agents need agent observability. You need to identify the tool handoffs and see into the decision-making process, and that is trajectory evaluation. Juan sees this as a practice more than a feature, one that carries into agent engineering the technical rigor software engineering has applied for years.
Trust decay and scoring
As you gather trajectories and the metrics inside them, you can manage the health of the system dynamically as it evolves. We call this the trust decay and scoring pattern.
You start with a specific level of confidence in an agent.
You keep scoring the agent as it carries out its tasks.
If the score drops, the gateway dynamically lowers that agent’s trust level.
The gateway routes traffic away from agents or nodes that are performing poorly.
It then either corrects them or sends the work to an agent that is performing well.
Of all the patterns here, Ali considers trust decay and scoring paramount. However your harness or loop is built, you are harnessing the agent and controlling its loop, and constant monitoring is the only way to succeed with agentic systems. As Juan puts it, it is not set and forget.
Synthetic data for golden datasets
Evaluations depend on a golden dataset that defines what good looks like, and we have had countless conversations with customers about exactly that. Our book introduces the preference-controlled synthetic data generation pattern to define and expand your golden dataset. You can then use that dataset for evaluations or for identifying high-value executions.
Building agentic systems is not a one-and-done task. It is a journey with multiple iterations, and you need to reach high-quality, desirable traces before you can start building your golden dataset from them.
From there, products on the market let you run your tests, identify what good looks like and evaluate against it. There are two kinds of evaluation.
Online evaluations happen while the agent is working.
Offline evaluations happen afterward, over the traces.
This may sound obvious, but teams often miss it. Looking only at the agent’s output is not enough. You have to define trajectory evaluations and evaluate the path the agent took.
Guardrails wrap around the container
Two questions from the session audience sharpened how these patterns fit together.
The first asked how to handle the growing risk of prompt injection. Our answer starts with the layers already covered.
Model-level filtering. A filter such as Model Armor screens everything before it reaches the model and after it leaves. It helps with automated control of injection and data exfiltration, keeping PII from leaving and data loss prevention.
The gateway. An agent gateway controls the external blast radius.
Least privilege. This principle matters just as much. Defining guardrails takes longer and requires extra effort, but these are the patterns we see in practice.
Red teaming. Red teaming is very useful when testing agents, and our book covers it as a pattern too.
Constant monitoring. Every agent has its own model, its own brain, and when a prompt injection attack lands, that agent can suddenly start behaving differently. Monitoring the confidence you have in its performance, which is what the trust decay and scoring pattern does, is how you catch the change.
All of it starts with defining the journey. That means knowing what your agentic building practice looks like, from agent inception through the first POCs, the tests, ongoing evaluations and optimizations.
The second question asked if anyone is developing hardened container images for agents. Ali took it first from the angle of the agents we ship at Google. We produce first-party agents such as a deep research agent, an ideation agent and a data science agent. These tend to be horizontal rather than vertical, so a financial services agent is not something Google would build, while deep research is. They are available in the Gemini Enterprise app or as an API you can call.
Juan added that our platform offers different levels of abstraction. Gemini Enterprise Agent Platform, which we introduced at Google Cloud Next, includes Agent Runtime, an optimized runtime that provides an optimized container for the agent.
But the guardrails you set up happen around the container rather than inside it. The container is essentially the runtime that interacts with the LLM, and since that interaction goes out to an external API, protection comes from how you wrap it.
Think about it the way you think about security models, as layers of an onion.
Prompt engineering might be your first barrier.
Running the agent in a sandbox adds a second.
Agent identity with the right permissions adds a third.
Model-level filtering adds a fourth.
The agent gateway adds a fifth.
Layered that way, the constraints come from what surrounds the container rather than from the container itself.
Long-running work starts with loop engineering
Another question asked how we handle long-running, asynchronous agent workflows that stretch across hours or days. The short answer is loop engineering. The longer answer has to do with the foundations of how you work. Loop engineering emerged in the first place because we needed controlled ways for agents to complete long-running asynchronous tasks. You design a loop, and your orchestrator moves through it.
Several components come together around that loop, each at its own layer.
The loop. Loop engineering defines how the agent moves through the task and works across its parts.
Agent-to-agent interaction. When two agents work together, the Agent2Agent (A2A) protocol supports some degree of task management, state management and asynchronous work.
The harness. The harness supports everything the loop needs. When the work involves long context, that includes context optimization such as context compaction and memory compaction.
The model. The base layer is the model that everything wraps around.
If you are getting started, pick a project and begin. Experiment with the different harnesses out there and look at how they differ. Ask why we use the things we use. Skills are popular right now, for example, and it is worth asking why you would use a skill instead of an agent or something else. Long-running tasks suffer from instruction bloat. That is why progressive disclosure exists, and progressive disclosure is why we use skills.
Some well-known patterns are in our book. You can find others by reading the source code of these harnesses, since most of them, apart from the proprietary ones, are open. Juan enjoys exploring the source of assistant agents like OpenClaw and Hermes, which form a category of their own.
Different kinds of agents serve different purposes.
Long-running assistant agents fit some use cases.
Workflow agents, Juan’s term for something more scoped, fit other workflows. You build them with a framework like ADK, LangGraph or CrewAI.
Cost, speed and quality trade off against each other, and you can reach similar results by different paths. Everything should start, and end, with the value you bring to the customer. Only when the work genuinely needs long-running tasks with sophisticated requirements should you start adding components. Some agent systems are the equivalent of a meeting that should have been an email.
Architects move from blueprints to harness design
The audience also asked how the architect’s role changes as engineers think more about architecture and less about each line of code. Ali hears a version of this often from students at the two universities where he teaches. They ask if they should even be in software engineering, and if there is still a career in it. His answer is yes, they are in the right profession.
The shift underway is from manual code authoring to high-level system orchestration, and it is reshaping software engineering. If a model handles the syntax, the boilerplate and the low-level implementation, you can focus on steering the system. To use the analogy we have been building, you are steering a self-driving car that takes you from where you are to where you want to go. Your attention moves to how the parts of the system fit together, how it handles edge cases and how it evolves over time.
In this era, the architect’s role evolves from a top-down blueprint designer into an orchestrator of dynamic, autonomous, variation-heavy environments. That lets you focus on systemic intent and harness design rather than static blueprints. Coming back to the trust decay and scoring pattern, you are developing a harness that enforces boundaries and crafts telemetry, using the OpenTelemetry standard we described earlier. You then test that harness through evaluation loops that continuously prove the generated system satisfies its nonfunctional requirements, such as latency, cost, quality and resilience.
Juan compares this moment to the industrial revolution, where the biggest change was the invention of the steam engine. The work shifted from the physical strain of humans and animals onto machines, and new roles came with it. The driver existed before, but not at the level of sophistication an engine required, and engineering grew into a role of its own out of the same shift.
He also believes that to succeed with agent-assisted coding, you need a background in coding. You may have seen the memes online about an amazing new product that does not even require authentication and is served from localhost. Those memes are a reminder of why that background still matters as the shift Ali describes takes hold.
The new world is producing people who will define the AI foundations and the AI strategy for their companies. It is also producing roles like forward deployed engineers, who deploy production-ready code at customer sites. All of these come together as an evolution of how we work.
Underneath all of it is the engineer orchestrating the swarms of agents that do the job. That engineer needs a clear notion of what good looks like, how to enforce it, and how to know the system is working the way it should. Juan’s golden question is whether you would trust your credit card to an agent.
Each maturity level has its own patterns
There are many ways to reach an objective, but starting haphazardly makes it very difficult to hill-climb and get better over time. That holds for the agentic systems you build, for the organizational structure around them, and for the architecture that supports each stage of maturity.
Each level of maturity has a corresponding set of patterns. Reaching the next level means looking at that pattern set, picking the ones that fit and implementing them, which gives you the capability to move up. We think that idea matters for anyone on an agentic AI journey.
The patterns in our book remain highly relevant today, and they apply directly to harness engineering and loop engineering. The biggest thing is understanding that there is a maturity model and that every organization is somewhere on it. Reaching the outcomes you want means identifying the most feasible, governed, enterprise-ready path to get there.
It is exciting to see real value arriving from agents, and some of what shows up in the news is scary, so this deserves to be taken seriously. We have been building software for more than fifty years. Those lessons still apply, and the industry has had this same conversation many times before.
There are multiple frameworks, harnesses and loop engineering techniques out there. What we wanted to offer is an abstraction you can work from, a way to map the things you need to think about when you build on any of them.
Our book’s GitHub repository includes a loan processing use case that illustrates many of these principles. It builds a compliance-aware loan origination agent three ways:
As a single agent on ADK.
As an orchestrator with specialist sub-agents for document validation, credit checks, risk assessment and compliance.
Again in CrewAI and LangGraph.
You can find more at ai-patterns.com, including skills for the patterns and a skills catalog you can browse by agent.
Dr. Ali Arsanjani and Juan Pablo Bustos are the authors of Agentic Architectural Patterns for Building Multi-Agent Systems, published by Packt, which also publishes Deep Engineering.










