Agent Skill Engineering Workshop
Getting an agent to work once is easy. Join Youssef Hosni to build reusable Agent Skills in Claude Code that load only when needed, compose into larger workflows, and hold up in production.
🗓️ Sat 17 Oct, 9 AM ET · One day, build-along
✍️ From the editor’s desk
Welcome to the 66th issue of Deep Engineering!
On 29 September, GitHub’s Spec Kit added an extension for OpenSpec, another spec-driven development framework. As these tools start plugging into each other, the choice of framework matters less than what you put inside it.
Alessandro Colla and Alberto Acerbis showed why at ARC 2026. A reasonable prompt gave them a green PR that touched 39 files and minted its own payment authorization id, because nothing in the repository said Payment owned that decision. Spec Kit with their domain rules still produced no Payment module. Only a harness of domain carriers, with guard agents at every step, fixed it.
Colla and Acerbis are software architects and the co-authors of Domain-Driven Refactoring, published by Packt. In today’s issue they take that PR apart and rebuild it, starting with a domain carrier that loads before the agent writes a line and ending with the one architecture rule that finally made Payment a real module.
Let’s get started.
🗓️ Sat 24 Oct, 10:30 AM ET · 25% off
Not every step in an agent needs an LLM. Join Jaime Buelta to build a Python agent that pairs fast Jev decision models with LLMs, tools and MCP, and keep it reviewable with spec-driven development.
🧠 Practical Deep Dive
Spec-Driven Development Needs a Harness, Not Just a Spec
AI promises to make software development easier, and we believe that. The hard part is making intent survive generation, and that is what this piece is about. At the start of last year, Alessandro caught himself doing what most of us did with coding agents, which was prompt and pray. This is how we stopped.
Everything below is in our repository, BrewUp/DWX26. We built it for an earlier delivery of this talk at DWX, which is where the name comes from. Each branch captures one run of the exercise, so you can follow along and diff them yourself.
A reasonable prompt, a green PR
BrewUp is the ERP we use to experiment with architectural and technical ideas. It is about beer, because we love beer. It is a modular monolith built on CQRS and event sourcing, and it already has Sales and Warehouse bounded contexts, plus a bit of payment.
The feature request came out of a conversation with the customer and the team, and it is a well-defined prompt.
Implement order confirmation for BrewUp.
An order can be confirmed when:
- the customer payment has been authorized;
- all requested beers are available in the warehouse.
When the order is confirmed, reserve the stock.Alessandro put it into his coding agent. The agent analyzed the code and came back with a solution that compiled, passed its tests and made absolute sense at first sight. It created a PaymentAuthorizationId and a StockReservationId, reused our libraries, and raised a SalesOrderConfirmed event after the invariant checks. One commit, 39 files changed, 951 lines added.
Then Alessandro asked Alberto whether he would approve the PR. Alberto said no. The code was fine, but the model was wrong.
One aggregate, three authorities
Three parts of BrewUp each own a different decision here. Sales owns the commercial commitment, Payment owns payment authorization, and Warehouse owns physical stock. The agent folded Payment into Sales and Warehouse. From a domain-driven design perspective, that left Payment implicit rather than explicit, and nothing in the solution had authority over it.
The ids are where it shows. The names are right, but look at where they come from.
// SalesOrderSaga.cs, no_sdd branch
internal void MarkCustomerBudgetAsVerified(CustomerJson customer, Guid correlationId)
{
var paymentAuthorizationId = new PaymentAuthorizationId(correlationId.ToString());
...
}
// SalesOrder.cs, no_sdd branch
internal void ConfirmOrder(
PaymentAuthorizationId paymentAuthorizationId,
StockReservationId stockReservationId,
Guid correlationId)
{
...
if (string.IsNullOrWhiteSpace(paymentAuthorizationId.Value))
throw new ArgumentException("Payment authorization id is required.", nameof(paymentAuthorizationId));
...
}The Sales Order checks that the authorization id is not empty. The saga guarantees it never is by minting one from its own correlation id. So the guard passes, and nobody ever authorized anything. An identifier is not evidence unless the owning authority produced the decision it refers to.
The process has the same problem. Authorizing payment, reserving stock and confirming the order collapsed into one transaction, so when payment is authorized and the reservation fails, there is no clean way back. Several policies are plausible. None is correct until someone with domain authority decides.
A longer prompt is not governance
The natural reaction is a better prompt, and then another one. We have both been there, seven hundred prompts in and sure the next one will be right.
The problem is that the LLM reads everything again in every new session, but nothing learned in prompt v1 carries into prompt v2. Every prompt is a new conversation. Synchronized copy and paste is not governance.
That is not how we work with a team. When we discover something about the domain, we write it down, and our knowledge grows over time. Spec-driven development applies the same discipline to agents.
Write the decision down once
Spec-driven development took shape as a methodology in 2025. We use GitHub Spec Kit as the example, but there are plenty of frameworks, including GSD, OpenSpec, Agent OS and BMAD, and you should choose the one you like most. If you work alone or in a small team, Spec Kit gives you clear steps, and GSD is lighter still. BMAD is more structured and closer to agile. It is really enterprise, and its token consumption is enterprise level too. Our repository has a BMAD version of the same exercise.
The backbone is the constitution. You write it once, at the start of the project, and then every feature repeats the same loop of specify, clarify, plan, tasks and implement.
The constitution holds the project’s non-negotiable principles. It is not a feature specification or a detailed architecture manual. It is the governance baseline. Ours covers domain-driven design, modular architecture, property-based testing, CQRS and event sourcing, how to split the layers, and quality gates. Two of its rules carry most of this story.
- Domain ownership MUST be explicit. A bounded context may depend on a decision produced by
another bounded context, but it MUST NOT own or reproduce that decision unless explicitly
assigned by the specification.
- Unknown business policy MUST remain visible as an open question or `[NEEDS CLARIFICATION]`.
Agents MUST NOT invent domain policy to make a model look complete.This is the first big difference between a prompt and a specification. A prompt asks for an output. A specification records decisions that subsequent outputs must respect. A prompt lives in a conversation and invites assumptions. A specification lives in the repository with the code and gets referenced and checked, which makes it closer to a decision you take with your team than a conversation with your LLM.
Domain carriers load the language first
Spec Kit is just a tool, and it is your job to customize it into a strong harness. Ours starts in .specify/memory, which holds three kinds of file.
.specify/memory/
constitution.md
architecture/brewup-module-structure.md
domain-carriers/brewup-sales-order-confirmation.mdWe split the architecture rules out of the constitution so we can change them for the next project. Sometimes Alberto adds a database file too, to make storage choices explicit. Once you write these files you reuse them for every feature in the project, and more or less half of them carry over to the next one.
Continue reading → In the rest of the deep dive, they wire domain carriers and six guard agents into Spec Kit’s hooks and show why Spec Kit alone still missed Payment.
New in Engineering Leadership
We interviewed engineering leaders on how teams keep control once agents start writing the code. Here’s what some of them told us.
Itamar Friedman, co-founder and CEO of Qodo, on why governance has to become infrastructure once agents write the code.
Dipanjan Sengupta of EY on treating agent autonomy as a governance decision.
Cassie Shum of RelationalAI on how AI acceleration exposes the friction developers used to absorb.
🛠️ Tool of the Week
ArchUnitNET - A .NET library for writing architecture rules as unit tests, so a dependency that crosses a layer or module boundary fails the build.
Fluent C# rules that read like sentences
Checks dependencies, naming, inheritance and attributes
Plugs into xUnit, NUnit and MSTest
Apache 2.0, with v0.13.4 released in August
📎 Tech Briefs
FTC opens AI probe - The FTC is investigating OpenAI, Anthropic and other AI companies over the product risks of their models.
EDG goes open source - The thirty-year-old commercial C++ front end is now public on GitHub, with The C++ Alliance as its nonprofit home.
AG-UI 1.0 - The open agent-to-app protocol freezes its spec, with TypeScript, Python and .NET SDKs generated from one JSON Schema.
What TLA+ can and can’t check - Hillel Wayne on why formal specs won’t rescue agentic coding alone, and which properties TLA+ can’t even express.
Marten 9.44.0 - Fixes a nested
Any()query that built valid SQL against the wrong key and silently returned no rows.
That’s all for this week. If you take one thing from Alessandro and Alberto, write one domain carrier before your next prompt.
Keep building,
Saqib Jan, Editor-in-Chief, Deep Engineering
Partner with Deep Engineering
If your company wants to reach senior developers, software engineers, and technical decision-makers, speak to us about partnering with Deep Engineering







