This deep dive is based on the ARC 2026 session on spec-driven development led by Packt authors, and has been edited in their own words.
AI promises to make software development easier, and we believe that. The hard part is making intent survive generation, and that is what this piece is about. At the start of last year, Alessandro caught himself doing what most of us did with coding agents, which was prompt and pray. This is how we stopped.
Everything below is in our repository, BrewUp/DWX26. We built it for an earlier delivery of this talk at DWX, which is where the name comes from. Each branch captures one run of the exercise, so you can follow along and diff them yourself.
Here are the slides from the session.
A reasonable prompt, a green PR
BrewUp is the ERP we use to experiment with architectural and technical ideas. It is about beer, because we love beer. It is a modular monolith built on CQRS and event sourcing, and it already has Sales and Warehouse bounded contexts, plus a bit of payment.
The feature request came out of a conversation with the customer and the team, and it is a well-defined prompt.
Implement order confirmation for BrewUp.
An order can be confirmed when:
- the customer payment has been authorized;
- all requested beers are available in the warehouse.
When the order is confirmed, reserve the stock.Alessandro put it into his coding agent. The agent analyzed the code and came back with a solution that compiled, passed its tests and made absolute sense at first sight. It created a PaymentAuthorizationId and a StockReservationId, reused our libraries, and raised a SalesOrderConfirmed event after the invariant checks. One commit, 39 files changed, 951 lines added.
Then Alessandro asked Alberto whether he would approve the PR. Alberto said no. The code was fine, but the model was wrong.
One aggregate, three authorities
Three parts of BrewUp each own a different decision here. Sales owns the commercial commitment, Payment owns payment authorization, and Warehouse owns physical stock. The agent folded Payment into Sales and Warehouse. From a domain-driven design perspective, that left Payment implicit rather than explicit, and nothing in the solution had authority over it.
The ids are where it shows. The names are right, but look at where they come from.
// SalesOrderSaga.cs, no_sdd branch
internal void MarkCustomerBudgetAsVerified(CustomerJson customer, Guid correlationId)
{
var paymentAuthorizationId = new PaymentAuthorizationId(correlationId.ToString());
...
}
// SalesOrder.cs, no_sdd branch
internal void ConfirmOrder(
PaymentAuthorizationId paymentAuthorizationId,
StockReservationId stockReservationId,
Guid correlationId)
{
...
if (string.IsNullOrWhiteSpace(paymentAuthorizationId.Value))
throw new ArgumentException("Payment authorization id is required.", nameof(paymentAuthorizationId));
...
}The Sales Order checks that the authorization id is not empty. The saga guarantees it never is by minting one from its own correlation id. So the guard passes, and nobody ever authorized anything. An identifier is not evidence unless the owning authority produced the decision it refers to.
The process has the same problem. Authorizing payment, reserving stock and confirming the order collapsed into one transaction, so when payment is authorized and the reservation fails, there is no clean way back. Several policies are plausible. None is correct until someone with domain authority decides.
A longer prompt is not governance
The natural reaction is a better prompt, and then another one. We have both been there, seven hundred prompts in and sure the next one will be right.
The problem is that the LLM reads everything again in every new session, but nothing learned in prompt v1 carries into prompt v2. Every prompt is a new conversation. Synchronized copy and paste is not governance.
That is not how we work with a team. When we discover something about the domain, we write it down, and our knowledge grows over time. Spec-driven development applies the same discipline to agents.
Write the decision down once
Spec-driven development took shape as a methodology in 2025. We use GitHub Spec Kit as the example, but there are plenty of frameworks, including GSD, OpenSpec, Agent OS and BMAD, and you should choose the one you like most. If you work alone or in a small team, Spec Kit gives you clear steps, and GSD is lighter still. BMAD is more structured and closer to agile. It is really enterprise, and its token consumption is enterprise level too. Our repository has a BMAD version of the same exercise.
The backbone is the constitution. You write it once, at the start of the project, and then every feature repeats the same loop of specify, clarify, plan, tasks and implement.
The constitution holds the project’s non-negotiable principles. It is not a feature specification or a detailed architecture manual. It is the governance baseline. Ours covers domain-driven design, modular architecture, property-based testing, CQRS and event sourcing, how to split the layers, and quality gates. Two of its rules carry most of this story.
- Domain ownership MUST be explicit. A bounded context may depend on a decision produced by
another bounded context, but it MUST NOT own or reproduce that decision unless explicitly
assigned by the specification.
- Unknown business policy MUST remain visible as an open question or `[NEEDS CLARIFICATION]`.
Agents MUST NOT invent domain policy to make a model look complete.This is the first big difference between a prompt and a specification. A prompt asks for an output. A specification records decisions that subsequent outputs must respect. A prompt lives in a conversation and invites assumptions. A specification lives in the repository with the code and gets referenced and checked, which makes it closer to a decision you take with your team than a conversation with your LLM.
Domain carriers load the language first
Spec Kit is just a tool, and it is your job to customize it into a strong harness. Ours starts in .specify/memory, which holds three kinds of file.
.specify/memory/
constitution.md
architecture/brewup-module-structure.md
domain-carriers/brewup-sales-order-confirmation.mdWe split the architecture rules out of the constitution so we can change them for the next project. Sometimes Alberto adds a database file too, to make storage choices explicit. Once you write these files you reuse them for every feature in the project, and more or less half of them carry over to the next one.
The domain carrier is the important one. It holds the ubiquitous language and the ownership rules for a feature before any specification exists. The rules are concrete. It is Sales Order, never generic Order. A Payment Authorization is an outcome produced by Payment. A Stock Reservation is an outcome produced by Warehouse, and availability is not a durable fact. Each ownership rule gets an id.
## BC-004 — Payment Authorization is an external decision
Sales may store `PaymentAuthorizationId` as evidence that Payment produced an authorization outcome.
Sales may react to a payment authorization outcome.
Sales may request payment authorization through an integration boundary.
Sales must not produce the payment authorization outcome.Spec Kit exposes hooks around each command in .specify/extensions.yml, and that is where the carrier gets loaded. When Alessandro types /speckit.specify, our agent fires first. It tells the model that it cannot take decisions by itself, and that it has to read these files and decide from them.
# .specify/extensions.yml (excerpt)
hooks:
before_specify:
- extension: brewup-sdd
command: speckit.brewup.load-domain-context
description: Load BrewUp Sales Order domain context before specificationPeople always ask us two things about this. First, we did not write every character of these files. An LLM drafted them and we corrected them, and if you know what you want, describing it is not a big effort. Second, in our experience all this markdown does not fill the context window, because each step loads only the files it needs, as long as you structure them correctly.
Let the agent find the ambiguity
/speckit.specify takes basically the same prompt we started with and iterates on it against the constitution and the carrier. The output is a spec.md where the language and the ownership become durable, with user stories and acceptance scenarios generated from it.
/speckit.clarify is where it gets bloody powerful. The agent may identify the ambiguity. The domain expert must resolve it. The specification must preserve the resolution. For order confirmation, clarify surfaced the questions we had deliberately left open, including the confirmation sequence, what happens when payment is authorized but stock cannot be reserved, and partial reservations.
Our answers are now recorded in the spec. Partial confirmation is allowed for the subset Warehouse can reserve. If one side succeeds and the other fails, the Sales Order stays unconfirmed with no compensation owned by Sales, because release, void and refund belong to Warehouse and Payment. Payment authorization and stock reservation are requested in parallel. Sales never interprets a payment-provider timeout and reacts only to the definitive outcomes that Payment emits.
This forces you to stay in the problem space longer before you move to the solution space, and that is a really good thing. It also means the architect or the domain expert answers these questions, never the LLM.
Spend tokens on thinking, not typing
Only after that do we talk about the tech stack, with /speckit.plan. We point the model at our own libraries, a small CQRS and event-sourcing boilerplate we wrote because we are lazy, and it infers our commands and interfaces from them. The plan covers the event store, the testing approach and the modular monolith. Every bounded context in it references a rule in the constitution, and it shows the structure the model intends to build, so we can validate it before any code exists.
/speckit.tasks turns the plan into tasks for each user story. Each task becomes a commit touching two or three files, five at most, so when Alessandro sends a PR, Alberto doesn’t shout at him anymore. Then /speckit.implement executes them, and you do not have to run everything at once. You can implement tasks 001 to 003, or just the first user story, and review a small piece each time. You can never trust the output of an LLM, so read every one.
Use a powerful frontier model for the constitution, specify, clarify, plan and tasks. Those steps need the strongest reasoning available, so watch out for tokens. Implementation can drop to a much cheaper model, because the guardrails leave it very little to invent. Alessandro uses a high-reasoning model up to the tasks step and a cheaper medium one for implementation, and one round of review usually fixes what is left. We have tried pushing implementation down to the smallest models, and the output doesn’t satisfy either of us yet.
Spec Kit alone still missed Payment
Here is the part we did not expect. The sdd_without_custom_agents branch is the same exercise run with Spec Kit and our domain carriers, but without our custom guard agents. Its specification got ownership right on paper. It says that producing the payment authorization is owned by the Payment authority, and then it calls that out of scope. The plan built exactly what the spec said. There is still no Payment module, only a saga passing the ids along through integration events.
The model preserved the conceptual ownership and still failed to create the physical structure BrewUp needs. BC rules answer who owns the decision, and we needed rules that answer where that ownership must live in the solution. The missing one was simple. If Payment is a business authority in scope for implementation, Payment must become a BrewUp module.
### AR-003 — New authorities must become modules when implementation is in scope
If a specification introduces a new authority and implementation is in scope, the plan MUST
create the corresponding module and solution folder before adding behavior.Rules alone are not enough, so the harness checks them at every step. Six custom agents hang off the Spec Kit hooks.
Before specify, one agent loads the domain context.
After specify, a domain guard checks ownership.
Before plan, a readiness check gates planning.
After plan, plan and module-structure guards review it.
After tasks, task and module-structure guards review them.
None of them is there to improve anything creatively. Each one verifies that the previous step really read our files and did what we wrote.
The sdd-with-harness branch is the result. Payment is a real module with the same projects as Sales and Warehouse, because that is how we want things, not because the LLM decided for us.
src/Payment/
BrewUp.Payment.Domain
BrewUp.Payment.Facade
BrewUp.Payment.Infrastructure
BrewUp.Payment.ReadModel
BrewUp.Payment.SharedKernel
BrewUp.Payment.TestsThe evidence is now earned. Payment’s own aggregate makes the decision, and Warehouse creates its own StockReservationId when it reserves stock.
// PaymentAuthorization.cs, sdd-with-harness branch
internal void Authorize(string salesOrderId, Price amount, Guid correlationId)
{
// Idempotency: already authorized/declined → no-op
if (!Equals(_status, PaymentAuthorizationStatus.Pending))
return;
if (amount.Value > 0)
RaiseEvent(new PaymentAuthorized(new PaymentAuthorizationId(Id.Value), correlationId, salesOrderId));
else
RaiseEvent(new PaymentDeclined(new PaymentAuthorizationId(Id.Value), correlationId, salesOrderId,
"Amount must be greater than zero"));
}It may not be the best solution in the world, but it is our solution. Because it follows our structure, we can navigate it, understand it and fix a bug without asking the LLM.
Make drift visible with analyze
SDD does not make the model deterministic. The tool is still probabilistic, and you can make mistakes writing the markdown too, because you are human. Even with a specification in place, an implementation can still manufacture the evidence instead of earning it, and /speckit.analyze is how we catch it.
/speckit.analyze
CRITICAL — Invented external decision evidence
PaymentAuthorizationId and StockReservationId are recorded, but their domain meaning is not
produced by the owning authorities recorded in the specification.
BC-003 Payment owns authorization outcomes.
BC-004 Payment Authorization is produced by Payment.
BC-006 Stock Reservation is produced by Warehouse.
BC-007 Warehouse owns physical stock and reservations.
BC-011 The agent must not invent reservation lifecycle, timeout, retry, void or release policy.SDD does not guarantee obedience. It makes architectural drift inspectable. Instead of telling a colleague an architecture feels wrong, you can tell them the plan violates BC-003, BC-007 and BC-009.
When implementation hits something the spec does not settle, it stops and marks the gap as [NEEDS CLARIFICATION] instead of guessing. You answer, and the flow regenerates the affected task. If someone pushes a hotfix outside the workflow, analyze can still search the code for breaches of the constitution, and because the code follows our structure we can read it immediately. The person who made the hotfix should at least update the documentation afterwards.
Requirements change, the spec moves first
Then a new requirement arrives. Approved wholesale customers may order using pre-approved payment terms, with no immediate payment authorization.
The specification changes first. A Sales Order is now confirmed when stock has been reserved and either payment is authorized or valid payment terms are recorded. Running analyze after that change shows the blast radius before anyone touches the code.
/speckit.analyze, after the change
HIGH — Requirement no longer covered
Task T-18 requires PaymentAuthorizationId for every Sales Order confirmation.
The evolved spec allows PaymentTermsApprovalId for eligible wholesale customers.A new feature goes through the same loop. Instead of 001 you get 002 with the name of your specification. When something changes, you go back through specification, plan and code again.
Architects govern the context now
None of the architect’s work goes away. We still design components and boundaries, communicate decisions, review design and code, and resolve ambiguity. With agents, each of those becomes more explicit. Boundaries get encoded as decision authority, decisions become persistent and checkable, review covers the whole artifact chain, and ambiguity gets exposed before execution instead of during it. The architect designs the context within which the AI is allowed to operate.
So stop opening prompts with “you are a developer with thirty years of experience.” AI is not your senior architect. It’s a very fast junior with no memory of your business, no taste, and no shame. Talk to it the way Alessandro talks to his niece, telling it what it can do and, just as important, what it cannot do.
There is a bonus we did not expect. These files are just documentation, and the agent forces you to write it. Working alone or in a small team, you used to be the only guardian of the knowledge. Writing it down was always part of an architect’s job, or it should have been.
A prompt describes a task. A specification records a decision. Architecture makes those decisions survive implementation. The prompt delegates hidden decisions to the model, and the specification makes them explicit, reviewable and ours.
If you want to try it, clone the repository and diff no_sdd against sdd-with-harness. Then write one domain carrier for your next feature before you write a single prompt.
Editorial note: This article is adapted from Alessandro Colla and Alberto Acerbis’s ARC 2026 session, Spec-Driven Development: Redefining the Software Architect in the AI Era.
The session recording and slides were condensed and reordered for print, and every code excerpt is taken verbatim from the branches of the BrewUp/DWX26 repository. Alessandro and Alberto are the authors of Domain-Driven Refactoring, published by Packt, which also publishes Deep Engineering.









