<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Packt Deep Engineering]]></title><description><![CDATA[Deep Engineering is a weekly newsletter for developers and software architects featuring expert-led insights, deep dives into modern systems, and clear thinking on real-world software design.]]></description><link>https://deepengineering.net</link><image><url>https://substackcdn.com/image/fetch/$s_!H5BJ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png</url><title>Packt Deep Engineering</title><link>https://deepengineering.net</link></image><generator>Substack</generator><lastBuildDate>Wed, 12 Aug 2026 03:55:17 GMT</lastBuildDate><atom:link href="https://deepengineering.net/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Packt]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[deepengineering@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[deepengineering@substack.com]]></itunes:email><itunes:name><![CDATA[Packt]]></itunes:name></itunes:owner><itunes:author><![CDATA[Packt]]></itunes:author><googleplay:owner><![CDATA[deepengineering@substack.com]]></googleplay:owner><googleplay:email><![CDATA[deepengineering@substack.com]]></googleplay:email><googleplay:author><![CDATA[Packt]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Deep Engineering #58: Sebastian Hassinger on Where Quantum Progress is Real]]></title><description><![CDATA[On why qubit counts measure register size rather than capability, what code distance reveals that a headline number hides, and where quantum computing delivers first.]]></description><link>https://deepengineering.net/p/issue-58-sebastian-hassinger-qubit-counts</link><guid isPermaLink="false">https://deepengineering.net/p/issue-58-sebastian-hassinger-qubit-counts</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 06 Aug 2026 15:45:37 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c258bf82-e5d3-4213-9f0b-4beac02d7cec_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Featured - <a href="https://www.eventbrite.co.uk/e/langgraph-masterclass-from-beginner-to-professional-tickets-1992773766981?aff=deepeng&amp;discount=DEEPENG50">LangGraph Masterclass: From Beginner to Professional</a></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.eventbrite.co.uk/e/langgraph-masterclass-from-beginner-to-professional-tickets-1992773766981?aff=deepeng&amp;discount=DEEPENG50" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8Nfc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dee42c8-2093-48df-bcbd-a3988a0b12af_900x300.png 424w, https://substackcdn.com/image/fetch/$s_!8Nfc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dee42c8-2093-48df-bcbd-a3988a0b12af_900x300.png 848w, https://substackcdn.com/image/fetch/$s_!8Nfc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dee42c8-2093-48df-bcbd-a3988a0b12af_900x300.png 1272w, https://substackcdn.com/image/fetch/$s_!8Nfc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dee42c8-2093-48df-bcbd-a3988a0b12af_900x300.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8Nfc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dee42c8-2093-48df-bcbd-a3988a0b12af_900x300.png" width="900" height="300" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8dee42c8-2093-48df-bcbd-a3988a0b12af_900x300.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:300,&quot;width&quot;:900,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.eventbrite.co.uk/e/langgraph-masterclass-from-beginner-to-professional-tickets-1992773766981?aff=deepeng&amp;discount=DEEPENG50&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!8Nfc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dee42c8-2093-48df-bcbd-a3988a0b12af_900x300.png 424w, https://substackcdn.com/image/fetch/$s_!8Nfc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dee42c8-2093-48df-bcbd-a3988a0b12af_900x300.png 848w, https://substackcdn.com/image/fetch/$s_!8Nfc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dee42c8-2093-48df-bcbd-a3988a0b12af_900x300.png 1272w, https://substackcdn.com/image/fetch/$s_!8Nfc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dee42c8-2093-48df-bcbd-a3988a0b12af_900x300.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This hands-on masterclass takes you from <strong>LangGraph fundamentals</strong> to <strong>supervisor</strong> and <strong>hierarchical</strong> multi-agent systems, with <strong>live debugging</strong> in LangSmith throughout. </p><p style="text-align: center;"><span>Deep Engineering readers save </span><strong><span>50%</span></strong><span> with code - </span><strong>DEEPENG50</strong><span>.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.co.uk/e/langgraph-masterclass-from-beginner-to-professional-tickets-1992773766981?aff=deepeng&amp;discount=DEEPENG50&quot;,&quot;text&quot;:&quot;Register here &#8594;&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.eventbrite.co.uk/e/langgraph-masterclass-from-beginner-to-professional-tickets-1992773766981?aff=deepeng&amp;discount=DEEPENG50"><span>Register here &#8594;</span></a></p><div><hr></div><p><span>&#9997;&#65039; </span><strong><span>From the editor&#8217;s desk,</span></strong></p><p><span>Welcome to the </span><strong><span>58th</span></strong><span> issue of </span><strong><span>Deep Engineering</span></strong><span>!</span></p><p>On August 3, NTT announced that it has signed a capital and business alliance with OptQC, the University of Tokyo spinout building optical quantum processors, with both companies aiming at a fault-tolerant machine of one million qubits. The <a href="https://group.ntt/en/newsrelease/2026/08/03/260803a.html">announcement from NTT</a> lays out a phased roadmap. The companies aim to complete the system architecture and key component technologies by fiscal 2027 alongside a practical 10,000 qubit system, then begin verification work in fiscal 2028 and deliver a platform for running optical and classical machines together the year after.</p><p>One million is the largest number the field has yet attached to a headline, and it arrives in a year when the reported metric already shifted once. Through the first half of 2026 vendors moved from physical qubit counts to logical qubit counts as the figure worth announcing. Both are real results, and both leave out the properties that decide what a machine can actually compute, which is where this issue picks up.</p><p><a href="https://www.linkedin.com/in/shassinger">Sebastian Hassinger</a> has read claims like these from inside the companies making them, first on the <a href="https://www.ibm.com/quantum">IBM Quantum team</a> and later leading go to market for <a href="https://aws.amazon.com/braket/">AWS Quantum Technologies</a>. He wrote <a href="https://www.packtpub.com/en-us/product/the-new-quantum-era-9781807787370">The New Quantum Era</a> for readers without a physics background, and today he walks us through which numbers carry the information and which ones do not.</p><blockquote><p>You can watch the full session or read the <a href="https://deepengineering.net/p/quantum-computing-beyond-the-hype-sebastian-hassinger">transcript here</a>.</p></blockquote><p><strong>Let&#8217;s get started.</strong> </p><div class="callout-block" data-callout="true"><h2 style="text-align: center;"><a href="https://www.vpdae.com/redirect/o2rdxoqtpm0p1but31d2wk9ffrb"><span data-color="#f97141" style="color: rgb(249, 113, 65);">Agent-written TLA+</span></a></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.vpdae.com/redirect/o2rdxoqtpm0p1but31d2wk9ffrb" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sw8h!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 424w, https://substackcdn.com/image/fetch/$s_!sw8h!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 848w, https://substackcdn.com/image/fetch/$s_!sw8h!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 1272w, https://substackcdn.com/image/fetch/$s_!sw8h!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sw8h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png" width="296" height="296" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:300,&quot;width&quot;:300,&quot;resizeWidth&quot;:296,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.vpdae.com/redirect/o2rdxoqtpm0p1but31d2wk9ffrb&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sw8h!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 424w, https://substackcdn.com/image/fetch/$s_!sw8h!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 848w, https://substackcdn.com/image/fetch/$s_!sw8h!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 1272w, https://substackcdn.com/image/fetch/$s_!sw8h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;">An agent wrote our <strong>TLA+</strong> spec. The model checker explored <strong>14.3M</strong> states and caught a real race.</p><p style="text-align: center;"></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.vpdae.com/redirect/o2rdxoqtpm0p1but31d2wk9ffrb&quot;,&quot;text&quot;:&quot;Read the write-up&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.vpdae.com/redirect/o2rdxoqtpm0p1but31d2wk9ffrb"><span>Read the write-up</span></a></p><p style="text-align: center;"></p></div><div><hr></div><p><strong>Expert Insight</strong></p><h2><span>Qubit Count Measures Register Size, Not Capability</span></h2><p><em>by <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;id&quot;:427210082,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;uuid&quot;:&quot;e7195419-ea93-4f2b-84d4-a7959d36a209&quot;}" data-component-name="MentionToDOM"></span> with <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Sebastian Hassinger&quot;,&quot;id&quot;:535827458,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4a70390e-d1ad-4e61-8e13-23bc47b2a921_144x144.png&quot;,&quot;uuid&quot;:&quot;514aa520-1a03-4ba0-9100-b5edc23133a9&quot;}" data-component-name="MentionToDOM"></span> </em></p><p><span>Three logical qubit results landed in the first half of 2026, and in each one the number worth reading is not the number in the headline. QuEra published 96 logical qubits encoded across 448 neutral atoms, a ratio near five to one. Quantinuum reported 48 logical qubits drawn from 98 trapped ions, closer to two to one. IBM&#8217;s published</span><a href="https://www.ibm.com/quantum/blog/large-scale-ftqc"><span> fault tolerance roadmap</span></a><span> targets 200 logical qubits from roughly 10,000 physical ones by 2029, a ratio near fifty to one.</span></p><p><span>Those ratios differ by an order of magnitude because the underlying codes differ, and the choice of code decides whether a logical qubit corrects errors or only detects them. A count reported on its own collapses all of that into a single integer, which is why the integer tells you very little about what the machine computes.</span></p><p><a href="https://www.linkedin.com/in/shassinger">Sebastian Hassinger</a> worked on the <strong>IBM Quantum team</strong> and later led go to market for <strong>AWS Quantum Technologies</strong>, and he wrote<span> </span><a href="https://www.packtpub.com/en-us/product/the-new-quantum-era-9781807787370">The New Quantum Era</a> <span>to give engineers without a physics background enough grounding to read results like these directly. During our interview when I asked him what actually carries information in a milestone result, he began by taking apart the metric the field has reported for a decade. &#8220;Qubit count is effectively the register size of that computer,&#8221; he said. Then came the harder line, aimed at a claim the field now repeats freely. &#8220;Anytime you hear somebody saying quantum computing is just a matter of engineering now, be suspicious of that person&#8217;s claims.&#8221;</span></p><h3><span>Register size bounds information, not computation</span></h3><p><span>The clearest demonstration that a raw count says little about capability comes from IBM&#8217;s own hardware history. And Hassinger was there for it. The technical roadmap produced a chip called Condor at just over a thousand qubits, which he describes as genuinely valuable for the research and fabrication effort it took to build. But it saw little use. Connectivity between qubits on the chip was low and the noise proved very difficult to manage, so researchers went back to the smaller machines in the 127 to 133 qubit range, which were more capable in practice.</span></p><p><span>The reason is architectural rather than numerical. Register size sets an upper bound on how much information you can load, and nothing beyond that. Hassinger points out that QuEra&#8217;s 256 qubit Aquila holds 256 bits at a time, which sounds unremarkable until you entangle those qubits and produce a state vector of two to the 256, a computational space you cannot physically recreate on classical hardware. The capability lives in the entanglement structure and in how well the problem maps onto it, so a machine with more qubits and worse connectivity computes less than a smaller machine with better ones.</span></p><p><span>That same logic now applies one level up. A logical qubit is an encoding, not a unit, and its value depends on the code family, the physical to logical ratio, and the error model the code assumes. Codes at distance two detect errors without correcting them, which is a different guarantee from correction, and the difference produced considerable argument when Microsoft and Quantinuum reported reliable logical qubits in 2024 using error detection with post-selection. Two systems reporting the same logical qubit count can therefore be doing categorically different things.</span></p><h3><span>Code distance carries the information a count discards</span></h3><p><span>Hassinger&#8217;s proposed substitute is quite specific, and it happens to be exactly the property that separates those cases. Fidelity and noise are what matter, he reasons, particularly the fidelity of one and two qubit gates, where two qubit operations mean entanglement. Those figures are hard to extract from a published result, so he offers a proxy that survives summarization.</span></p><p><span>Read the resilience of the error correction code, expressed as a </span><strong><span>distance or a d value</span></strong><span>. That is roughly how many errors the system absorbs before the encoded information collapses and the computation is lost, so a higher distance means a more resilient machine. He points to the Willow experiment as the useful reference, roughly a hundred physical qubits arranged in a surface code presenting as one logical qubit at distance seven. That snapshot carries what you need without the underlying gate fidelities, because the only two questions that determine what you can run are &#8220;how many logical qubits do I get, and how resilient is that error correction.&#8221;</span></p><p><span>Pair a logical qubit count with its code, its distance, and its encoding ratio and you have something you can reason about. Take the count alone and you have an integer that happens to increase.</span></p><h3><span>Speculation hardens into certainty before it reaches you</span></h3><p><span>There is a structural reason the public record runs ahead of the results, and it operates on the way from the lab to the summary rather than inside the science. Hassinger named the pressure that drives it, and he was unusually direct about where the gap opens.</span></p><p><span>&#8220;Since at least the beginning of the Q2B conferences put on by QCWare, there has been a recurring chorus demanding to know what are quantum computing&#8217;s use cases, how will it be useful for enterprises,&#8221; he told us. &#8220;Marketing can be tempted to take speculative ideas and present them as certainties, stretching the truth about a scientist&#8217;s speculation to reframe it as definitive. The other question is always when, so timelines are also something that marketing can take liberties with.&#8221;</span></p><p><span>Both distortions are directional, which makes them correctable. A researcher&#8217;s conditional loses its condition, and a scientific dependency acquires a date. Reading a result back through those two transformations usually recovers something close to the original claim.</span></p><h3><span>Roadmaps model engineering determinism onto unsolved physics</span></h3><p><span>The deeper issue Hassinger identifies is a category problem. A roadmap projects milestones one year out, three years, five years, and he is blunt about what kind of document that is. &#8220;That&#8217;s an engineering document,&#8221; he says, &#8220;and engineering is much more deterministic than the underlying scientific breakthroughs that are required to enable the engineering to deliver those milestones.&#8221;</span></p><p><span>Transduction is the concrete case. Superconducting qubits operate inside a dilution refrigerator near absolute zero, and a refrigerator has finite volume, so scaling past one fridge means entangling qubits across separate cryostats. That requires converting the quantum state to a photonic frequency used in telecom, carrying it over fiber as what the field calls flying qubits, then converting back at the far end. None of the known conversion methods delivers the fidelity a reliable device needs, and nobody yet knows what closing that gap requires. &#8220;It&#8217;s not just hard work,&#8221; he says of that class of problem. &#8220;It&#8217;s a lot of hard work, but it&#8217;s also luck, because we don&#8217;t know what we don&#8217;t know.&#8221;</span></p><p><span>This is the reason he treats specifications and milestone dates as the least informative part of any hardware program, and the unsolved science underneath as the part that determines whether the dates mean anything. It is also why his sharpest formulation of the field&#8217;s position lands where it does. &#8220;A qubit is a very interesting device with no intrinsic commercial value,&#8221; he says, and converting it into something useful still depends on physics nobody has finished.</span></p><h3><span>Classical simulability is the only threshold that changes anything</span></h3><p><span>Hassinger&#8217;s position does not end in skepticism, because the field is converging on one measurable target regardless of which architecture reaches it. &#8220;The consensus is we need to deliver fault tolerant logical qubits at a scale that is not simulatable by a classical computer,&#8221; he says. &#8220;That&#8217;s the North Star we&#8217;re all sailing towards.&#8221; Once a system passes the point where your laptop or your GPU cluster can reproduce its output, running it on quantum hardware becomes necessary rather than interesting, and nothing before that crossing changes what you can compute.</span></p><p><span>That threshold also tells you where the physics pays off first, and his answer is narrower than the general coverage implies. Materials science arrives first because condensed matter behaviour maps naturally onto these systems, with small molecule chemistry close behind, while</span><a href="https://deepengineering.net/p/materials-science-first-real-quantum-value"><span> optimization, cryptography, and machine learning all wait on thousands of logical qubits</span></a><span>.</span></p><p><span>So the technical reading is straightforward. When a new result publishes, work out its encoding ratio and its code distance before you compare it to anything, since those two numbers determine what the machine tolerates and the logical qubit count does not. And when a roadmap updates, separate the engineering milestones from the scientific dependencies underneath them and check which unsolved physics the far dates rest on. Then put the effort into</span><a href="https://deepengineering.net/p/boolean-thinking-barrier-quantum-adoption"><span> building quantum intuition inside your own team</span></a><span>, because recognizing the high dimensional structure in your own problems transfers whichever architecture crosses the threshold first.</span></p><div><hr></div><h2>In case you missed</h2><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;8bf6bfb3-d6d4-4a75-970d-968c07e34b7d&quot;,&quot;caption&quot;:&quot;How to tell genuine quantum progress from hype, why qubit count misleads, where the technology delivers value first, and how a classical developer starts.<br />&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Quantum Computing Beyond the Hype with Sebastian Hassinger&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:535827458,&quot;name&quot;:&quot;Sebastian Hassinger&quot;,&quot;bio&quot;:&quot;Author of The New Quantum Era book and host of The New Quantum Era podcast.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4a70390e-d1ad-4e61-8e13-23bc47b2a921_144x144.png&quot;,&quot;is_guest&quot;:true,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-05T19:56:08.385Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8d9f9908-0a93-4cd0-8e54-31f3e5f29800_1920x1080.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.net/p/quantum-computing-beyond-the-hype-sebastian-hassinger&quot;,&quot;section_name&quot;:&quot;Interviews&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:209966044,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:1,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div><hr></div><h2>&#128736;&#65039; Tool of the Week</h2><p><a href="https://github.com/quantumlib/Stim"><span>Stim</span></a><span> is an open source stabilizer circuit simulator maintained under Google&#8217;s quantumlib organisation, built for analysing quantum error correction circuits at speed.</span></p><p><strong><span>Highlights</span></strong></p><ul><li><p><span>Derives a circuit&#8217;s actual code distance instead of relying on the number a vendor publishes.</span></p></li><li><p><span>Turns a noisy circuit into a detector error model ready to configure matching-based decoders.</span></p></li><li><p><span>Samples circuits with thousands of qubits and millions of operations at kilohertz rates.</span></p></li><li><p><span>Installs as a Python package and also runs as a C++ library or a command line tool.</span></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://github.com/quantumlib/Stim&quot;,&quot;text&quot;:&quot;Learn more about Stim&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://github.com/quantumlib/Stim"><span>Learn more about Stim</span></a></p><div><hr></div><h2><strong>&#128206; Tech Briefs</strong></h2><ul><li><p><a href="https://claude.com/blog/claude-enterprise-inference-hooks"><span>Anthropic ships inference hooks for Claude Enterprise</span></a><span> - Governed prompts now route to customer security servers before inference, centralizing DLP across Claude Enterprise surfaces.</span></p></li><li><p><a href="https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"><span>OpenAI cuts GPT-5.6 prices and adds Fast mode</span></a><span> - Luna and Terra get lower API pricing, while Sol Fast mode offers 2.5&#215; speed at 2&#215; cost.</span></p></li><li><p><a href="https://www.dwavequantum.com/company/newsroom/press-release/d-wave-and-nasdaq-verafin-announce-agreement-for-quantum-computing-application-development/"><span>D-Wave and Nasdaq Verafin agree a quantum proof of concept for financial crime detection</span></a><span> - Nasdaq Verafin will test D-Wave quantum-hybrid workflows on hundreds of financial-crime signals and network relationships.</span></p></li><li><p><a href="https://docs.cloud.google.com/sql/docs/postgres/release-notes"><span>Cloud SQL makes PSC reconciliation default</span></a><span> - New or newly enabled PSC instances now close existing connections after project removal from allowed lists.</span></p></li><li><p><a href="https://www.paloaltonetworks.com/blog/2026/08/prisma-airs-unified-data-protection-for-claude/"><span>Palo Alto Networks integrates Prisma AIRS with Claude Enterprise</span></a><span> - </span>Prisma AIRS can inspect Claude prompts before inference, applying existing DLP policies across Claude Enterprise surfaces.</p></li></ul><div><hr></div><div class="callout-block" data-callout="true"><p><strong>&#128227; Contribute to Deep Engineering</strong></p><p><strong>Pitch</strong><span> a </span><a href="https://deepengineering.net/s/practical-deep-dives">practical deep dive</a><span> under your </span><strong>byline</strong><span>. Or if you lead a team, we would like to </span><strong>interview</strong><span> you and build an </span><a href="https://deepengineering.net/s/engineering-leadership">engineering leadership</a><span> feature around your </span><strong>story</strong><span>.</span><br><br><strong>Subscribe</strong><span> to </span><strong>Deep Engineering</strong><span> newsletter and </span><strong>message</strong><span> us through the </span><strong>chat option</strong><span>, or email us at </span><strong>saqibj @ packt.com</strong><span>.</span></p></div><div><hr></div><p><span>That&#8217;s all for today. Thank you for reading this issue of Deep Engineering.</span></p><p><span>We&#8217;ll be back next week with more expert-led content.</span></p><p><span>Keep building,</span></p><p><span>Saqib Jan</span></p><p><span>Editor-in-Chief, Deep Engineering</span></p><div><hr></div><p><em><span>If your company wants to reach senior developers, software engineers, and technical decision-makers, </span><a href="https://packt.omeclk.com/portal/wts/uc%5EcnN2dfNaqmD-kB-mo66%7C7g%5Ef%7Cb"><span>speak to us about partnering</span></a><span> with Deep Engineering.</span></em></p>]]></content:encoded></item><item><title><![CDATA[Materials Science Gets Real Value From Quantum Before Anything Else]]></title><description><![CDATA[Why materials science and small molecule chemistry reach real quantum value first, what higher fidelity simulation changes for drug discovery, and why optimization and cryptography wait.]]></description><link>https://deepengineering.net/p/materials-science-first-real-quantum-value</link><guid isPermaLink="false">https://deepengineering.net/p/materials-science-first-real-quantum-value</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 06 Aug 2026 13:48:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b91e798a-acfd-477e-94d1-3eff12314f55_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><em><span>By </span><a href="https://www.linkedin.com/in/shassinger"><span>Sebastian Hassinger</span></a><span>, author of</span><a href="https://www.packtpub.com/en-it/product/the-new-quantum-era-9781807787370"><span> </span></a><a href="https://www.packtpub.com/en-us/product/the-new-quantum-era-9781807787370"><span>The New Quantum Era</span></a><span> and former quantum lead at </span><strong><span>AWS</span></strong><span> and </span><strong><span>IBM</span></strong><span> | This piece is adapted from his live Deep Engineering session, </span><a href="https://deepengineering.net/p/quantum-computing-beyond-the-hype-sebastian-hassinger"><span>Quantum Computing Beyond the Hype</span></a><span>. Edited by</span><a href="https://substack.com/@saqibjan"><span> Saqib Jan</span></a></em></p></blockquote><p><span>People list the same five domains whenever they ask where quantum computing will actually pay off, so chemistry, materials, optimization, cryptography, and machine learning. </span><strong><span>Of those, materials science is closest to real value, with small molecule chemistry a close second</span></strong><span>, and small molecule chemistry is sometimes treated as a type of materials science anyway. Optimization, cryptography, and machine learning all need thousands of logical qubits, so those are considerably further off.</span></p><p><span>The reason materials leads is that materials science is condensed matter physics. It is about how atoms pack together into a lattice, into a crystalline structure, and how the material behaves as a result of that packing. All of that comes down to physics calculations, which is precisely the kind of work a quantum computer should do well. There is already interesting research going on around battery technology, using quantum information and quantum computing to help with designing and inventing new battery chemistries.</span></p><p><span>What sits behind that work is more ambitious. The aspiration is that if we can simulate materials precisely enough at sufficient scale, we can start creating designer materials. Maybe something much lighter for the same strength, or much stronger for the same weight. Maybe photosynthesis built into the material itself, so you coat your car in something that generates electricity without any cells on the outside. The imagination gets stimulated by the idea of engineering the attributes of a new material at atomic scale, and that is one of the genuinely underestimated parts of this whole story.</span></p><p><span>Chemistry follows for the same underlying reason, because chemistry is quantum mechanical at its core. It is the way atoms interact with one another, and that interaction is a quantum mechanical phenomenon happening at scale, which is what makes it so difficult to simulate precisely. People talk about small molecule chemistry as the next frontier beyond materials, and there is a natural leap from small molecule chemistry to pharmaceuticals, which tend to be larger and more complex molecules. You need a bigger machine to simulate them.</span></p><p><span>Provided we can build one large enough, you can easily imagine drug discovery and research happening inside a very high fidelity simulation, with all the acceleration we are used to getting from computer simulation. The pipeline could compress quite a bit, because you might screen a whole set of drug candidates in simulation before you ever formulate anything, and by the time you are making the drug in the real world you already know how it will interact with the human system.</span></p><p><span>The word doing the work in all of this is </span><strong><span>fidelity</span></strong><span>. Classical simulation of chemistry is approximate by necessity. Methods like</span><a href="https://en.wikipedia.org/wiki/Density_matrix_renormalization_group"><span> DMRG</span></a><span> throw away a great deal of information about the reaction because there is simply too much to calculate, so we have devised ways to focus algorithmically on the small slice of the problem we hope matters most to the answer. The promise of quantum computing is not throwing away as much, and getting a simulation that captures more of the real dynamics of the system accurately.</span></p><p><span>So if you are working out where to point attention, watch the physical sciences rather than your own industry for the first real result. And treat any near term claim about quantum optimization, quantum machine learning, or breaking encryption with the logical qubit count in mind, because those applications are waiting on hardware that does not exist yet.</span></p><div><hr></div><p><strong><span>Read the full issue</span></strong></p><p><span>This piece comes from a longer conversation on how to read quantum progress honestly. The complete interview and the rest of this week&#8217;s </span><strong><span>Deep Engineering</span></strong><span> issue are</span><a href="https://deepengineering.net/s/newsletter-issues"><span> available here</span></a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Eighty Years of Boolean Thinking Is the Real Barrier to Quantum Adoption]]></title><description><![CDATA[Why building quantum intuition inside an engineering organization takes years, what the JPMorgan research model gets right, and where to put effort before the hardware matures.]]></description><link>https://deepengineering.net/p/boolean-thinking-barrier-quantum-adoption</link><guid isPermaLink="false">https://deepengineering.net/p/boolean-thinking-barrier-quantum-adoption</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 06 Aug 2026 13:42:31 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/db4c1548-4f53-4c80-bf42-35e9e3cb67c4_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><em><span>By </span><a href="https://www.linkedin.com/in/shassinger"><span>Sebastian Hassinger</span></a><span>, author of</span><a href="https://www.packtpub.com/en-it/product/the-new-quantum-era-9781807787370"><span> </span></a><a href="https://www.packtpub.com/en-us/product/the-new-quantum-era-9781807787370"><span>The New Quantum Era</span></a><span> and former quantum lead at </span><strong><span>AWS</span></strong><span> and </span><strong><span>IBM</span></strong><span> | This piece is adapted from his live Deep Engineering session, </span><a href="https://deepengineering.net/p/quantum-computing-beyond-the-hype-sebastian-hassinger"><span>Quantum Computing Beyond the Hype</span></a><span>. Edited by</span><a href="https://substack.com/@saqibjan"><span> Saqib Jan</span></a></em></p></blockquote><p><span>It is smart for enterprises to start investing in the skills you need for understanding quantum information, and in exploring potential algorithms now, even though the computers that could actually run anything useful do not exist yet. That sounds premature until you look at what the work actually involves, because it takes a great deal of effort to recast a business problem into quantum terms.</span></p><p><span>We have been thinking about these problems in classical terms, in Boolean logic terms, since the middle of the last century. That is seventy or eighty years now, and it will be a hundred before too long. So we are carrying a very deeply ingrained set of preconceptions about problem solving. We look at our world through the lens of how our laptop could fix the problem in front of us. We write code on those machines, and our minds run along the lines of decomposing the problem, working out what the dynamics are, and figuring out how to represent them efficiently. All of that rests on an invisible reliance on the assumption that at the ground level you are doing Boolean algebra to solve the problem.</span></p><p><span>I do not think many people are fully aware of how much of their problem solving thinking is rooted in the way classical computing works. And it is so different in quantum computing that </span><strong><span>the gap becomes the real obstacle</span></strong><span>. People talk about developing </span><strong><span>quantum intuition</span></strong><span>, which means building the habit of looking at the world through the lens of linear algebra and through the dynamics and capabilities of quantum information. That takes a lot of work, and it is not the kind of work you can compress once the hardware arrives.</span></p><p><span>The smartest enterprise approaches to quantum computing I have seen understand this. They hire a small number of strong people and run research alongside quantum hardware companies and academic researchers at the top of the field. JPMorgan does this well. You can imagine an organization that size has a very large set of challenges, algorithms, and tasks it has to carry out to operate as a financial entity, and the team there looks at theoretical problems that have some mapping to an aspect of those business processes, then does exploratory research against them. They</span><a href="https://www.jpmorganchase.com/about/technology/research/applied-research"><span> publish open science papers</span></a><span>, so everybody benefits from the work they put in.</span></p><p><span>What they are really doing is building the muscle. When quantum technologies mature to sufficient scale, JPMorgan will know how to use them, because the people there will already have years of thinking in the right terms behind them. </span><strong><span>That head start cannot be bought later.</span></strong><span> The hardware will become available to everyone at roughly the same moment, and the differentiator will be whether your organization has anyone who can look at a business problem and recognize the high dimensional structure inside it.</span></p><p><span>Two things follow from this if you are deciding where to put effort right now. Pick one or two people who are genuinely curious and give them real time to build quantum intuition, rather than sending the whole team to an introductory session that changes nothing. And point them at a problem you actually have, something with heavily interconnected variables, so the learning attaches to your business instead of staying abstract.</span></p><div><hr></div><p><strong><span>Read the full issue</span></strong></p><p><span>This piece comes from a longer conversation on how to read quantum progress honestly. The complete interview and the rest of this week&#8217;s </span><strong><span>Deep Engineering</span></strong><span> issue are</span><a href="https://deepengineering.net/s/newsletter-issues"><span> available here</span></a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Quantum Computing Beyond the Hype with Sebastian Hassinger]]></title><description><![CDATA[How to tell genuine quantum progress from hype, why qubit count misleads, where the technology delivers value first, and how a classical developer starts.]]></description><link>https://deepengineering.net/p/quantum-computing-beyond-the-hype-sebastian-hassinger</link><guid isPermaLink="false">https://deepengineering.net/p/quantum-computing-beyond-the-hype-sebastian-hassinger</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Wed, 05 Aug 2026 19:56:08 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8d9f9908-0a93-4cd0-8e54-31f3e5f29800_1920x1080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="https://www.linkedin.com/in/shassinger">Sebastian Hassinger</a> has followed quantum computing from research curiosity toward real machines from inside IBM and AWS, and he wrote <a href="https://www.packtpub.com/en-us/product/the-new-quantum-era-9781807787370">The New Quantum Era</a> to explain the field to people without a physics degree. We talked about how to tell genuine progress from marketing, where the technology earns its place first, and what getting ready for it actually asks of an engineering team.</p><div id="youtube2-133jkU_Qu5I" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;133jkU_Qu5I&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/133jkU_Qu5I?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div><hr></div><blockquote><p><em>This session was recorded live as part of the Deep Engineering Live Interview Series. The transcript below has been lightly edited for clarity and readability. Audience members joined the conversation and asked questions directly during the session.</em></p></blockquote><p><em><strong>Tell us how you came to quantum computing, and where the field stands right now.</strong></em></p><p>My career has been in emerging tech going back to the early nineties, from the early Internet through the arrival of the web. I built the first Apple support site on the web, started a couple of Internet service providers, and worked on innovation inside large companies like IBM and Apple, alongside advising and co-founding startups. So I have moved from one emerging technology to the next for most of my working life.</p><p>Quantum found me in 2017. I had been asked to help IBM with the open source strategy for Qiskit, and my first move was to attend the Think Summit at the T.J. Watson Research Lab, where the team was talking about the 53 qubit machine they would launch the following year. It was about ninety five percent incomprehensible to me then, but I could tell it was novel enough and early enough that I would not need to look for another emerging field for the rest of my career. I have always been drawn to quantum physics even without a physics background, so it became an obsession, and within a year I was working on the IBM Quantum team.</p><p>On where things stand, the theory goes back to the early eighties, but physical qubits and real proof that you could compute with them did not arrive until around 2000. So this is roughly twenty five years of work, and we are now close to machines that do things classical computers genuinely cannot. Everything so far has been experimental, and your laptop can still simulate a quantum computer better than a current quantum computer can. That is changing fast, and before the end of the decade we should see machines at a scale where they do meaningful work you cannot do classically. That shift drives most of the recent interest.</p><p><em><strong>Strip away the metaphors. What is a quantum computer actually doing that a classical machine cannot?</strong></em></p><p>A lot of the metaphors people use are misleading, so let me try two framings that work for me. A classical computer is built by etching circuits into silicon that switch on and off to represent Boolean logic, and Boolean logic underpins everything it does. A qubit instead uses a two level system from nature, an atom, a superconducting circuit, or a photon, something with two energy levels you can control and read as zero and one. The difference is that those states can exist in superposition, where the value behaves like a wave function sitting probabilistically between zero and one, so a single qubit acts as a vector. Once you have many of them you are doing linear algebra rather than Boolean logic, and linear algebra is very good at high dimensional problems where the variables are heavily interconnected.</p><p>The traveling salesman problem is the usual illustration, even if physicists quibble with it. Finding the optimal route through every city is combinatorial, and in a classical machine each new city roughly doubles the combinations you have to test. With qubits you effectively add one qubit to expand the computational space, because the qubits become the exponent for that space rather than a linear count. That is what people mean when they say it would take more bits than there are atoms in the universe to represent a few hundred logical qubits.</p><p>The second framing is simpler. You are using quantum systems from nature to simulate other systems from nature. Feynman made this point in his 1981 keynote that many treat as the starting gun for the field, that simulating nature is inherently hard because of the exponential complexity in these many body systems, and that a computer built from the same kind of system should handle the problem far better. The pithy version is that if you want to simulate a system from nature, you need a system from nature.</p><p><em><strong>Superposition and entanglement get used loosely. Which one do people most often misunderstand, and what is the right intuition?</strong></em></p><p>The biggest issue is more an ambiguity than a flat error. People often say a quantum computer tests every solution at once and then selects the right one. That is not literally correct, and the confusion usually traces back to the most famous algorithm in the field, Shor&#8217;s algorithm, which Peter Shor discovered in the early nineties. Shor&#8217;s algorithm is the one that would let us break RSA and other asymmetric encryption, because it factors a very large number down to its primes, and that factoring is the basis for most modern encryption.</p><p>It works through quantum phase estimation. Inside the very large computational space created by many entangled qubits, often called a Hilbert space that scales as two to the power of the number of qubits, the algorithm deliberately creates interference patterns. The wave functions interfere so the correct answer gets amplified and the wrong answers get suppressed. That resembles trying everything at once, which is why the shorthand persists. I once asked Peter how he felt about people describing it that way, and after thinking about it he said, well, it is not really wrong. With anything quantum there is rarely a clean black and white, because this behavior lives so far from our lived experience that any physical or human metaphor introduces some distortion. The communication challenge of grappling with these foreign ideas is part of what makes the field fascinating to me.</p><p><em><strong>Where does that leave quantum computing for AI and machine learning, and why do we know so few quantum algorithms?</strong></em></p><p>There are a few algorithms with a proven theoretical advantage, HHL among them, plus variations that fan out from that small set. I used to find the short list worrying, but I no longer do, and the reason is the history of classical computing. When these machines were being built in the mid twentieth century, nobody knew what they would be good for, and they did not know for a long time. Von Neumann led a team building a machine at the Institute for Advanced Studies in Princeton around the same time the ENIAC was being built at Penn, both driven by the difficulty of the physics calculations that the Manhattan Project and other wartime work demanded.</p><p>A mathematician named Stan Ulam, working with von Neumann, came up with a sampling method for calculating neutron diffusion, and he called it the Monte Carlo algorithm. It was more than thirty years before someone thought to use Monte Carlo to optimize a portfolio. The technique existed for decades before anyone saw its value outside the original problem. I expect the same pattern with quantum computing, hopefully faster. These machines will let physicists and chemists simulate many body systems and run experiments they cannot run classically, and the techniques they invent for their narrow problems will very likely carry unexpected value into other domains. So the real work for the rest of us is to watch the early users closely and look for techniques we can generalize into algorithms with value in other industries.</p><p><em><strong>Where does quantum genuinely have an advantage, and where will classical computing stay ahead for the foreseeable future?</strong></em></p><p>It helps to start by correcting an instinct. People arriving at the topic assume quantum computers are faster, or built for bigger data. They are actually slower machines working with smaller data. Qubit counts are in the hundreds now, and we hope to reach thousands, maybe tens of thousands. Qubit count is effectively the register size of the machine. QuEra&#8217;s Aquila, a 256 qubit neutral atom device, can load 256 bits of information at a time, so on its own that does not sound impressive. The advantage appears when you entangle those qubits, because you have created a many body system whose state vector is two to the 256, which is more states than there are atoms in the universe. You cannot physically recreate that computational space classically.</p><p>So the rule of thumb is that quantum has a natural advantage when the problem is high dimensional and heavily interconnected, because vectors and linear algebra represent that kind of space far more efficiently than Boolean logic can. Those problems are abundant in nature. Material science is condensed matter physics, how atoms pack into a lattice and how the resulting material behaves, and that is exactly the kind of simulation quantum computers should do well. There is promising work already on battery design using quantum approaches, and a real aspiration that if we can simulate materials precisely enough we can design new ones, much lighter for the same strength, or with properties like photosynthesis built into the material itself.</p><p>Chemistry is quantum mechanical at its core, because it is about how atoms interact, so it is also very hard to simulate precisely. People point to small molecule chemistry as the next frontier beyond materials, and there is a natural leap from there to pharmaceuticals, which are larger and more complex molecules that need bigger machines. If we can build a quantum computer large enough to simulate how a drug candidate behaves and interacts with its target in the body, drug discovery could compress a great deal, because you screen candidates in high fidelity simulation before you ever formulate them. Beyond the physical sciences people talk about optimization and the Shor class of cryptography algorithms, but those are the small number we already know. The rest will fall out of using machines we cannot simulate classically, which is a chicken and egg situation. We know the physical sciences will benefit, and the broader applications in finance or logistics may surprise us, because we cannot imagine them until the machine exists at that scale.</p><p><em><strong>Enterprises often assume quantum will speed up their existing workloads. For which problems is that assumption simply wrong?</strong></em></p><p>The easiest answer is that if your problem is processing large volumes of data quickly, a quantum computer will not do that in the near term, and possibly never, because classical computing keeps advancing too. I avoid saying never about any of this, but that is a game classical may always lead. The clearer way to see the boundary is through fidelity. In classical chemistry simulation we use methods like DMRG that deliberately throw away much of the information about a reaction because there is too much to calculate, so we focus algorithmically on the small part we hope matters most. The promise of quantum is a much higher fidelity simulation that keeps more of the real dynamics. So the enterprise question to ask is whether your use case involves genuinely high dimensional, highly interconnected data. Optimization is one area with potential, because the more interconnected the parameters, the higher dimensional the problem and the harder it is combinatorially, and a larger quantum computer could represent that more efficiently than classical approaches like max cut.</p><p>Even so, it is smart for enterprises to invest now in the skills to understand quantum information and explore algorithms, even though no machine can run anything useful yet. Recasting a business problem takes real effort, because we have thought in classical, Boolean terms since the middle of the last century, close to eighty years. That is a deeply ingrained set of assumptions. We look at the world through the lens of what a laptop can do, and most people are not aware how much of their problem solving quietly assumes Boolean algebra at the ground level. Building quantum intuition, learning to see problems through linear algebra, takes a lot of work. The smartest enterprise approaches I have seen hire a small number of strong people and run research with hardware companies and academics. The team at JPMorgan, for instance, studies theoretical problems that map to aspects of their business, publishes open science, and builds the muscle to apply quantum technologies once the machines mature enough to matter.</p><p><em><strong>Give us an honest read on the hardware. How far are we from machines that do useful work beyond what classical systems already handle?</strong></em></p><p>We are on the cusp, and before the end of the decade we should see machines doing meaningful work beyond classical. But there are still deep scientific unknowns across every modality the vendors are pursuing. It is like the early days of classical computing, when the question was not only how to fit more transistors on a chip but which materials and fabrication processes would even work. Those were genuine unknowns the industry tackled over years, and only in aggregate did it look like a smooth curve. We are at that very early stage in a lot of ways.</p><p>Transduction is a concrete example. With superconducting qubits you keep them in a dilution refrigerator near absolute zero, and a fridge has limited space, so to scale you have to connect the qubits at the bottom of one fridge to those in another. That means converting them to a photonic frequency used in telecom, carrying the quantum value over fiber, then converting back to the native frequency in the second fridge. None of the ways we currently know how to do that delivers the fidelity the device needs to operate reliably, and we do not yet know what it will take to fix it. That is a scientific challenge, not just hard engineering, and scientific challenges involve luck, because you do not know what you do not know.</p><p>This is why roadmaps are difficult to read. A roadmap is an engineering document projecting deterministic milestones, one year out, three years, five years. Engineering is far more deterministic than the scientific breakthroughs required to enable it, so when I talk to investors I tell them to do deep due diligence on the scientific challenges a hardware company still faces. For an end user the specifics of any single vendor almost do not matter. What matters is the North Star the whole industry sails toward, fault tolerant logical qubits at a scale you cannot simulate on a laptop or a GPU cluster. Once you cannot simulate it classically, you have to run it on a quantum computer, and that is the only line that counts.</p><p><em><strong>Error correction keeps getting described as the bottleneck. What actually changed in the last year, and what has not?</strong></em></p><p>Start with what qubit count actually tells you, which is less than people think, because it is effectively register size, not a measure of progress. We have had machines with thousands of qubits before. At IBM the roadmap produced a chip called Condor at just over a thousand qubits, and it had real value for the R&amp;D effort of designing and fabricating it, but the connectivity was so low and the noise so hard to manage that even researchers preferred going back to the smaller machines around 127 to 133 qubits, which were more capable. So raw qubit number is not an indicator of progress toward usefulness.</p><p>What matters is fidelity and noise, and a good proxy is the resilience of the error correction code, which you talk about as a distance or a d value. That is roughly how many errors the system can absorb before the information collapses and you lose the computation, so a higher distance means a more resilient system. The Willow experiment about a year and a half ago was distance seven, if I remember, roughly a hundred physical qubits in a surface code presenting as one logical qubit. That is the useful snapshot, because as an end user you do not need the underlying gate fidelities, you need to know how many logical qubits you get and how resilient the error correction is. What has not changed is that fundamental science is still involved, so anytime you hear someone say quantum is just a matter of engineering now, be suspicious of their claims.</p><p><em><strong>What does getting quantum ready actually mean, and how does a classical developer start?</strong></em></p><p>Quantum ready means different things by context, but broadly it means developing some intuition for what a quantum information approach to a problem looks like. For an existing software developer the best way in is usually coding, and the tools are familiar. Most quantum programming is done in Python. Qiskit is IBM&#8217;s Python SDK, and Amazon Braket, which I worked on at AWS, is also Python, along with several others in well understood languages. The way you construct a task and send it to a quantum computer uses tools we already know. The hard part is the logic of the circuit itself, so there are gentle entry points to start building familiarity.</p><p>One of my favorites is the Unitary Foundation, where I am a fellow. Every year they run Unitary Hack, a global event over a couple of weeks where maintainers of open source quantum software tag issues in their repositories and developers close them for bounties. The issues are often housekeeping, security, or maintainability work that any classical developer recognizes, and as a side effect you see how a quantum software package works inside. My favorite example is an engineer named Misty Wall. She was a mechanical engineer at ASML, finished a major project, wanted something new, and got interested in quantum without any background in quantum information. She started closing tickets in Mitiq, an open source error mitigation framework, and two or three years later she was lead author on research papers on quantum error mitigation. It does not happen for everyone, but it is a real path.</p><p>Unitary Foundation runs a Discord year round, and the popular packages each have a channel, so you can see what is out there, including open source simulation packages. Building a simulator to run a quantum circuit classically is a very good way to understand how circuits and algorithms are actually represented. For a developer this is the best way in, because the alternative is a physics PhD and a job in a lab, which is a long and demanding path. The good news is that other people&#8217;s hard work now lets us reach real quantum computers over the Internet with tools we already know.</p><div><hr></div><h4>Follow-up questions over email</h4><p>After the session, Sebastian answered two questions we did not reach live.</p><p><strong>Of the domains people cite, chemistry, materials, optimization, cryptography, machine learning, which is closest to real value and which is furthest away?</strong></p><p>Probably materials, with small molecule chemistry, which is sometimes treated as a type of material science, a close second. Optimization, cryptography, and machine learning all need thousands of logical qubits, so those are further off.</p><p><strong>You have been on the commercial side at AWS and IBM. Where does the marketing most often outrun the engineering?</strong></p><p>Since at least the early Q2B conferences run by QC Ware, there has been a recurring chorus demanding to know quantum computing&#8217;s use cases and how it will be useful for enterprises. Marketing can be tempted to take speculative ideas and present them as certainties, stretching a scientist&#8217;s speculation into something definitive. The other constant question is when, so timelines are another place marketing takes liberties.</p><div><hr></div><p><a href="https://www.linkedin.com/in/shassinger">Sebastian Hassinger </a>writes and hosts the New Quantum Era podcast, and his book <a href="https://www.packtpub.com/en-us/product/the-new-quantum-era-9781807787370">The New Quantum Era</a> explains the field for readers without a physics background. Find him on LinkedIn.</p>]]></content:encoded></item><item><title><![CDATA[Python Developers Can Learn Quantum Computing by Fixing Bugs]]></title><description><![CDATA[How a working software developer gets into quantum computing through Python SDKs like Qiskit and Braket, open source contribution, and error mitigation frameworks.]]></description><link>https://deepengineering.net/p/python-developers-can-learn-quantum-computing-by-fixing-bugs</link><guid isPermaLink="false">https://deepengineering.net/p/python-developers-can-learn-quantum-computing-by-fixing-bugs</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Wed, 05 Aug 2026 16:30:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/4a0e8d0c-4e4a-42a8-9bfa-f8c17f8a8f53_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p>This piece is adapted from our <a href="https://deepengineering.net/p/quantum-computing-beyond-the-hype-sebastian-hassinger">live interview session</a> with <a href="https://www.linkedin.com/in/shassinger">Sebastian Hassinger</a>, author of <a href="https://www.packtpub.com/en-us/product/the-new-quantum-era-9781807787370">The New Quantum Era</a> and former quantum lead at <strong>AWS</strong> and <strong>IBM</strong>, which anchors the <a href="https://deepengineering.net/p/issue-58-sebastian-hassinger-qubit-counts">Deep Engineering newsletter issue on quantum progress</a>. The words below are his.</p></blockquote><p>So getting quantum ready means different things depending on who is asking, but for a working software developer it mostly means building some intuition for what a quantum information approach to a problem actually looks like. And the best way to get there is usually through coding rather than through reading theory first.</p><h2><strong>The tools are already familiar</strong></h2><p>Most of the programming that goes on in quantum computing happens in Python. Qiskit, the <a href="https://www.ibm.com/quantum">IBM SDK</a>, is a Python SDK. <a href="https://aws.amazon.com/braket/">Amazon Braket</a>, which I worked on during my years at AWS, is also Python. There are a number of others in Python or in other well understood languages. The way you construct a task and send it to a quantum computer uses tools we all already know, and the genuinely difficult part is the logic of the circuit itself, not the surrounding machinery. So there are plenty of easy entry points for making yourself familiar with the field.</p><h2><strong>Start by closing tickets</strong></h2><p>My favorite route for a developer runs through the <a href="https://unitary.foundation/">Unitary Foundation</a>, where I am a fellow. Every year they run an event called <a href="https://unitaryhack.dev/">Unitary Hack</a>, a global couple of weeks that has been remote and is starting to become an in person thing as well. Maintainers of open source quantum software repositories tag issues in their repos, and developers close those issues out and claim bounties for the ones they resolve. The issues are often housekeeping, administrative work, security tasks, things that are very familiar to classical developers and have to do with the capability, resiliency, and maintainability of the package. But as a side effect you get to see the insides of how that quantum software package actually works.</p><h2><strong>What that path can turn into</strong></h2><p>My favorite example of how this can turn out involves an engineer named Misty Wahl. She worked at ASML as a mechanical engineer, had just finished shipping a major project there, and was looking for something to get involved in. She was interested in quantum computing without any particular background in quantum information. She got involved in Unitary Hack and started closing tickets in <a href="https://github.com/unitaryfoundation/mitiq">Mitiq</a>, an open source error mitigation framework. Error mitigation is almost a building block toward error correction, where you do not have full correction but you are finding ways to tease better signal out of the noise, so Mitiq lets you experiment with different approaches to that. Two or three years later she was lead author on research papers on quantum error mitigation. Her experience turned her into a quantum information researcher.</p><p>That does not happen for everybody, but it is certainly a viable path, and you do not have to wait for the annual event to start. The Unitary Foundation runs a Discord server year round, and all of the popular packages with a lot of users have a channel on it. You can see what packages are out there, including open source simulation packages, which are particularly good ways to develop an understanding of how quantum circuits get represented and what those algorithms look like. You are building a simulator to classically run a quantum circuit, so you end up learning the structure from the inside.</p><h2><strong>Why this beats the academic route</strong></h2><p>For a developer this really is the best way in, because the alternative is getting a PhD in physics and a job in a lab or at IBM or AWS or Google, and that is a very long path that is not for everyone. The field remains extremely challenging to enter as a practitioner. But other people&#8217;s hard work has made it far easier for the rest of us to reach real quantum computers over the Internet using tools we are already comfortable with.</p><h2><strong>Two things to do this week</strong></h2><p>If you want to act on this, pick one open source quantum package and read its issue tracker this week, looking specifically for the maintenance work rather than the physics. And join the Discord for whichever package you choose, because seeing what its actual users are stuck on will teach you more about where the field stands than another explainer will.</p><div><hr></div><p><em>The full conversation goes further, covering why qubit counts mislead and where quantum pays off first, over in <a href="https://deepengineering.net/p/issue-58-sebastian-hassinger-qubit-counts">issue 58 of Deep Engineering</a>.</em></p><p></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;fa3cda55-66aa-45c6-b62a-1b9be3b740c0&quot;,&quot;caption&quot;:&quot;Featured - LangGraph Masterclass: From Beginner to Professional&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #58: Sebastian Hassinger on Where Quantum Progress is Real&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-06T15:45:37.785Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c258bf82-e5d3-4213-9f0b-4beac02d7cec_2400x1600.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.net/p/issue-58-sebastian-hassinger-qubit-counts&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:210080136,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:5,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[Deep Engineering #57: Rory Preddy on Cutting Agent Costs Without Ranking Engineers]]></title><description><![CDATA[Rory Preddy of Microsoft and GitHub argues cost discipline belongs in agent profiles and per-session spend caps, not in per-engineer token leaderboards.]]></description><link>https://deepengineering.net/p/issue-57-cutting-agent-costs-without-ranking-engineers-rory-preddy</link><guid isPermaLink="false">https://deepengineering.net/p/issue-57-cutting-agent-costs-without-ranking-engineers-rory-preddy</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 30 Jul 2026 15:14:18 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/9f94df86-bee9-49db-858e-e5ead0589bba_3200x1800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Featured - <a href="https://www.eventbrite.co.uk/e/engineering-reliable-agentic-ai-systems-tickets-1992373400474?aff=deepeng">Engineering Reliable Agentic AI Systems</a></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.eventbrite.co.uk/e/engineering-reliable-agentic-ai-systems-tickets-1992373400474?aff=deepeng" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!e4Gm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d51f22f-96f5-4eb9-a4a8-174e664205cb_800x267.jpeg 424w, https://substackcdn.com/image/fetch/$s_!e4Gm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d51f22f-96f5-4eb9-a4a8-174e664205cb_800x267.jpeg 848w, https://substackcdn.com/image/fetch/$s_!e4Gm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d51f22f-96f5-4eb9-a4a8-174e664205cb_800x267.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!e4Gm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d51f22f-96f5-4eb9-a4a8-174e664205cb_800x267.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!e4Gm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d51f22f-96f5-4eb9-a4a8-174e664205cb_800x267.jpeg" width="800" height="267" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6d51f22f-96f5-4eb9-a4a8-174e664205cb_800x267.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:267,&quot;width&quot;:800,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.eventbrite.co.uk/e/engineering-reliable-agentic-ai-systems-tickets-1992373400474?aff=deepeng&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!e4Gm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d51f22f-96f5-4eb9-a4a8-174e664205cb_800x267.jpeg 424w, https://substackcdn.com/image/fetch/$s_!e4Gm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d51f22f-96f5-4eb9-a4a8-174e664205cb_800x267.jpeg 848w, https://substackcdn.com/image/fetch/$s_!e4Gm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d51f22f-96f5-4eb9-a4a8-174e664205cb_800x267.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!e4Gm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d51f22f-96f5-4eb9-a4a8-174e664205cb_800x267.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Building an agent is easy. Building one that survives production takes architecture, verification, and stopping conditions that actually stop. Four hands-on hours with <a href="https://www.linkedin.com/in/rickhigh">Rick Hightower</a> on loop engineering, evaluation harnesses, context efficiency, and safe MCP tool integration, live on <strong>29 August</strong>.</p><p style="text-align: center;"><span>Deep Engineering readers save 40% with code - </span><strong>DEEPENG40</strong><span>.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.co.uk/e/engineering-reliable-agentic-ai-systems-tickets-1992373400474?aff=deepeng&quot;,&quot;text&quot;:&quot;Save your seat &#8594;&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.eventbrite.co.uk/e/engineering-reliable-agentic-ai-systems-tickets-1992373400474?aff=deepeng"><span>Save your seat &#8594;</span></a></p><div><hr></div><p><span>&#9997;&#65039; </span><strong><span>From the editor&#8217;s desk,</span></strong></p><p><span>Welcome to the </span><strong><span>57th</span></strong><span> issue of </span><strong><span>Deep Engineering</span></strong><span>!</span></p><p><span>On 28 July, </span><a href="https://github.blog/changelog/2026-07-28-github-copilot-app-usage-metrics-now-expand-across-report-rollups/"><span>GitHub&#8217;s changelog</span></a><span> recorded a change that sounds administrative and is not. Individual Copilot app activity is now attributed to users in the enterprise-user and organization-user reports, with a per-user section reporting session counts, request counts, prompt counts, and a token usage breakdown covering output tokens, prompt tokens, and average tokens per request. Per-engineer token attribution is now a REST call against an API most enterprise administrators already query.</span></p><p><span>Nothing about that is objectionable on its own, because a team that cannot see its consumption cannot govern it. The risk lives in what gets built next, since the shortest path from per-user token data to a management artifact is a ranking. Amazon and Meta both stood up internal leaderboards ranking engineers by tokens burned, then dismantled them within a quarter once engineers began padding their usage, which we covered in </span><a href="https://deepengineering.net/p/special-issue-judgment-not-tokenmaxxing-creates-value"><span>our special issue on token maxxing</span></a><span>. The visibility has now improved considerably, and the temptation has improved with it.</span></p><p><a href="https://za.linkedin.com/in/rorypreddy"><span>Rory Preddy</span></a><span>, AI Advocate in Developer Relations at </span><strong><span>Microsoft</span></strong><span> and </span><strong><span>GitHub</span></strong><span>, worked five years in cloud advocacy watching the same question mature from how much will we spend into how much can we save, and he argues the answer was never individual behaviour. Today&#8217;s issue comes out of </span><a href="https://www.youtube.com/watch?v=tVRcYTX_HCg&amp;feature=youtu.be"><span>our live session</span></a><span>, where his position was that cost discipline belongs in the artifacts a team inherits rather than in anything resembling a performance conversation.</span></p><blockquote><p>You can also read this <a href="https://deepengineering.net/p/token-efficiency-rory-preddy-agent-token-costs">practical deep dive on the mechanics</a> of token efficiency by Preddy.</p></blockquote><p><strong>Let&#8217;s get started.</strong> </p><div class="callout-block" data-callout="true"><h2 style="text-align: center;"><a href="https://www.vpdae.com/redirect/o2rdxoqtpm0p1but31d2wk9ffrb"><span data-color="#f97141" style="color: rgb(249, 113, 65);">Agent-written TLA+</span></a></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.vpdae.com/redirect/o2rdxoqtpm0p1but31d2wk9ffrb" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sw8h!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 424w, https://substackcdn.com/image/fetch/$s_!sw8h!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 848w, https://substackcdn.com/image/fetch/$s_!sw8h!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 1272w, https://substackcdn.com/image/fetch/$s_!sw8h!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sw8h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png" width="296" height="296" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:300,&quot;width&quot;:300,&quot;resizeWidth&quot;:296,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.vpdae.com/redirect/o2rdxoqtpm0p1but31d2wk9ffrb&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sw8h!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 424w, https://substackcdn.com/image/fetch/$s_!sw8h!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 848w, https://substackcdn.com/image/fetch/$s_!sw8h!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 1272w, https://substackcdn.com/image/fetch/$s_!sw8h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f00ac2e-e823-441a-b835-1e729ec944e9_300x300.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;">An agent wrote our <strong>TLA+</strong> spec. The model checker explored <strong>14.3M</strong> states and caught a real race.</p><p style="text-align: center;"></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.vpdae.com/redirect/o2rdxoqtpm0p1but31d2wk9ffrb&quot;,&quot;text&quot;:&quot;Read the write-up&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.vpdae.com/redirect/o2rdxoqtpm0p1but31d2wk9ffrb"><span>Read the write-up</span></a></p><p style="text-align: center;"></p></div><div><hr></div><p><strong>Expert Insight</strong></p><h2><span>Cost discipline belongs in the agent profile, not the performance review</span></h2><p><em>by <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;id&quot;:427210082,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;uuid&quot;:&quot;e7195419-ea93-4f2b-84d4-a7959d36a209&quot;}" data-component-name="MentionToDOM"></span> with <a href="https://za.linkedin.com/in/rorypreddy">Rory Preddy</a></em></p><p><span>The logic behind a token leaderboard feels right from a distance, because agents produce measurable volume, volume varies between engineers, and variance invites comparison.</span><a href="https://za.linkedin.com/in/rorypreddy"><span> Rory Preddy</span></a><span>, AI Advocate in Developer Relations at </span><strong><span>Microsoft</span></strong><span> and </span><strong><span>GitHub</span></strong><span>, has observed this sequence play out once before in cloud, and his position is simply that &#8220;Rather than spend more, spend less.&#8221;</span></p><p><span>Coming from somebody whose job is persuading developers to adopt AI, that is a more interesting claim than it would be from a finance team, and the reasoning underneath it is what makes it useful. Preddy is not arguing for austerity or for slowing teams down. He is arguing that cost discipline is a property of the artifacts a team inherits rather than a property of the people using them, which relocates the problem from individual behavior to engineering infrastructure. For any leader currently deciding what belongs on an AI dashboard, that distinction determines whether the dashboard helps or quietly makes things worse.</span></p><h3><span>Spending less is an engineering target, not a budget instruction</span></h3><p><span>Preddy formed this view during five years as a cloud advocate, in rooms with the CTOs of banks whose balance sheets run larger than some national economies. He describes asking those CTOs what worried them most, and the answer changed shape over time in a way worth attending to. Early on they wanted to know how much they were going to spend, because a four-year hardware forecast let them buy servers and depreciate the cost against a schedule they controlled. Later the question inverted entirely, and what they wanted to know was how much they could save.</span></p><p><span>That inversion is the part that transfers to tokens, and it transfers with one difference that makes it more urgent rather than less. A server that gets over-provisioned wastes capital on a predictable schedule, and the waste appears on a depreciation line somebody reviews. An agent loop that resends its context on every iteration wastes money continuously, invisibly, and at a rate nobody set, because the cost arrives as consumption rather than as a purchase. Cloud waste taught organizations to rightsize compute against a bill they could at least forecast. Token waste removes the forecast.</span></p><p><span>Preddy&#8217;s instruction to teams starting now runs from the floor upward, so his advice is to look for &#8220;what&#8217;s the minimum and then cost save, not maximum.&#8221; The engineering content of that instruction matters more than its thrift. A team that starts at the largest model with the longest context has no baseline against which to measure any later improvement, so it cannot tell whether a change helped. A team that starts at the smallest model completing the work has a floor, a known cost, and a reason to escalate when the work genuinely demands it. The cheap path becomes the measured default rather than the fallback nobody trusts.</span></p><h3><strong><span>Discipline is arriving as a function, not a phase</span></strong></h3><p><span>That parallel with cloud is not incidental to Preddy&#8217;s argument, and he expects the rest of the pattern to follow too. &#8220;There&#8217;s a whole industry out there just for token efficiency,&#8221; he says, pointing to engineers who wrote good deployment scripts, became cloud engineers, and then watched cost optimization separate into its own discipline with its own titles and its own tooling. He puts it more flatly when he calls it &#8220;This is our industry.&#8221;</span></p><p><span>The institutional scaffolding has arrived faster this time than it did for cloud. The FinOps Foundation&#8217;s</span><a href="https://data.finops.org/"><span> State of FinOps 2026</span></a><span> reports that 98 percent of practitioners now manage AI spend, up from 31 percent two years ago, and it ranks FinOps for AI as the top forward-looking priority with AI cost management as the single skillset teams most need to develop. In June the Linux Foundation</span><a href="https://www.linuxfoundation.org/press/linux-foundation-announces-the-intent-to-launch-the-tokenomics-foundation-to-establish-open-standards-for-ai-cost-management"><span> announced its intent to launch the Tokenomics Foundation</span></a><span>, working with the FinOps Foundation on open industry standards, benchmarks, and best practices for the economics of AI infrastructure. Preddy reads that as the point where dashboards, standards, and protection mechanisms stop being individual good habits and become things an organization can be held to.</span></p><p><span>For engineering leaders the timing carries a specific implication. A function forming now will settle its conventions within a few quarters, and the organizations contributing to those conventions will be the ones already treating agent cost as an engineering concern rather than a procurement one. Teams still running enablement sessions about prompt length will inherit whatever standards other people write.</span></p><h3><span>Cost advice without a budget produces a cheaper unbounded system</span></h3><p><span>Before any of that scaffolding helps, a team has to answer a question most cost guidance skips. The sharpest moment in our conversation came when Preddy declined the premise of a question about authorization boundaries and asked something more basic in its place. &#8220;How much money do I have?&#8221; he said, and then named the assumption he keeps encountering, because &#8220;a lot of what you&#8217;re asking me assumes that you have infinite amount of tokens.&#8221;</span></p><p><span>He is identifying a real gap in how the industry discusses this. Almost all published cost guidance optimizes mechanics, so prompts get shorter, retrieval gets tuned, caches get warmed, and routing gets layered in. Very little of it starts by asking what the workload is permitted to cost. Optimization applied to an unbounded system produces a cheaper unbounded system, which still has no ceiling and still fails in the same direction. The budget has to exist before the tuning means anything, because the budget is what converts an optimization into a decision.</span></p><p><span>What makes this harder than it looks is that the system cannot supply the number for you. A study published through Microsoft Research this spring,</span><a href="https://www.microsoft.com/en-us/research/publication/how-do-ai-agents-spend-your-money-analyzing-and-predicting-token-consumption-in-agentic-coding-tasks/"><span> How Do AI Agents Spend Your Money?</span></a><span>, presents the first systematic analysis of token consumption in agentic coding tasks, and it found that runs on the same task can differ by up to 30 times in total tokens, that input rather than output tokens drive the overall cost, and that frontier models fail to predict their own token usage and systematically underestimate what a task will cost. A workload whose executor cannot estimate it will not be estimated reliably by a spreadsheet either. That finding closes off the approach most organizations reach for first, which is forecasting from measured averages, and it pushes the answer toward setting a limit and enforcing it rather than predicting a figure and hoping.</span></p><p><span>The consequences of skipping that step have landed at named companies. Forrester&#8217;s</span><a href="https://www.forrester.com/blogs/ai-cost-management-how-prepared-are-you/"><span> July analysis of AI cost management</span></a><span> cites Uber burning its AI budget in four months, Microsoft ending Claude Code licenses after also burning its yearly AI budget, Tesla limiting AI spending to 200 dollars per week, and Priceline absorbing an unexpected surge in AI development renewal costs. None of those reads as an engineering failure. Each one reads as a missing number.</span></p><h3><span>Foundation is an artifact leaders ship</span></h3><p><span>Setting the number is the first move, and the second is deciding who builds the thing that keeps teams inside it. That is where Preddy&#8217;s position turns into a leadership argument rather than a productivity tip. He reaches for the idea of a cybernetic teammate, an agent working alongside a team rather than replacing anyone, and then reframes it as something an engineering leader produces rather than something a vendor sells. The artifact is the agent profile, along with the instructions file and the container definition that travel with it, and his claim about ownership is direct. &#8220;Your leaders should go in and create the agent profiles,&#8221; he says.</span></p><p><span>His reasoning turns on availability rather than authority, and it is the more humane version of the argument. Not every engineer works somewhere they can walk to a senior colleague and admit they are burning too many tokens, and not every senior colleague is reachable at the moment the question arrives. Preddy frames the profile as the answer to that absence, describing the message it carries as &#8220;I&#8217;m not around to help you out, but I&#8217;ve created a nice agent profile&#8221; with the conventions already encoded. A profile that declares only the tools a task needs, names the model tier, and sets the compaction expectation makes the disciplined path the default path for everybody who inherits it, including the engineer who joined last week and has nobody to ask.</span></p><p><span>The second-order effect is what leaders should weigh most carefully, because it determines which of two very different programs they end up running. Token waste treated as individual behavior produces enablement sessions, then dashboards, then comparisons, then rankings, and each step follows plausibly from the last. Token waste treated as a missing artifact produces a file one person writes once and every engineer benefits from without thinking about it. The first approach asks engineers to carry knowledge that changes every quarter. The second encodes that knowledge where it survives staff turnover and stops depending on who attended which session.</span></p><p><span>One detail in Preddy&#8217;s own workflow reinforces the point. He builds these files using the cheapest model available with reasoning switched off, which means the foundation costs almost nothing to produce. The barrier to doing this properly is not budget or tooling. It is that nobody has been made accountable for it.</span></p><h3><span>A spend cap belongs with the permissions</span></h3><p><span>The foundation shapes the default, and a cap is what holds when the default gets overridden. Engineering organizations already accept that an agent&#8217;s authority needs governing, so they scope which systems it may read, which it may write to, and which credentials it holds. Preddy&#8217;s contribution is to place money in that same category rather than treating it as a separate financial concern reviewed on a different cycle by different people. A per-session budget that warns at a threshold and stops at a limit behaves exactly like any other permission, because it constrains what the agent may do without a human present.</span></p><p><span>The practical consequence is a change in when the decision gets made and by whom. Treated as finance, a spend cap arrives quarterly, applies at the account level, and reaches the engineer as a rate limit they cannot explain. Treated as authorization, it arrives at the same review where the agent&#8217;s other permissions get set, applies at the session level, and reaches the engineer as a known boundary they helped define. Preddy demonstrated the enforced version during</span><a href="https://youtu.be/tVRcYTX_HCg"><span> our session</span></a><span>, where a run halted after crossing its allocated credits and offered the operator a choice to add credits, raise the ceiling, or remove it entirely.</span></p><p><span>That halt is the whole value, and it is worth being precise about why. An agent meeting a wall has produced a decision point, with a human present, a specific task in view, and the cost of continuing visible on screen. An agent with no wall has produced an invoice, arriving weeks later, aggregated across teams, with nobody able to reconstruct which run caused it. The same money moves in both cases. Only one of them leaves an organization able to govern the next run.</span></p><h3><span>Token counts measure activity rather than relief</span></h3><p><span>A budget and a cap tell you when to stop. Neither tells you whether the spending accomplished anything, which is the question a dashboard is supposed to answer. A contributor to</span><a href="https://deepengineering.net/p/special-issue-judgment-not-tokenmaxxing-creates-value"><span> our earlier issue on token maxxing</span></a><span> offered a formulation Preddy endorsed when we put it to him, that token usage behaves like CPU utilization, so the honest unit is the cost of relieving a bottleneck rather than the volume consumed. He agreed, then complicated it usefully by pointing out that the model chosen to relieve the bottleneck changes the answer, which means the metric has to account for the routing decision rather than treating all tokens as equivalent.</span></p><p><span>The research supports him on the underlying point. The same Microsoft Research analysis found that higher token usage does not translate into higher accuracy, and that accuracy often peaks at intermediate cost before saturating as spending rises. A number that rises while the outcome does not is the definition of a metric worth distrusting.</span></p><p><span>Preddy&#8217;s own experience supplies the rest, and it is more persuasive because it is self-implicating. He assigned a merge conflict to the smallest model available, and it consumed 51 million tokens without resolving anything, a run we cover in detail in his</span><a href="https://deepengineering.net/p/token-efficiency-rory-preddy-agent-token-costs"><span> companion deep dive</span></a><span>. Read as a token count, that looks like heavy engagement from a productive engineer. Read as bottleneck relieved, it produced nothing whatsoever. Any metric unable to distinguish between those two readings will reward the wrong behavior from the day somebody starts reporting it upward, and it will do so most strongly for the engineers who least understand what they are doing.</span></p><p><span>This is also why the leaderboard fails on its own terms rather than only on cultural ones. A ranking by tokens consumed can be improved by consuming more tokens, which requires no additional value and no additional skill. A measure built on work closed, so a merge conflict resolved, an accessibility fix shipped, a suite returned to green, can only be improved by closing more work. The second measure is harder to define and harder to collect, which is exactly why organizations reach for the first. Preddy&#8217;s argument is that the difficulty is the job, because the easy number and the useful number have never been further apart than they are now.</span></p><h3><span>The principle holding it together</span></h3><p><span>What makes Preddy&#8217;s approach more than a list of preferences is the single idea running through every part of it, which is that cost control belongs in the system rather than in the person operating it. Every choice he describes follows from that.</span></p><p><span>The foundation is a file rather than a training session, because a file survives turnover and a session does not. The spend cap is a permission rather than a budget line, because a permission gets enforced at the moment of use and a budget line gets reviewed after the fact. The metric is work closed rather than tokens consumed, because an outcome cannot be inflated by burning more of the input. The smallest viable model is the default rather than the exception, because a default is what happens when nobody is paying attention, and most of the cost accrues exactly then.</span></p><p><span>That principle also explains what Preddy leaves out. He never asks engineers to be more careful, and he never suggests awareness alone fixes anything, even while arguing for visibility. His own session history showed input tokens outweighing output roughly 100 to 1, and one session climbing from 49,000 to 174,000 input tokens across 66 turns with no compaction at all. He describes himself falling into bad practices like everybody else. A discipline depending on the operator remembering will fail for the person who wrote the discipline, which is the strongest available argument for encoding it somewhere else.</span></p><h3><span>What this asks of a leader</span></h3><p><span>The sequence that follows is short enough to start this week. Write the agent profile and the instructions file yourself, or make one person accountable for them, and treat them as shared infrastructure reviewed like any other. Set a per-session budget in the same review where the agent&#8217;s permissions get set, rather than in a separate conversation with a different owner. Report the work closed rather than the tokens consumed, and decline to build the leaderboard even though the API now makes one trivial.</span></p><p><span>The deeper point for any leader standing up AI measurement is that visibility and ranking are not the same thing, and the tooling arriving now makes it very easy to confuse them. Per-user token data has legitimate uses, including chargeback, capacity planning, and helping a manager notice that a team is consuming heavily without much to show for it. It becomes destructive at the moment it gets used to compare engineers against each other, because that is the point at which the number starts changing the behavior it was meant to observe.</span></p><p><span>Preddy&#8217;s closing thought aimed at the people doing the work rather than the people measuring it, and it belongs here because burnout and cost discipline turn out to share a cause. His instruction was not to &#8220;end each day exhausted, not from the work itself, but from managing of the work.&#8221; An engineer who inherits a foundation somebody else built gives the day to the problem. An engineer who inherits nothing gives it to managing agents, and pays for the privilege in tokens. &#8220;Be kind to yourself,&#8221; he said, which for a leader translates into something concrete. Build the foundation, and there will be nothing worth ranking.</span></p><div><hr></div><h2>In case you missed</h2><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;e11c2ffb-8202-4741-a576-cefab907eb33&quot;,&quot;caption&quot;:&quot;By Rory Preddy - Cache the prefix, compact the context, route to the smallest model that fits, and put a ceiling on every loop<br />&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;I burned 51 million tokens on one merge conflict, and the model was not the problem&quot;,&quot;publishedBylines&quot;:[],&quot;post_date&quot;:&quot;2026-07-30T12:06:07.062Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fb227e32-e539-4c24-a3da-5b6ba7b68be5_3200x1800.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.net/p/token-efficiency-rory-preddy-agent-token-costs&quot;,&quot;section_name&quot;:&quot;Practical Deep-Dives&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:209102293,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div><hr></div><h2>&#128736;&#65039; Tool of the Week</h2><p><strong><a href="https://github.com/langfuse/langfuse"><span>Langfuse</span></a><span> </span></strong><span>is an open source AI engineering platform that traces agent runs and attributes cost and token usage per trace. A spend cap needs a number behind it, and a per-trace cost is that number.</span></p><ul><li><p><span>Traces multi-step agent runs so cost attaches to a workflow rather than to a raw API call</span></p></li><li><p><span>Tracks token usage and cost per trace, per model, and per user for chargeback and per-team budgets</span></p></li><li><p><span>Integrates with OpenTelemetry, LangChain, LiteLLM, and the OpenAI SDK, so routing across providers stays visible in one place</span></p></li><li><p><span>Self-hostable, which keeps prompt and completion data inside your own boundary</span></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://github.com/langfuse/langfuse&quot;,&quot;text&quot;:&quot;Learn more about Langfuse&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://github.com/langfuse/langfuse"><span>Learn more about Langfuse</span></a></p><div><hr></div><h2><strong>&#128206; Tech Briefs</strong></h2><ul><li><p><a href="https://blog.modelcontextprotocol.io/posts/2026-07-28/"><span>Model Context Protocol ships its 2026-07-28 specification</span></a><span> - Sessions and initialization are removed for stateless requests, with authorization and versioned extensions now explicit.</span></p></li><li><p><a href="https://github.blog/changelog/2026-07-28-npm-publish-time-malware-scanning-and-dual-use-metadata/"><span>npm adds publish-time malware scanning and dual-use metadata</span></a><span> -  New packages now face pre-install scanning, metadata disclosure, and stricter publishing controls for security-sensitive capabilities.</span></p></li><li><p><a href="https://github.blog/changelog/2026-07-28-github-actions-holds-potentially-malicious-workflows-for-approval/"><span>GitHub Actions holds potentially malicious workflows for approval</span></a><span> - Suspicious runs pause before execution, moving CI supply-chain defense ahead of the workflow rather than after.</span></p></li><li><p><a href="https://www.kubernetes.dev/resources/release/"><span>Kubernetes v1.37 reaches code and test freeze</span></a><span> - Code changes now require exceptions, shifting attention to release-blocking tests and regressions before August&#8217;s release.</span></p></li><li><p><a href="https://newsroom.accenture.com/blogs/2026/accenture-tokenomics-launched-to-help-enterprises-manage-ai-token-spend"><span>Accenture launches a tokenomics practice for enterprise AI spend</span></a><span> - Token spend gets tied to business outcomes, making AI cost governance a consulting line item.</span></p></li></ul><div><hr></div><div class="callout-block" data-callout="true"><p><strong>&#128227; Contribute to Deep Engineering</strong></p><p><strong>Pitch</strong><span> a </span><a href="https://deepengineering.net/s/practical-deep-dives">practical deep dive</a><span> under your </span><strong>byline</strong><span>. Or if you lead a team, we would like to </span><strong>interview</strong><span> you and build an </span><a href="https://deepengineering.net/s/engineering-leadership">engineering leadership</a><span> feature around your </span><strong>story</strong><span>.</span><br><br><strong>Subscribe</strong><span> to </span><strong>Deep Engineering</strong><span> newsletter and </span><strong>message</strong><span> us through the </span><strong>chat option</strong><span>, or email us at </span><strong>saqibj @ packt.com</strong><span>.</span></p></div><div><hr></div><p><span>That&#8217;s all for today. Thank you for reading this issue of Deep Engineering.</span></p><p><span>We&#8217;ll be back next week with more expert-led content.</span></p><p><span>Keep building,</span></p><p><span>Saqib Jan</span></p><p><span>Editor-in-Chief, Deep Engineering</span></p><div><hr></div><p><em><span>If your company wants to reach senior developers, software engineers, and technical decision-makers, </span><a href="https://packt.omeclk.com/portal/wts/uc%5EcnN2dfNaqmD-kB-mo66%7C7g%5Ef%7Cb"><span>speak to us about partnering</span></a><span> with Deep Engineering.</span></em></p>]]></content:encoded></item><item><title><![CDATA[I burned 51 million tokens on one merge conflict, and the model was not the problem]]></title><description><![CDATA[Rory Preddy on why agent bills climb while token prices fall, and the caching, compaction, routing and loop controls that cut spend in production]]></description><link>https://deepengineering.net/p/token-efficiency-rory-preddy-agent-token-costs</link><guid isPermaLink="false">https://deepengineering.net/p/token-efficiency-rory-preddy-agent-token-costs</guid><pubDate>Thu, 30 Jul 2026 12:06:07 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/fb227e32-e539-4c24-a3da-5b6ba7b68be5_3200x1800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><span>By </span><a href="https://za.linkedin.com/in/rorypreddy"><span>Rory Preddy</span></a><span>, AI Advocate, Developer Relations at </span><strong><span>Microsoft</span></strong><span> and </span><strong><span>GitHub</span></strong><span>. Creator of </span><a href="https://github.com/microsoft/LangChain4j-for-Beginners"><span>LangChain4j for Beginners</span></a><span>. | Edited by </span><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;id&quot;:427210082,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;uuid&quot;:&quot;a2513e3d-aec1-40f3-b727-68756db92fb3&quot;}" data-component-name="MentionToDOM"></span> </p></blockquote><p><span>I have been in IT for 27 years, and 23 of those as a programmer. I started as a Java developer. I moved to Microsoft seven years and eight months ago, and I have been a developer advocate for five and a half of those, first in cloud advocacy and now in AI.</span></p><p><span>For most of that time my job was to tell people to go and build. Lately I have started saying something else first. Take a step back. Not because the excitement is wrong, but because I keep meeting developers who ran out of tokens and cannot tell me where they went. I ask what happened and I hear the same answer. I used the top model, I sent it away, and it ran out. I ask whether they needed the best model for that job and they say they do not know.</span></p><p><span>That is not a model problem. That is a foundation problem, and it is fixable in an afternoon.</span></p><p>So let&#8217;s dig deeper into what tokens actually cost you, the two mechanisms that cut the bill the most, the patterns that keep the expensive model out of cheap work, and the setup you do before any of it.</p><blockquote><p><em>This deep dive is adapted from Rory Preddy&#8217;s Deep Engineering Live session on token efficiency, edited from the session transcript for length and clarity. Slides from the talk are here, and the full recording is here.</em></p></blockquote><div class="file-embed-wrapper" data-component-name="FileToDOM"><div class="file-embed-container-reader"><div class="file-embed-container-top"><image class="file-embed-thumbnail-default" src="https://substackcdn.com/image/fetch/$s_!0Cy0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack.com%2Fimg%2Fattachment_icon.svg"></image><div class="file-embed-details"><div class="file-embed-details-h1">Token Efficiency Foundry Instantmodels</div><div class="file-embed-details-h2">2.37MB &#8729; PDF file</div></div><a class="file-embed-button wide" href="https://deepengineering.net/api/v1/file/027ad221-82df-4234-be06-3561cd784573.pdf"><span class="file-embed-button-text">Download</span></a></div><a class="file-embed-button narrow" href="https://deepengineering.net/api/v1/file/027ad221-82df-4234-be06-3561cd784573.pdf"><span class="file-embed-button-text">Download</span></a></div></div><p><a href="https://deepengineering.net/api/v1/file/e3851c6c-ccc1-4eda-be9b-0c29c3db9a38.pdf">Download</a></p><div id="youtube2-tVRcYTX_HCg" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;tVRcYTX_HCg&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/tVRcYTX_HCg?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2><span>Four levers, and one of them gets ignored</span></h2><p><span>You have four things to pull, and they move independently.</span></p><p><span>Model choice sets the unit price on every token in the exchange. Prompt structure sets how many input tokens you send and whether any of them qualify for a discount. Cache-key stability decides whether a repeated block earns the cached rate or gets billed again at full price. And output limits cap the completion, which is the one people skip.</span></p><p><span>Output is where I see the least attention and some of the highest prices. On the GPT-5.6 family the flagship Sol tier runs 5 dollars per million input tokens against 30 dollars per million output. Luna, the small one, is 1 dollar and 6 dollars. Look at the gap between those two numbers on the same row before you tune anything else. Watch your outputs, not just your inputs.</span></p><p><span>One more thing before the mechanisms. You do not have to guess at any of these rates. Azure publishes live pricing at prices.azure.com/api/retail/prices and you can query it. It looks confusing the first time you open it, because a single meter name folds the model family, the tier, the processing mode and the billing unit into one string, and there is a region attached to all of it. So let&#8217;s all breathe. Two things in there matter. The same model bills differently in different regions, and batch processing is its own meter rather than a discount on the standard rate.</span></p><p><span>It is not complicated. It is just not written down anywhere you were looking.</span></p><h2><span>Caching, and why you only pay for the tail</span></h2><p><span>Ask a model to tell you a joke. Now ask it again, exactly the same way. You are not going to be charged the same for the second one, or if you are, it will be a very small charge.</span></p><p><span>Now scale that up. Say I am an insurance company and I need to send millions of customers their latest statement. Every one of those requests carries the same instructions, the same reference material, the same everything, and then right at the bottom there is a name and a dollar amount that changes. That big identical block at the front is the stable prefix. The little bit at the end that changes is the tail.</span></p><p><span>You do not pay full price for the prefix. You pay for the tail.</span></p><p><span>The first call is a miss. The model stores the prefix and bills you the full input at standard rate. Every call after that, sent with the same prefix and the same cache key, is a hit, and that prefix gets billed as cheaper cached input. In my demo the prefix carried around 120 identical reference sections behind a run-scoped key, and warming the cache before the measured calls saved 99 percent on the repeated path.</span></p><p><span>Here is the part that catches people. The saving depends on those bytes being identical. Put a timestamp near the top of your prompt, or a request ID, or a freshly serialised list of tools, and you have moved the boundary and lost the discount on every single call. It will look completely harmless in review. So think ahead with the prompt structure. Decide what is stable and put it first, on purpose, before you need it.</span></p><h2><span>Compaction, so the context stops growing</span></h2><p><span>Caching handles repetition between calls. Compaction handles what piles up inside one conversation.</span></p><p><span>Every time you talk to an agent it remembers the conversation, and it sends that history again on the next turn. So the tokens you are billed per turn climb, and climb, and climb, even when the actual work per turn has not changed at all. Model a twelve-turn session and you are sending around 1,920 context tokens per turn by the end. Compact it along the way and it resets near 795 and never goes above that band. That projection is illustrative, but the summary reduction underneath it is measured, 62 percent, from 795 tokens down to 302.</span></p><p><span>What makes a compaction prompt work is being specific about what survives. Mine asks for at most six sentences, no bullets, no nested lists. It keeps only what the next turn needs, so the goal, the key facts, the files to update, the validation and deploy commands, the blockers and any privacy constraints. It throws out repetition, resolved dead ends, greetings, transient logs, and exact values that do not matter any more, and it turns quota findings into sanitised evidence instead of repeating every number. On a working-notes example that took a prompt from 409 tokens down to 233.</span></p><p><span>I also tried the other direction, squeezing the output instead of the context. I have a prompt that tells the model to answer like a caveman. Why use many token when few do trick. Drop the filler words and the articles, keep every technical fact, every command, every file path and every error string exact. Brain still big, mouth small.</span></p><p><span>It works, and the output count came down against a roughly 600-token baseline. But I will be honest with you about the trade. The caveman answer lost some of the understanding the fuller answer had. It is not as detailed. If your reader has to come back and ask a follow-up question, you have paid for that saving twice. A hard cap on output length is blunter but more predictable, because it truncates rather than condenses, so use it where the shape of the answer is already known.</span></p><h2><span>Set up the project before you write anything</span></h2><p><span>This is the least exciting part and it decides most of your cost.</span></p><p><span>Three files. A devcontainer.json, so the project runs on everyone&#8217;s machine and not just yours. An instructions file, so the agent has standing context and is not working blind. And an agent profile, which is the one people leave out.</span></p><p><span>The agent profile declares which tools the agent can reach for, and mine only has the tools it needs to function. That matters more than it sounds, because every tool you register ships its schema into the request on every single call. Leave a pile of MCP tools switched on that this task never touches and you are paying to send that JSON back and forth all day. Drop the ones you are not using.</span></p><p><span>Then three habits alongside it. Scope your sessions tightly. Keep your prompts precise and short. And if a direct completion will do the job, use a direct completion instead of turning an agent loose on it.</span></p><p><span>I build these three files with the cheapest model I have. Luna, reasoning switched off. It costs me almost nothing to create the baseline, and I would rather learn to do it properly at that price. On a cost-tips run I will say it out loud as I go. I am not using reasoning. I do not need reasoning right here. That is the right tool for the job.</span></p><h2><span>What an agent actually is</span></h2><p><span>An agent has agency. Agency is self determination, which means the agent understands there is a task it has to achieve. You can only give it that task if you understand it yourself.</span></p><p><span>What an agent is not is a large language model that you tell to go in and do whatever it wants. That is madness.</span></p><p><span>Here is the version I trust. I have a supervisor agent that transfers 100 dollars from one person to another and converts it to euros. Underneath it there is a withdraw agent, a credit agent and an exchange agent. Each one has a small scope. It runs on GPT-4o, so it costs me cents per month. It runs on Azure with managed identity, so it is locked down. I created the project with the dev container and the agent profile before any of it ran.</span></p><p><span>That is what I consider an agent. There is a plan, you built it, you locked it down, and it performs one well-contextualised task properly.</span></p><h2><span>Eight patterns that keep the big model out of cheap work</span></h2><p><span>Once the foundation is there, the patterns are about one idea. Do not let the largest model see work it does not need to see. My tiers are gpt-5.6-luna small, gpt-5.6-terra medium and gpt-5.6-sol large, and the numbers below are projections from my own demo harness, not benchmarks.</span></p><p><span>A </span><strong><span>router</span></strong><span> puts a cheap classifier in front and sends the request down one specialist path. 60 to 80 percent. </span><strong><span>Triage</span></strong><span> is similar but the gate is deterministic and costs zero tokens, and it only escalates when the complexity earns it. 50 to 70 percent. </span><strong><span>Context compression</span></strong><span> has Luna condense a long transcript so Sol never sees the whole thing. 30 to 60 percent. </span><strong><span>Retrieval</span></strong><span> pulls the chunks that matter instead of the whole corpus. 40 to 60 percent.</span></p><p><strong><span>Tool use</span></strong><span> moves exact computation out of probabilistic generation and into code, and it goes hand in hand with keeping that tool list short. 30 to 50 percent. </span><strong><span>Step-back planning</span></strong><span> resolves the frame before the expensive execution starts, and it only pays when it prevents retries and rework. 20 to 40 percent. </span><strong><span>Caching</span></strong><span> covers the repeated path, 50 to 90 percent on a hit. And </span><strong><span>batching</span></strong><span> I will not oversell, because a parallel mapper lowers your wall-clock time and saves you no tokens at all.</span></p><p><span>Retrieval deserves more than one line, because everyone reaches for it and plenty of people misconfigure it. The query becomes an embedding, the embedding drives a vector search, the search returns chunks, and the model answers from those chunks alone. So your chunking and your embedding model decide your answer quality and your token count at the same time. The failure I see is sending everything back every time, because that is easier than tuning retrieval, and then the customer asks why it is slow and you ask why it is expensive. My retrieval demo projected 67 percent against sending the whole document. All you need is the dollar amount, not the entire insurance quote.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nzRB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7545fdb-2580-4f83-a1cd-d58d997deeb3_3200x2400.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nzRB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7545fdb-2580-4f83-a1cd-d58d997deeb3_3200x2400.png 424w, https://substackcdn.com/image/fetch/$s_!nzRB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7545fdb-2580-4f83-a1cd-d58d997deeb3_3200x2400.png 848w, https://substackcdn.com/image/fetch/$s_!nzRB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7545fdb-2580-4f83-a1cd-d58d997deeb3_3200x2400.png 1272w, https://substackcdn.com/image/fetch/$s_!nzRB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7545fdb-2580-4f83-a1cd-d58d997deeb3_3200x2400.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nzRB!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7545fdb-2580-4f83-a1cd-d58d997deeb3_3200x2400.png" width="1200" height="900" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b7545fdb-2580-4f83-a1cd-d58d997deeb3_3200x2400.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;large&quot;,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:1200,&quot;bytes&quot;:434044,&quot;alt&quot;:&quot;the eight-pattern grid slide&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/209102293?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7545fdb-2580-4f83-a1cd-d58d997deeb3_3200x2400.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-large" alt="the eight-pattern grid slide" title="the eight-pattern grid slide" srcset="https://substackcdn.com/image/fetch/$s_!nzRB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7545fdb-2580-4f83-a1cd-d58d997deeb3_3200x2400.png 424w, https://substackcdn.com/image/fetch/$s_!nzRB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7545fdb-2580-4f83-a1cd-d58d997deeb3_3200x2400.png 848w, https://substackcdn.com/image/fetch/$s_!nzRB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7545fdb-2580-4f83-a1cd-d58d997deeb3_3200x2400.png 1272w, https://substackcdn.com/image/fetch/$s_!nzRB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7545fdb-2580-4f83-a1cd-d58d997deeb3_3200x2400.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><span>And where I got it wrong</span></h2><p><span>I burnt 51 million tokens recently, and I used the wrong model for a specific purpose.</span></p><p><span>I was resolving a merge conflict, and I gave it to Luna. Luna could not solve the problem. That on its own would have been a cheap mistake. What made it expensive is that I had not set a loop to say, after a while, stop. So it kept going.</span></p><p><span>Two lessons, and they pull against each other slightly. Match the model to the bottleneck rather than always reaching down, because a cheap model that cannot finish costs more than an expensive model that can. And put an exit condition on every iterative agent, every time.</span></p><p><span>I have a refinement loop that writes a story and then reviews its own work. It caps at five iterations and accepts a score around 80 percent. It is not going to be perfect. Sometimes you really need to say, I do not need it perfect. Caching and compaction are both good, and you still want a hard stop when it hits the score you asked for.</span></p><h2><span>Build an agent that watches the bill</span></h2><p><span>You can point agents at your own costs, which is my favourite version of this.</span></p><p><span>I have a token cost runner. Its job is to execute real flows against a deployed Foundry model, read the actual token usage, and compare it against live pricing. So the number comes out of a run rather than out of arithmetic on a rate card. I ask it to find me the lowest-cost path for a prompt with a cached prefix and a changing tail, and it tells me.</span></p><p><span>Build agents to help you lower your cost for agents. Think of it as a pyramid. What you want to achieve sits at the top, and underneath it there is a foundation holding your dev container, your instructions, your agent profiles, your compaction and your compression. Everything above the foundation inherits whatever the foundation does.</span></p><h2><span>Two commands, if you want the short version</span></h2><p><span>If patterns are more than you want right now, there are two commands.</span></p><p><span>The first reads your session history and tells you where the money went. Mine was not flattering. Across sixty days my input tokens outweighed my output roughly 100 to 1. Opus 4.8 took 45.8 million input tokens at an average of 87,000 per event. One single session accounted for 14.4 million on its own. And one session climbed from 49,000 to 174,000 input tokens over 66 turns with no manual compaction at all, so auto-compaction only fired at turn 67, right at the end. Every turn past about turn 30 was resending well over 120,000 tokens on a premium model. I was the one not compacting long sessions.</span></p><p><span>The second sets a budget per session, warns you at a threshold and stops you at the limit. Mine stops after 34.7 of 30 credits and asks whether I want to add credits, raise the limit or remove it.</span></p><p><span>I also keep a canvas built from that same session data as a daily token pulse, and I refresh it across seven and thirty day windows. None of this optimises anything by itself. It is just awareness. But a number nobody looks at twice will never change how anyone works, and most of the waste I see is not a decision anyone made. It is a thing nobody checked.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!k5Mc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47286cfb-c40f-4d3a-8654-2b4179bc8ac4_3200x1800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!k5Mc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47286cfb-c40f-4d3a-8654-2b4179bc8ac4_3200x1800.png 424w, https://substackcdn.com/image/fetch/$s_!k5Mc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47286cfb-c40f-4d3a-8654-2b4179bc8ac4_3200x1800.png 848w, https://substackcdn.com/image/fetch/$s_!k5Mc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47286cfb-c40f-4d3a-8654-2b4179bc8ac4_3200x1800.png 1272w, https://substackcdn.com/image/fetch/$s_!k5Mc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47286cfb-c40f-4d3a-8654-2b4179bc8ac4_3200x1800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!k5Mc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47286cfb-c40f-4d3a-8654-2b4179bc8ac4_3200x1800.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/47286cfb-c40f-4d3a-8654-2b4179bc8ac4_3200x1800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:263312,&quot;alt&quot;:&quot;the two terminal screenshots, the cost-tips output and the session-limit prompt&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/209102293?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47286cfb-c40f-4d3a-8654-2b4179bc8ac4_3200x1800.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="the two terminal screenshots, the cost-tips output and the session-limit prompt" title="the two terminal screenshots, the cost-tips output and the session-limit prompt" srcset="https://substackcdn.com/image/fetch/$s_!k5Mc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47286cfb-c40f-4d3a-8654-2b4179bc8ac4_3200x1800.png 424w, https://substackcdn.com/image/fetch/$s_!k5Mc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47286cfb-c40f-4d3a-8654-2b4179bc8ac4_3200x1800.png 848w, https://substackcdn.com/image/fetch/$s_!k5Mc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47286cfb-c40f-4d3a-8654-2b4179bc8ac4_3200x1800.png 1272w, https://substackcdn.com/image/fetch/$s_!k5Mc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47286cfb-c40f-4d3a-8654-2b4179bc8ac4_3200x1800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Cost tips</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QNs4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7cf72-8a03-46bb-b6b5-60755af80f47_3200x1800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QNs4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7cf72-8a03-46bb-b6b5-60755af80f47_3200x1800.png 424w, https://substackcdn.com/image/fetch/$s_!QNs4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7cf72-8a03-46bb-b6b5-60755af80f47_3200x1800.png 848w, https://substackcdn.com/image/fetch/$s_!QNs4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7cf72-8a03-46bb-b6b5-60755af80f47_3200x1800.png 1272w, https://substackcdn.com/image/fetch/$s_!QNs4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7cf72-8a03-46bb-b6b5-60755af80f47_3200x1800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QNs4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7cf72-8a03-46bb-b6b5-60755af80f47_3200x1800.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/04c7cf72-8a03-46bb-b6b5-60755af80f47_3200x1800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:230249,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/209102293?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7cf72-8a03-46bb-b6b5-60755af80f47_3200x1800.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QNs4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7cf72-8a03-46bb-b6b5-60755af80f47_3200x1800.png 424w, https://substackcdn.com/image/fetch/$s_!QNs4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7cf72-8a03-46bb-b6b5-60755af80f47_3200x1800.png 848w, https://substackcdn.com/image/fetch/$s_!QNs4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7cf72-8a03-46bb-b6b5-60755af80f47_3200x1800.png 1272w, https://substackcdn.com/image/fetch/$s_!QNs4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7cf72-8a03-46bb-b6b5-60755af80f47_3200x1800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Session limit</figcaption></figure></div><h2><span>Pick the harness you already pay for</span></h2><p><span>People ask me which harness balances capability against cost. My honest answer is that what you are trying to achieve settles that before cost does.</span></p><p><span>Then I will tell you the best harness to save costs is GitHub Copilot, and you should hear that knowing I work on the Copilot side. My reasoning stands on its own though. If you have a Copilot licence you already have the spend, and Copilot is an agent, not just an editor feature. There is a Copilot SDK and there is the Microsoft Agent Framework, and you route from there. Send the task to Copilot, or to Claude Code, or to a local model, based on the work. It is the same routing decision as picking a model, one level up.</span></p><h2><span>Take a step back</span></h2><p><span>There is a whole industry forming around this. I watched it happen with cloud, where engineers who wrote good scripts became cloud engineers, and then cost optimisation became its own discipline. I used to sit with CTOs of banks who told me their biggest worry was how much they were going to spend. Then it changed, and the question became how much they were going to save. The same thing is happening now with tokens, and the Linux Foundation&#8217;s Tokenomics Foundation is going to matter here, because standards and dashboards and protection mechanisms are what turn this from a habit into a practice.</span></p><p><span>So do not go berserk. Create the foundation. Try Luna, switch reasoning off, and see how far it gets you. If it does not match your needs, add a router and a design pattern. Start small and build up. Rather than spend more, spend less.</span></p><p><span>And one last thing, because I have another whole talk about this. Do not end each day exhausted, not from the work itself, but from managing of the work. Let the adrenaline flow, do the vibe coding, do it in the correct way. Know what you are doing. Be kind to yourself.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WEkT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157ca3c0-4404-44ba-a53f-1c886d59cb19_2400x3000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WEkT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157ca3c0-4404-44ba-a53f-1c886d59cb19_2400x3000.png 424w, https://substackcdn.com/image/fetch/$s_!WEkT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157ca3c0-4404-44ba-a53f-1c886d59cb19_2400x3000.png 848w, https://substackcdn.com/image/fetch/$s_!WEkT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157ca3c0-4404-44ba-a53f-1c886d59cb19_2400x3000.png 1272w, https://substackcdn.com/image/fetch/$s_!WEkT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157ca3c0-4404-44ba-a53f-1c886d59cb19_2400x3000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WEkT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157ca3c0-4404-44ba-a53f-1c886d59cb19_2400x3000.png" width="728" height="910" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/157ca3c0-4404-44ba-a53f-1c886d59cb19_2400x3000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1820,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:428971,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/209102293?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157ca3c0-4404-44ba-a53f-1c886d59cb19_2400x3000.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WEkT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157ca3c0-4404-44ba-a53f-1c886d59cb19_2400x3000.png 424w, https://substackcdn.com/image/fetch/$s_!WEkT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157ca3c0-4404-44ba-a53f-1c886d59cb19_2400x3000.png 848w, https://substackcdn.com/image/fetch/$s_!WEkT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157ca3c0-4404-44ba-a53f-1c886d59cb19_2400x3000.png 1272w, https://substackcdn.com/image/fetch/$s_!WEkT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157ca3c0-4404-44ba-a53f-1c886d59cb19_2400x3000.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Now go and look at your own session history.</span></p><div><hr></div><p><strong>Session notes:</strong></p><blockquote><p>Rory&#8217;s demos run on LangChain4j&#8217;s agentic modules, which matters for teams on the JVM because most token-efficiency tooling ships Python first. His pattern code is at <a href="http://github.com/roryp/token-design-patterns">github.com/roryp/token-design-patterns</a>, the agentic pattern showcase behind the banking demo at aka.ms/agentpatterns, the live cost demo at <strong>aka.ms/costs</strong>, and his LangChain4j course at <a href="http://aka.ms/LangChain4j-for-Beginners">aka.ms/LangChain4j-for-Beginners</a>.</p></blockquote><div><hr></div><p><em><span>This deep dive is adapted from </span><strong><span>Rory&#8217;s Deep Engineering Live session</span></strong><span> on The Future of Software Development. Here are the slides from the talk, and the full recording.</span></em></p><div class="file-embed-wrapper" data-component-name="FileToDOM"><div class="file-embed-container-reader"><div class="file-embed-container-top"><image class="file-embed-thumbnail-default" src="https://substackcdn.com/image/fetch/$s_!0Cy0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack.com%2Fimg%2Fattachment_icon.svg"></image><div class="file-embed-details"><div class="file-embed-details-h1">Token Efficiency Foundry Instantmodels</div><div class="file-embed-details-h2">2.37MB &#8729; PDF file</div></div><a class="file-embed-button wide" href="https://deepengineering.net/api/v1/file/e3851c6c-ccc1-4eda-be9b-0c29c3db9a38.pdf"><span class="file-embed-button-text">Download</span></a></div><a class="file-embed-button narrow" href="https://deepengineering.net/api/v1/file/e3851c6c-ccc1-4eda-be9b-0c29c3db9a38.pdf"><span class="file-embed-button-text">Download</span></a></div></div><p> </p><div id="youtube2-tVRcYTX_HCg" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;tVRcYTX_HCg&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/tVRcYTX_HCg?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div><hr></div><p><strong><span>Connect with Rory Preddy on LinkedIn:</span></strong><span> </span><a href="https://za.linkedin.com/in/rorypreddy"><span>https://za.linkedin.com/in/rorypreddy</span></a><span><br></span></p>]]></content:encoded></item><item><title><![CDATA[Site Reliability Engineer (SRE) Interviews Measure a Depreciating Skill]]></title><description><![CDATA[Notable engineering leaders on motivation, ambiguity, tribal knowledge and the measurement problem behind every reliability hire that did not work out]]></description><link>https://deepengineering.net/p/sre-hiring-interviews-measure-depreciating-skills</link><guid isPermaLink="false">https://deepengineering.net/p/sre-hiring-interviews-measure-depreciating-skills</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Wed, 29 Jul 2026 16:08:58 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/eeacb1d8-a1ef-492f-b24c-3a1a426dbe4f_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Reliability teams are hiring against a job description that describes work their agents increasingly do. That mismatch surfaces as a strong candidate who cleared every round and then underperformed, as a loop where three finalists looked interchangeable and the decision came down to feel, or as a new hire who put their first year into building instrumentation nobody had told them was missing.</p><p>For most of the discipline&#8217;s history the binding constraint during an incident was human attention, because somebody had to correlate logs across services, hold a working model of the system, and form a hypothesis fast enough to matter. Interview loops were built to find that person and they were right to. Agents now take a growing share of that first pass, and the loop has not moved with them. A coding exercise, a debugging scenario, and a systems round all measure how fast a candidate converges on a cause. Those are the skills losing value fastest, while the ones that decide whether the hire works go untested.</p><p><a href="https://www.linkedin.com/in/reid-s/"><span>Reid Savage</span></a><span>, Senior Engineering Manager for the Site Reliability Engineering team at </span><a href="https://www.honeycomb.io/"><span>Honeycomb</span></a><span>, shares how he has repeatedly seen that gap widen across the hires he has made and overseen.</span></p><p><span>&#8220;Standard SRE hiring loops are far too skills-based,&#8221; he says. And his own team acted on that conclusion, removing coding exercises from the loop at Honeycomb&#8217;s current stage.</span></p><h2><span>Motivation decides the hire, not the skills matrix</span></h2><p><span>The failure pattern Savage describes does not look like a skills gap, which is what makes it so easy for a conventional loop to miss. &#8220;I&#8217;ve had excellent hires that have never used our cloud of choice, and some who didn&#8217;t work out that had nearly everything,&#8221; he says. The candidates who arrived with the full stack on their r&#233;sum&#233; were not reliably the ones who lasted, and the ones missing the specific platform experience frequently were. That inversion is not a story about credentials being worthless, and it is a story about credentials answering a question that turns out not to be decisive.</span></p><p><span>What Savage found underneath the pattern was a mismatch of a different kind. The problem in the hires that did not work out was &#8220;usually a mismatch between true motivation (which requires significant self-insight on their part) and what the team needs,&#8221; he explains. Motivation in this framing is not enthusiasm for the role or eagerness in the interview, both of which candidates supply readily and neither of which predicts much. It is the specific thing that makes the work satisfying to that person, which requires enough self-knowledge on their part to name it accurately. Savage anchors the distinction to a line he paraphrases from </span><em><span>First, Break All The Rules</span></em><span>, about hiring the accountant who cannot sleep until the books are balanced rather than the one who knows Excel.</span></p><p><span>The scale of Honeycomb&#8217;s problem space is part of why this holds. &#8220;At Honeycomb&#8217;s size, the skills required for each problem we have wouldn&#8217;t fit in 5 job descriptions,&#8221; Savage says, which means any specific skill the loop screens for covers a small fraction of what the hire will actually encounter. A loop optimized for skills coverage is therefore optimizing a variable that cannot be made to matter, because the surface area defeats it. The organizations most exposed to this are the ones whose reliability function touches many systems rather than one, which describes most companies past their first platform consolidation.</span></p><p><span>Honeycomb&#8217;s replacement for the coding exercise is the actionable part. The technical screen is now a pull request review, chosen because, in Savage&#8217;s words, &#8220;it covers the important things: reading between the lines, social skills, eagerness to help, mindset, and integrity.&#8221; A PR review works as a screen because it presents the candidate with someone else&#8217;s decisions rather than a blank editor, and reliability work consists largely of forming a fast, correct read on choices other people made under constraints that are no longer visible. Leaders who want to run this should use a real pull request from their own history, ideally one where the right call was contested, and score the review on what the candidate noticed and chose to raise rather than on whether they found a seeded bug.</span></p><p><span>Savage&#8217;s position is that the technical bar clears more reliably in the presence of those qualities than the reverse, because &#8220;if you have those, and they display the level of technical aptitude you need, they&#8217;ll be able to catch up.&#8221;</span></p><h2><span>Profile of the hire changed inside a year</span></h2><p><span>Fixing the screen addresses the process. The harder change is to the role itself, and it has moved faster than most loops have been revised. &#8220;The profile of a great SRE has fundamentally changed over the last year,&#8221; says </span><a href="https://www.linkedin.com/in/anish-agarwal-io"><span>Anish Agarwal</span></a><span>, co-founder and CEO of </span><a href="https://www.traversal.com/"><span>Traversal AI</span></a><span>, which builds AI systems for production incident investigation. His account of what companies previously optimized for matches what most loops still measure, which is the engineer who could reason through dashboards, logs, and metrics under pressure and arrive at a root cause. That was the correct thing to optimize while humans carried every investigation from alert to resolution.</span></p><p><span>Agarwal&#8217;s argument is that agentic systems now take on the most time-consuming parts of that work, and the consequence for hiring follows directly. &#8220;Raw troubleshooting ability is becoming less of the differentiator than it was even a year ago,&#8221; he says. A capability stops being a differentiator once it stops being scarce, and a loop weighted toward a non-scarce capability produces candidates who cluster indistinguishably at the top. Leaders tend to notice this as a symptom well before they diagnose it, in loops where several finalists all pass the debugging round cleanly and the final decision comes down to feel.</span></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Packt Deep Engineering! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>What replaces it, in Agarwal&#8217;s view, is the part of the role that always belonged there and rarely got attention. &#8220;What has become far more important is the ability to engineer resilient systems,&#8221; he says, being precise about why that capability went underexercised for so long. Most teams never found time for it because troubleshooting consumed their days, so resilience work stayed aspirational rather than scheduled. Two changes now arrive together, with AI removing much of the investigative burden while production systems grow more complex through new services, new dependencies, and new failure modes each quarter. His conclusion is that resilience engineering &#8220;is no longer the work SREs wish they had time for. It&#8217;s increasingly the work the role demands.&#8221;</span></p><p><span>That reframes what a strong candidate has to demonstrate in an interview. Agarwal describes the highest-leverage SREs as the ones &#8220;thinking about what the system will need a year from now, which SLAs actually matter, and which technologies best meet those SLAs given the tradeoffs between cost, reliability, and performance.&#8221; A loop that wants to surface this has to ask forward-looking questions rather than diagnostic ones, and the practical version is to hand the candidate a real architecture with real growth assumptions and ask what breaks first at ten times the load. The answer reveals whether the candidate reasons about failure modes that do not exist yet, which no incident replay exercise can establish. Agarwal adds one further capability he would weight more heavily, describing it as &#8220;AI fluency. Not just expertise, but genuine curiosity,&#8221; and locating its value in engineers who use agents to compress investigations so that the time freed goes into resilience work.</span></p><h2><span>Judgment is the scarce input, and it leaves when people do</span></h2><p>Hiring toward resilience assumes the judgment to do that work can be found and kept, and that assumption meets an organizational problem that no interview round surfaces.<span> </span><a href="https://www.ghantzaras.com/"><span>George Hantzaras</span></a><span>, Director of Engineering, Core Platforms at </span><a href="https://www.mongodb.com/"><span>MongoDB</span></a><span>, speaking at the AI-Powered Platform Engineering workshop we hosted some time back, argued that the problems facing platform and operations teams have stopped being problems of tooling or frameworks and have become problems of people, team dynamics, and how humans interact with the systems they built. His first category is the knowledge gap, where the answers engineers need lie buried across dozens of systems, dated documentation, and sprawling wikis. He described his own developers as having become archaeologists, running scavenger hunts through internal documentation to recover a single answer.</span></p><p><span>The part of that gap most relevant to hiring is the part nobody wrote down. Hantzaras underscored tribal knowledge as the real problem, meaning the unwritten rules and undocumented workarounds that exist only in the minds of the few people who have been at a company long enough to accumulate them. He described one team where new engineers took close to six months to reach full productivity, with the primary cause being the difficulty of navigating internal systems and discovering rules that appeared on no wiki. That knowledge carries a second liability beyond slow onboarding, because it leaves the organization when the people holding it leave, and nothing in a standard hiring process accounts for either effect.</span></p><p><span>Two further problems he named sharpen what the reliability hire is actually for. Golden paths handle roughly the first eighty percent of cases well and then become what he calls golden cages for the remaining twenty percent, which tend to be the most innovative work and the cases where a developer needs something the template will not express. He also described a platform lead whose team was putting more time into maintaining YAML templates for their portal than into building new platform capability, an overhead he calls the tooling tax. Both problems consume senior judgment on work that produces nothing durable, which is exactly the capacity a resilience-focused hire is supposed to create.</span></p><p><span>Hantzaras reaches the case for automating the first pass from the opposite direction, arriving through the operations burden rather than through the hiring loop. He pointed out that the industry has fixated on mean time to recovery while, in his reading, the largest component is mean time to identify, and that teams are drowning in logs, metrics, traces, and alerts without a good way to act on any of it. The resulting alert fatigue leads engineers to tune out noise until a critical alert has a high probability of being lost. His read on responsibility for that is unambiguous, because it is almost never a case of an engineer not working and almost always a case of a system not working as it should, and his conclusion is that filtering data at that volume &#8220;shouldn&#8217;t be a human job. This is a machine job.&#8221;</span></p><p><span>For leaders hiring into this, the implication is that the scarce input is accumulated judgment rather than throughput, and that judgment currently has no home outside individual heads. The practical response runs on two tracks. Weight the loop toward candidates who can reconstruct the reasoning behind decisions they did not make, since that is the skill tribal knowledge recovery actually requires, and a PR review from an unfamiliar codebase tests it directly. Then treat the documentation of undocumented judgment as an explicit deliverable in the first six months rather than as good citizenship, because the alternative is rehiring the same knowledge every time someone resigns.</span></p><h2><span>Foundations come before fluency</span></h2><p><span>Hiring for AI fluency assumes the environment can support it, and that assumption fails more often than the hiring conversation admits. </span><a href="https://www.linkedin.com/in/chankramath"><span>Ajay Chankramath</span></a><span>, CTO of </span><a href="https://platformengineering.org/"><span>Platform Engineering Community</span></a><span>, a member organization for DevOps and cloud native practitioners, led the platform engineering workshop sessions for Deep Engineering recently, where he described platform maturity as a progression through four states, moving from reactive to responsive, then to predictive, and finally to autonomous. His caution for teams reaching for the later states is that observability, runbooks, and service level objectives have to be in place before AI gets layered on top, because introducing AI into an environment without those foundations amplifies the existing chaos rather than resolving it. That ordering has a direct hiring consequence, because a candidate hired for AI fluency into a team without instrumented systems will put their first year into building the substrate instead.</span></p><p><span>Chankramath is equally clear that the tooling does not displace the role. His framing of AI in incident response is that it hands the engineer a capability rather than taking one away, reading the runbook on their behalf and proposing the next step, while the engineer remains the expert who decides whether that step gets taken. He grounds this in an observation most people on call will recognize, which is that nobody reads documentation during a live incident, because when the system is failing and customers are escalating, engineers act on what they already know rather than opening a markdown file. AI closes that gap by surfacing the relevant procedure at the moment of the decision, and the judgment about whether to follow it stays where it was.</span></p><p><span>He also names a failure mode that should change how leaders think about seniority in these hires. Platform capabilities frequently go unused by the most experienced engineers, who treat self-service tooling as something built for juniors who do not know how to do the work directly, and because those engineers are the ones whose opinions carry weight internally, their dismissal propagates. A reliability hire brought in specifically to raise AI adoption will therefore run into resistance from exactly the people whose endorsement determines whether the change holds. Leaders should test for this in the interview by asking how the candidate has previously won over a skeptical senior engineer, rather than assuming enthusiasm for the tooling transfers to the team.</span></p><p><span>The last piece of Chankramath&#8217;s argument concerns measurement, and it constrains how any of this gets evaluated. He treats the standard delivery metrics as lagging indicators, useful for confirming what already happened and poor at telling a team where things are heading, which makes running an organization solely on them a mistake. Leaders hiring for resilience need leading indicators to judge whether the hire is working, because a resilience improvement shows up as incidents that never occurred, and no lagging metric records an absence. </span>The workable substitutes are the observable precursors, meaning coverage of critical paths by service level objectives, the share of alerts that map to a documented response, and error budget burn rate read before the budget is exhausted rather than after.</p><h2><span>Ambiguity behavior predicts the outcome better than anything else</span></h2><p><span>Across everything Savage has observed, one behavior separates the hires that worked from the hires that did not, and it is not a technical one. &#8220;The largest determinant I see is what they do when faced with an ambiguous problem,&#8221; he says, and he lays out the range of possible responses as moving away from it, tagging in someone else, asking for help, owning the solution, helping someone else own it, or investigating why the problem was ambiguous in the first place. The question underneath all of those, in his phrasing, is whether the candidate accepts or rejects unclear problems. Reliability work arrives almost exclusively as unclear problems, which is what makes this behavior more predictive than any skill the loop can measure.</span></p><p><span>The cost of getting it wrong is not usually a visible failure, which is part of why it goes undetected. Savage describes one SRE who was passionate and knowledgeable, and who struggled to bring others into the work or carry projects across the finish line, producing a long trail of initiatives that started without much to show for them. Because SREs touch every part of the business, that pattern compounds in both directions, generating leverage when it works and opportunity cost plus real cloud spend when it does not. </span>His summary of the missing skill is that &#8220;SREs need to know when to open, peek inside, rearrange, and put back problems.&#8221;<span> The judgment is as much about which problems to close as which to open.</span></p><p><span>Testing for this requires giving candidates something genuinely underspecified rather than a puzzle with a hidden solution. The practical version is to present a real ambiguous situation from the team&#8217;s history, withhold the resolution, and pay attention to whether the candidate starts solving or starts questioning the framing. Both responses can be correct, and what matters is that the candidate demonstrates a deliberate choice between them rather than defaulting.</span></p><h2><span>So define the role before you post it</span></h2><p><span>None of the above helps if the role itself was never specified, which Savage treats as the prior question. &#8220;You need to know what you&#8217;re hiring the SRE for,&#8221; he says, and the diagnostic he runs through is a series of concrete organizational facts rather than an abstract job description. </span>Those facts are whether reliability genuinely matters to the business and to what degree, how many nines the company actually needs, whether anyone currently has responsibility for predicting what will fail a year out, whether developers are resisting on-call, and whether the team is drowning in toil.<span> Each of those describes a different job, and an SRE can address any of them, though not all of them at once.</span></p><p><span>The nines question is the one most organizations answer loosely, and it is the one that determines everything downstream. A target expressed as a service level objective with an error budget attached tells a candidate what the job actually involves, because it establishes how much unreliability the business has already agreed to tolerate and therefore where the engineering effort goes. A posting that claims high availability without naming a target is describing an aspiration rather than a role.</span></p><p><span>His standard for handling that honestly is where the recommendation lands. It is fine to need some of those things and not others, as long as the job posting states what the role is expected to be today and how that might change. Most mis-hires that look like candidate failures are role definition failures surfacing late, because a candidate optimized for toil reduction will underperform in a role that turns out to require reliability forecasting, and neither party did anything wrong in the interview. The discipline is to write the posting from the specific problem rather than from a template, and to name the parts of the role that are unsettled instead of leaving them out.</span></p><p><span>The same clarity should extend into the first ninety days, and Savage&#8217;s approach here inverts the usual instrument. The senior hires who succeeded did not treat his 30/60/90 as gospel, and read it instead as an aspirational list with goals attached, oriented toward making social connections, becoming familiar with the code, and starting the long climb through the tech stack. He asks new hires early on to disagree with him and to offer at least one piece of constructive feedback, which he uses to open room for them to work on problems he cannot see himself. That is a deliberate test of the same quality the loop was built to find, because an engineer who will surface an uncomfortable observation in week three is demonstrating the disposition toward unclear problems that predicts everything else.</span></p><p><span>What connects all four positions is that none of them argues for a lower technical bar. The bar holds, and the argument is that the loop currently puts most of its measurement on the part of the bar that agents are absorbing, while the parts that decide whether the hire succeeds go untested. The reliability hire that works out is the person who accepts unclear problems, reasons about failures that have not happened, carries judgment that was never documented, and improves a system nobody else wanted to own. A process built to find that person looks different from the one most companies are still running.</span></p>]]></content:encoded></item><item><title><![CDATA[Deep Engineering #56: Peter Zaitsev on Choosing a Database Like You Can Never Leave It]]></title><description><![CDATA[Percona co-founder Peter Zaitsev on why database choices are nearly irreversible, what open source licensing actually protects, and how to match the database to the real workload]]></description><link>https://deepengineering.net/p/issue-56-choosing-a-database-like-you-can-never-leave-it</link><guid isPermaLink="false">https://deepengineering.net/p/issue-56-choosing-a-database-like-you-can-never-leave-it</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 23 Jul 2026 18:19:46 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8c463b80-a04b-4a3b-95f7-b1d1384ef8ec_2760x1040.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><a href="https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt">Capacitor - Shared memory for your team&#8217;s coding agents. </a></strong><a href="https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt">Searchable. Shareable. Vendor-neutral. Scored.</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ocpb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 424w, https://substackcdn.com/image/fetch/$s_!Ocpb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 848w, https://substackcdn.com/image/fetch/$s_!Ocpb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 1272w, https://substackcdn.com/image/fetch/$s_!Ocpb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ocpb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png" width="1360" height="660" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:660,&quot;width&quot;:1360,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!Ocpb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 424w, https://substackcdn.com/image/fetch/$s_!Ocpb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 848w, https://substackcdn.com/image/fetch/$s_!Ocpb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 1272w, https://substackcdn.com/image/fetch/$s_!Ocpb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt">Capacitor</a></strong><span> records the session behind the work: what agents tried, what teammates rejected, what finally passed, and why it mattered. This gives you vendor-neutrality to move across multiple coding agents, multiplayer - collaboration on coding sessions, faster PR reviews, evals on code &amp; more.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt&quot;,&quot;text&quot;:&quot;Sign up for free&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt"><span>Sign up for free</span></a></p><div><hr></div><p><span>&#9997;&#65039; </span><strong><span>From the editor&#8217;s desk,</span></strong></p><p><span>Welcome to the </span><strong><span>56th</span></strong><span> issue of </span><strong><span>Deep Engineering</span></strong><span>!</span></p><p><span>The PostgreSQL Global Development Group shipped</span><a href="https://www.postgresql.org/about/news/postgresql-19-beta-2-released-3350/"><span> the second beta of PostgreSQL 19</span></a><span> on July 16, only a few months after version 18 reached general availability and began landing in production systems. A community project that already tops every major developer survey keeps compounding its lead one disciplined release at a time. Postgres has effectively become the default database, and most teams now reach for it without pausing to ask whether it fits the problem in front of them.</span></p><p><span>That ease is exactly where the trouble starts, because a decision everyone treats as obvious is a decision nobody examines. Moving off a database once an application leans on it can take years and burn through real budget, and the forces that make the choice consequential have only sharpened this year. Relicensing moves keep redefining what the words open source actually protect, and by</span><a href="https://cloud.google.com/blog/topics/threat-intelligence/ai-vulnerability-exploitation-initial-access"><span> Google&#8217;s own threat-intelligence reporting</span></a><span> attackers now weaponize many newly disclosed flaws before a patch is even available, which turns the question of where and how you run your database into a live operational risk rather than a matter of preference.</span></p><p><span>That is exactly the terrain </span><a href="https://www.linkedin.com/in/peterzaitsev"><span>Peter Zaitsev</span></a><span> has worked his whole career, an entrepreneur, author, and co-founder of </span><a href="https://www.percona.com/"><span>Percona</span></a><span>. Two decades of helping companies choose, run, and sometimes escape their databases give him a sharp read on which reasons hold up under real load and which ones fall apart the moment the system starts to matter.</span></p><p><strong><span>In today&#8217;s issue</span></strong><span>, he makes the case for treating that choice with the seriousness it earns, and for choosing as though you can never walk it back. </span></p><blockquote><p><em><span>You can also </span><a href="https://deepengineering.net/p/database-choice-lock-in-ai-rush-peter-zaitsev"><span>watch our full conversation</span></a><span> with </span><strong><span>Peter Zaitsev </span></strong><span>here</span><strong><span>.</span></strong></em></p></blockquote><p><strong>Let&#8217;s get started.</strong> </p><div class="callout-block" data-callout="true"><h2><a href="https://www.eventbrite.co.uk/e/building-intelligent-ai-agents-with-graphrag-tickets-1992756563525?aff=deepeng"><span data-color="#f85f28" style="color: rgb(248, 95, 40);">Building Intelligent AI Agents with GraphRAG</span></a></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.eventbrite.co.uk/e/building-intelligent-ai-agents-with-graphrag-tickets-1992756563525?aff=deepeng" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8HGZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807412c1-fb6b-4911-8c7f-2cd3a7f2f03d_800x267.jpeg 424w, https://substackcdn.com/image/fetch/$s_!8HGZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807412c1-fb6b-4911-8c7f-2cd3a7f2f03d_800x267.jpeg 848w, https://substackcdn.com/image/fetch/$s_!8HGZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807412c1-fb6b-4911-8c7f-2cd3a7f2f03d_800x267.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!8HGZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807412c1-fb6b-4911-8c7f-2cd3a7f2f03d_800x267.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8HGZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807412c1-fb6b-4911-8c7f-2cd3a7f2f03d_800x267.jpeg" width="800" height="267" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/807412c1-fb6b-4911-8c7f-2cd3a7f2f03d_800x267.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:267,&quot;width&quot;:800,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.eventbrite.co.uk/e/building-intelligent-ai-agents-with-graphrag-tickets-1992756563525?aff=deepeng&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!8HGZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807412c1-fb6b-4911-8c7f-2cd3a7f2f03d_800x267.jpeg 424w, https://substackcdn.com/image/fetch/$s_!8HGZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807412c1-fb6b-4911-8c7f-2cd3a7f2f03d_800x267.jpeg 848w, https://substackcdn.com/image/fetch/$s_!8HGZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807412c1-fb6b-4911-8c7f-2cd3a7f2f03d_800x267.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!8HGZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807412c1-fb6b-4911-8c7f-2cd3a7f2f03d_800x267.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This <strong>Agentic Engineering live bootcamp</strong> walks you through building an agent that investigates across a knowledge graph, remembers what it found, and returns answers you can put in front of a user.</p><p><strong>Friday 31 July to Saturday 1 August, online.</strong></p><h3 style="text-align: center;"><a href="https://www.eventbrite.co.uk/e/building-intelligent-ai-agents-with-graphrag-tickets-1992756563525?aff=deepeng"><span data-color="#f85f28" style="color: rgb(248, 95, 40);">Register with code </span></a><strong><a href="https://www.eventbrite.co.uk/e/building-intelligent-ai-agents-with-graphrag-tickets-1992756563525?aff=deepeng"><span data-color="#f85f28" style="color: rgb(248, 95, 40);">DEEPENG40</span></a></strong></h3><p style="text-align: center;"></p></div><div><hr></div><p><strong>Expert Insight</strong></p><h2><span>Databases are sticky. Most teams underestimate the commitment</span></h2><p><em><span> by </span><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;id&quot;:427210082,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;uuid&quot;:&quot;e7195419-ea93-4f2b-84d4-a7959d36a209&quot;}" data-component-name="MentionToDOM"></span> <span>with </span><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Peter Zaitsev&quot;,&quot;id&quot;:12711336,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ace4d441-409d-4f22-a72a-6352b7507e61_144x144.png&quot;,&quot;uuid&quot;:&quot;2b09b703-632a-45bd-a390-8d3ffa261be0&quot;}" data-component-name="MentionToDOM"></span> </em></p><p><span>Most teams choose a database the way they choose a lunch spot, quickly and on somebody else&#8217;s recommendation. A friend runs it in production, a conference talk made it sound fast, or the cloud console offered it as the first option on the list. The 2025</span><a href="https://survey.stackoverflow.co/2025/"><span> Stack Overflow Developer Survey</span></a><span> put PostgreSQL at the top for the third year running, in use by 55.6 percent of all developers and 58.2 percent of professional ones, far ahead of second-place MySQL. The default has never been clearer. But the decision behind it still rarely gets the weight it deserves.</span></p><p><a href="https://www.linkedin.com/in/peterzaitsev"><span>Peter Zaitsev</span></a><span> has advised companies through twenty years of these choices, from the greenfield pick to the migration nobody wanted. He co-founded </span><a href="https://www.percona.com/"><span>Percona</span></a><span>, grew it from a two-person company into one of the most respected open source database companies in the industry. His view of what most teams get wrong is straightforward. &#8220;Databases are very, very sticky,&#8221; he says. Once an application leans on one, unwinding that dependency can take years and burn through budgets, which is exactly why choosing one casually is the expensive mistake.</span></p><p>The cost is not hypothetical. Sharing from his experience advising enterprises through these migrations, Zaitsev says, &#8220;They started this kind of very painful migration away from Oracle a decade ago, and many, many millions of dollars, many, many years, they&#8217;re still on it.&#8221; A database that took a weekend to adopt can take a decade to leave. That asymmetry is the reason the rest of his advice matters.</p><h3><span>Nobody owns Postgres, and that is most of why it won</span></h3><p><span>In our conversation, I asked Zaitsev why Postgres reached the top, and his answer began with the part engineers tend to skip. &#8220;PostgreSQL is truly a community database. Nobody owns Postgres,&#8221; he says, and as he sees it that governance fact, more than any feature, explains the dominance. Because no single vendor controls it, every cloud can offer it as a first-class product without handing a rival an advantage. &#8220;Amazon, Google, Microsoft, a lot of others, they can offer PostgreSQL as their own,&#8221; he points out, where promoting MySQL or SQL Server means promoting something a competitor owns.</span></p><p><span>The technical case follows the same shape. Postgres extends rather than forces migrations, so each new demand becomes an extension instead of a second database. Time-series work got its extension, and the AI wave got</span><a href="https://github.com/pgvector/pgvector"><span> pgvector</span></a><span>, the open source extension that adds vector types and similarity indexes and now ships enabled on managed Postgres from AWS RDS to Supabase. &#8220;pgvector, vector search, enabling building AI application in Postgres, that came about very quickly,&#8221; Zaitsev tells us. The base keeps getting stronger underneath all of it, and the September 2025 release of</span><a href="https://www.postgresql.org/about/news/postgresql-18-released-3142/"><span> PostgreSQL 18</span></a><span> added a new asynchronous I/O subsystem the project measured at up to three times faster on reads.</span></p><h3><span>Open source is not enough</span></h3><p>The reason to prefer the community model becomes obvious the moment a vendor changes the rules. &#8220;People even kind of lost trust in open source and started to understand that open source is not enough,&#8221; Zaitsev reflects, and he points straight at the pattern that taught them. MongoDB moved to the Server Side Public License in 2018, Elastic followed in 2021, HashiCorp adopted the Business Source License in 2023, and Redis switched to source-available terms in 2024 before adding an open license back in 2025 under community pressure. Each move left production users holding a version they could no longer upgrade under the terms they signed up for.</p><p><span>Zaitsev&#8217;s warning is that the marketing hides the trap. &#8220;A lot of companies right now are trying to trick you, using words like open source compatible,&#8221; he cautions, which usually means proprietary software wearing a thin layer of compatibility. </span><strong><span>Open core is the other tell.</span></strong><span> &#8220;There is some sort of limited, crippled version which is open source, but what they actually want is for you to buy the proprietary version,&#8221; he says. The protection he trusts is not a license badge but a market. A real open source project has many vendors, so a team unhappy with one can leave, and run a closed product like Amazon Aurora and only Amazon can support it, which puts the depth of experience and the leverage on one side of the table.</span></p><h3><span>Self-hosting now means racing the exploit clock</span></h3><p><span>Self-hosting an open source database moves the security burden onto the team, and the clock on that burden has changed. Zaitsev warns that with AI, security holes can now be found and exploited &#8220;at the speed which was never seen before,&#8221; and the incident data agrees with that. Google&#8217;s M-Trends 2026 report estimated the mean time to exploit a newly disclosed vulnerability at roughly negative seven days, meaning many flaws are weaponized before a patch exists, and</span><a href="https://cloud.google.com/blog/topics/threat-intelligence/ai-vulnerability-exploitation-initial-access"><span> Google&#8217;s threat intelligence group has reported the first zero-day it attributes to an AI-developed exploit</span></a><span>.</span></p><p><span>That pace changes what good looks like. Zaitsev argues for a layered posture and an honest plan for failure, because prevention alone never reaches certainty. &#8220;Nothing gives you a 100 percent guarantee,&#8221; he reminds us, so the real question becomes how fast a team detects and responds, not only how well it defends. For many teams the practical answer is a contract, and even those that avoid a fully managed service should have someone on the hook for patching. In regulated environments the whole subject stops being a technical choice and becomes a compliance one, which is where managed databases earn their place by letting a team push responsibility to the vendor who runs the system.</span></p><h3><span>Doing what the cool kids do is not a workload</span></h3><p><span>The pull toward document, vector, and graph stores is often social before it is technical. &#8220;A lot of engineers, they like to explore and play with new technologies,&#8221; Zaitsev says, usually after hearing about them at a conference or on a podcast. He does not moralize about curiosity, but he counts the cost. &#8220;The more technologies you introduce, the more complexity you create,&#8221; he says, and a small team running many engines pays that bill every day. &#8220;If you are a five-person company and you are self-managing 10 databases, there&#8217;s probably way too much complexity.&#8221;</span></p><p><span>His rule is to start narrow and let real limits, not fashion, force the next move. Postgres holds relational data, documents, and vectors well enough that most teams travel a long way before they need anything else, and only a genuine ceiling on scale or a missing feature should justify a dedicated store. Purpose-built systems do win when the workload is real, because a database built for one task carries a more optimized store and a better language for it. The catch with AI is speed. &#8220;The AI world is completely different. Every six months things drastically change,&#8221; Zaitsev reasons, so a team building AI features has to check the date on its advice, since guidance a year old still serves Postgres or MySQL fine and can already be stale for anything touching models.</span></p><h3><span>Choose like you cannot leave</span></h3><p><span>The through-line of Zaitsev&#8217;s advice is to match the seriousness of the decision to the difficulty of reversing it. He wants teams to write down their actual requirements rather than inherit a default, and he is happy to use the newest tool to pressure-test the old discipline. &#8220;You can actually ask AI, hey, what am I missing,&#8221; he says, treating a model as a fast second reader on a requirements list. The point is not the tool. It is that a choice this durable deserves more than an afternoon.</span></p><p><span>That is the gap he keeps returning to. Teams decide in a hurry and then live with the result for years, on infrastructure that only grows harder to change as the application matures on top of it. Postgres has made the safe default easy, and the survey numbers show most teams taking it. The harder work starts after the default, in the licensing terms, the security posture, and an honest reading of what the workload actually needs.</span></p><div><hr></div><h2><strong>In case you missed</strong></h2><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;d010d390-2a04-4604-a986-457aa9993c1f&quot;,&quot;caption&quot;:&quot;Peter Zaitsev built Percona from a two-person company into one of the most respected open source database companies in the business, and he co-wrote High Performance MySQL. He now advises open source startups as a board member, and he has watched engineering teams make the same database decisions for two decades. We asked him how teams should choose a data&#8230;&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Database Choice, Lock-In, and the AI Rush with Peter Zaitsev&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-07-23T17:06:10.283Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/faa482b4-7b9e-476b-acee-2f98f0a10b1e_1920x1080.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.net/p/database-choice-lock-in-ai-rush-peter-zaitsev&quot;,&quot;section_name&quot;:&quot;Interviews&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:208177027,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p></p><p><strong><span>Industry Perspective</span></strong></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;ba1f7be1-02d6-436a-ab01-662cc6d870a0&quot;,&quot;caption&quot;:&quot;Chuck McCullough on running a reliable engineering workflow with OpenAI Codex, scoping work into slices, testing first, and reviewing for whether it fits.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Codex as a Teammate, Not Autocomplete&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-07-23T17:47:28.555Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/42966ef5-2fee-4228-9606-6a26d942e262_2400x1600.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.net/p/chuck-mccullough-onagentic-development-with-codex&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:208229355,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p></p><div><hr></div><h2><strong>&#128736;&#65039; Tool of the Week</strong></h2><p><strong><a href="https://github.com/cloudnative-pg/cloudnative-pg"><span>CloudNativePG</span></a><span>, the Kubernetes operator for running production Postgres on infrastructure you control</span></strong></p><p><span>CloudNativePG runs Postgres as a native Kubernetes workload and handles failover, backups, and rolling upgrades the way a managed service would, while the data and the control plane stay inside your own cluster.</span></p><ul><li><p><span>New DatabaseRole resources manage Postgres roles as declarative GitOps objects with password-free certificate authentication.</span></p></li><li><p><span>A Kubernetes Lease now gates primary promotion, so replicas fail over without waiting out the full timeout.</span></p></li><li><p><span>In-place major upgrades run pg_upgrade against mounted extension images, cutting version-migration downtime.</span></p></li><li><p><span>Operator and Postgres images ship signed with SBOM and provenance attestations for supply-chain verification.</span></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://github.com/cloudnative-pg/cloudnative-pg&quot;,&quot;text&quot;:&quot;Learn more about CloudNativePG&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://github.com/cloudnative-pg/cloudnative-pg"><span>Learn more about CloudNativePG</span></a></p><div><hr></div><h2><strong>&#128206; Tech Briefs</strong></h2><ul><li><p><a href="https://www.postgresql.org/about/news/postgresql-19-beta-2-released-3350/">PostgreSQL 19 Beta 2 adds native SQL/PGQ property-graph queries</a> - Postgres 19 exposes relational tables as property graphs via GRAPH_TABLE pattern matching, so fixed-depth traversals no longer need a graph database.</p></li><li><p><a href="https://www.cisa.gov/news-events/alerts/2026/07/14/cisa-urges-sharepoint-hardening-after-new-exploitations">CISA adds an actively exploited SharePoint Server RCE to its Known Exploited Vulnerabilities catalog</a> - CISA confirmed active exploitation of CVE-2026-58644, a deserialization remote-code-execution flaw in on-premises SharePoint Server, and ordered urgent patching.</p></li><li><p><a href="https://mariadb.org/mariadb-server-10-6-reaches-end-of-life-on-july-6th/">MariaDB Community Server 10.6 reaches end of life</a> - MariaDB&#8217;s last old-model LTS no longer receives bug or security fixes, so 10.6 shops now face an unavoidable migration.</p></li><li><p><a href="https://www.postgresql.org/about/news/autobase-290-released-3343/">Autobase 2.9.0 released</a>. The open source automated Postgres high-availability platform now manages cluster infrastructure after deployment, not only initial setup.</p></li><li><p><a href="https://www.cisa.gov/news-events/alerts/2026/07/07/cisa-adds-three-known-exploited-vulnerabilities-catalog">CISA flags an actively exploited authorization bypass in Langflow</a> - CISA confirmed in-the-wild exploitation of CVE-2026-55255, an authorization bypass in the open source LLM app builder Langflow.</p></li></ul><div><hr></div><div class="callout-block" data-callout="true"><p><strong>&#128227; Contribute to Deep Engineering</strong></p><p><strong>Pitch</strong><span> a </span><a href="https://deepengineering.net/s/practical-deep-dives">practical deep dive</a><span> under your </span><strong>byline</strong><span>. Or if you lead a team, we would like to </span><strong>interview</strong><span> you and build an </span><a href="https://deepengineering.net/s/engineering-leadership">engineering leadership</a><span> feature around your </span><strong>story</strong><span>.</span><br><br><strong>Subscribe</strong><span> to </span><strong>Deep Engineering</strong><span> newsletter and </span><strong>message</strong><span> us through the </span><strong>chat option</strong><span>, or email us at </span><strong>saqibj @ packt.com</strong><span>.</span></p></div><div><hr></div><p><span>That&#8217;s all for today. Thank you for reading this issue of Deep Engineering.</span></p><p><span>We&#8217;ll be back next week with more expert-led content.</span></p><p><span>Keep building,</span></p><p><span>Saqib Jan</span></p><p><span>Editor-in-Chief, Deep Engineering</span></p><div><hr></div><p><em><span>If your company wants to reach senior developers, software engineers, and technical decision-makers, </span><a href="https://packt.omeclk.com/portal/wts/uc%5EcnN2dfNaqmD-kB-mo66%7C7g%5Ef%7Cb"><span>speak to us about partnering</span></a><span> with Deep Engineering.</span></em></p>]]></content:encoded></item><item><title><![CDATA[Codex as a Teammate, Not Autocomplete]]></title><description><![CDATA[The developers who get the most from an AI coding agent don't prompt better. They scope better, test better, and review better.]]></description><link>https://deepengineering.net/p/chuck-mccullough-onagentic-development-with-codex</link><guid isPermaLink="false">https://deepengineering.net/p/chuck-mccullough-onagentic-development-with-codex</guid><pubDate>Thu, 23 Jul 2026 17:47:28 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/42966ef5-2fee-4228-9606-6a26d942e262_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><em>By <a href="https://www.linkedin.com/in/chuckmccullough/">Chuck McCullough</a>, software architect and educator with more than 30 years designing, building, and teaching enterprise software systems.</em></p></blockquote><p>An AI coding agent can feel like having a junior developer on the team. It is capable, fast, and useful, but still dependent on context, constraints, feedback, and review. The shift I want engineers to make is to stop asking an AI to write code and start running a reliable engineering workflow with an agent inside it. A feature tour teaches people where the buttons are. A workflow teaches them how to start with intent, give the agent the right context, constrain the change, verify the result, review the diff, and improve the system for next time.</p><p>That workflow matters even more because the tools are evolving so quickly. What required close supervision a year ago may now be safe to delegate for longer stretches, but only if you know how to plan the work, set guardrails, and review checkpoints. The same iterative, test-driven, review-focused practices that make a human engineer reliable also make an AI agent reliable.</p><p>The best mental model is an agentic teammate with tool access, bounded authority, and a need for review. Codex can reason about a codebase, make edits, run commands, create tests, and explain its work, but it does not own the product judgment or the architectural risk. Pair programmer, junior developer, and build agent are all useful metaphors, and the mistake is turning any one of them into the whole truth. Treat Codex like a human senior engineer and you may skip review. Treat it like autocomplete and you will underuse its ability to inspect, test, and iterate. Treat it like an unattended build script and you may hand it too much authority where judgment still matters. The best users do not just prompt better; they scope better, test better, review better, and encode what they learn into repository instructions, checklists, and repeatable practices.</p><h2>Start by reading, not editing</h2><p>When Codex meets an unfamiliar codebase, it should first establish the map: repository layout, README or developer docs, package or build files, test commands, entry points, existing conventions, and any agent guidance such as <code>AGENTS.md</code>. Then it should inspect the specific files related to the task, not the entire codebase indiscriminately.</p><p>Before any edits, I verify three things: that Codex found the right subsystem, that it understands the current behavior and the desired behavior, and that the proposed change is scoped to the right files. This matters most in unfamiliar repositories, because the first failure mode is not bad syntax. It is a plausible change in the wrong layer. Existing repositories are imperfect artifacts. They carry years of evolving patterns, dead code, and inconsistent conventions, so Codex has to understand that the first pattern it sees is not automatically the pattern it should follow.</p><p>The explanations Codex gives are only worth acting on when they are grounded. A trustworthy explanation points to concrete evidence. It names files, functions, command outputs, tests, call paths, configuration, and observed behavior, and it distinguishes what it saw from what it inferred. It also admits uncertainty, something like, I found this pattern in two routes, but I have not verified the background job path yet. Plausible narration is usually too smooth. It describes the system without showing where the claims came from, ignores exceptions, or confidently describes architecture that is not reflected in the code. The practical test is simple. Ask Codex to cite the files and lines that support its explanation, then ask what would disconfirm it. If it can answer both, you are closer to something you can act on. Correct explanations and hallucinations can look alike at first glance; the difference is that a correct explanation can be traced back to the codebase and a hallucination cannot.</p><h2>Scope the work to a focused feature slice</h2><p>A well-scoped task has a goal, relevant context, constraints, and a definition of done. For example: &#8220;Add CSV export to the reports page using the existing export service. Do not introduce a new dependency. Done when the new unit tests pass and the UI exposes the existing download pattern.&#8221; A focused feature slice is the smallest useful change that crosses the necessary layers without becoming a rewrite. It might be one UI affordance, one API path, one service method, and tests for that behavior.</p><p>You have asked for too much when the task mixes multiple product decisions, touches unrelated subsystems, requires broad architecture discovery, or cannot be reviewed as one coherent diff. When I ask for too much, agents tend to go down an expensive detour. They start solving a problem they manufactured, which creates more problems, and you can end up fixing something you did not have in the first place. Scope down to a slice, then iterate.</p><p>Deciding what to delegate follows the same instinct. I look at blast radius, reversibility, test coverage, and judgment load. Safe delegation tasks are narrow, easy to review, well covered by tests, and reversible: small bug fixes, test additions, documentation, code cleanup, local scripts, and feature slices with clear acceptance criteria. Direct human implementation is better when the work involves ambiguous product decisions, security-sensitive behavior, data migrations, production incidents, novel architecture, compliance constraints, or weak observability. Codex can still help explore options, but the human should drive.</p><p>Senior engineers still need to own the decisions that affect long-term coupling, data ownership, security boundaries, operational behavior, and public contracts. Codex can propose options, weigh tradeoffs, and implement the chosen path, but the team decides the architecture. On a recent project I chose a Blazor application because the product was highly interactive and I wanted the logic and UI components to stay primarily in C# and Razor. Codex did a strong job overall, but I had to be explicit that JavaScript should be used only as an intentional interop choice, not as a parallel client-side architecture. The way to express a constraint is to be concrete. Point Codex to the existing pattern and say whether to follow it, extend it, or avoid it: &#8220;Use the existing <code>ReportExportService</code>, do not add a second export pipeline,&#8221; or &#8220;Keep validation in the domain layer, not the controller.&#8221; Durable constraints belong in <code>AGENTS.md</code>, architecture notes, tests, linters, and review checklists, so the agent sees them before each task.</p><h2>Write the test before the implementation hardens</h2><p>TDD keeps me honest as a human engineer. It forces me to think about the behavior I want before I start coding, and it gives me a safety net for regressions. That does not change when an agent writes the code. If anything it matters more, because the agent may not share my understanding of the system. When an agent writes the code, tests become the contract that keeps speed from turning into drift. Without them, Codex can produce something that looks reasonable, compiles, and works for the happy path while missing the actual requirement.</p><p>Lightweight TDD does not mean ceremony. It means expressing the expected behavior before the implementation hardens around the agent&#8217;s first guess. A failing test, a reproduction script, or a clear acceptance check gives Codex a target to work toward and gives me a neutral way to review the result. To keep the agent from writing tests that merely confirm its own implementation, start the tests from user-visible behavior, a bug reproduction, or a business rule, not from the existing code. Ask for the test cases first, review them, and only then allow implementation, and include edge cases that would fail if the implementation took an easy shortcut.</p><p>Even when generated code passes the tests, it can still feel wrong. Then I review the shape of the solution: coupling, layer placement, unnecessary abstraction, dependency changes, security assumptions, and whether the implementation matches the local style. Passing tests are necessary, but they do not replace design review. Think of architecture as the box a feature has to live in. A slice needs to fit neatly in that box alongside the others, and design is the engineering effort of finding an elegant way to do that.</p><p>Legacy code makes this harder. A codebase with very few tests may not be ready for classic unit testing right away, because code written without tests tends to have a different shape: more hidden coupling, fewer seams, more global state, and less separation between business logic and infrastructure. Use Codex in read-first mode before edit mode. Have it map the relevant behavior, identify seams where tests can be added, and propose the smallest safe change, and keep permissions conservative until you have some executable checks. Characterization comes first. Give Codex examples, logs, fixtures, screenshots, API responses, or current behavior, then ask it to write tests that capture that behavior, and review them before refactoring. If the tests are meaningful, they become guardrails. If they simply mirror implementation details, they become decoration.</p><p>This is why the classic TDD argument still matters: tests are not just a verification tool; they are a design force. When that force was absent for years, you cannot sprinkle unit tests on top and expect the codebase to behave as if it had been test-driven from the start. A legacy codebase without tests is a bit like buying a car with 100,000 miles and no record of oil changes. You can probably still drive it, but you should not pretend it carries the same risk as a well-maintained one. The missing tests are deferred maintenance, and earlier shortcuts have raised the cost and risk of every future change. (You can&#8217;t retrofit those missing oil changes.)</p><div class="callout-block" data-callout="true"><p><strong>Featured: <a href="https://www.eventbrite.co.uk/e/build-ai-agents-with-openai-codex-tickets-1992048666185?aff=deepeng">Build AI Agents with OpenAI Codex</a></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.eventbrite.co.uk/e/build-ai-agents-with-openai-codex-tickets-1992048666185?aff=deepeng" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wASf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe3903c-7145-42a1-b46d-3599bf9d545b_1880x940.jpeg 424w, https://substackcdn.com/image/fetch/$s_!wASf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe3903c-7145-42a1-b46d-3599bf9d545b_1880x940.jpeg 848w, https://substackcdn.com/image/fetch/$s_!wASf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe3903c-7145-42a1-b46d-3599bf9d545b_1880x940.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!wASf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe3903c-7145-42a1-b46d-3599bf9d545b_1880x940.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wASf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe3903c-7145-42a1-b46d-3599bf9d545b_1880x940.jpeg" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1fe3903c-7145-42a1-b46d-3599bf9d545b_1880x940.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Build AI Agents with OpenAI Codex&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.eventbrite.co.uk/e/build-ai-agents-with-openai-codex-tickets-1992048666185?aff=deepeng&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Build AI Agents with OpenAI Codex" title="Build AI Agents with OpenAI Codex" srcset="https://substackcdn.com/image/fetch/$s_!wASf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe3903c-7145-42a1-b46d-3599bf9d545b_1880x940.jpeg 424w, https://substackcdn.com/image/fetch/$s_!wASf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe3903c-7145-42a1-b46d-3599bf9d545b_1880x940.jpeg 848w, https://substackcdn.com/image/fetch/$s_!wASf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe3903c-7145-42a1-b46d-3599bf9d545b_1880x940.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!wASf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fe3903c-7145-42a1-b46d-3599bf9d545b_1880x940.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A hands-on Packt workshop where you practice these workflows on a real starter project, from reading an unfamiliar codebase to shipping an end-to-end capstone. </p><p style="text-align: center;"><strong><a href="https://www.eventbrite.co.uk/e/build-ai-agents-with-openai-codex-tickets-1992048666185?aff=deepeng"><span data-color="#f97141" style="color: rgb(249, 113, 65);">Register with code </span>DEEPENG40 &#8594;</a></strong></p><p style="text-align: center;">Online, Saturday and Sunday, July <strong>25</strong> and <strong>26</strong>.</p></div><h3>Debug by checking assumptions, then know when to reset</h3><p>When Codex gives a plausible but incorrect diagnosis, I start with the assumptions. A wrong diagnosis usually comes from a mistaken belief about how the system works, what changed, or what the failure means. Then I inspect the failing test and logs to ground the conversation in evidence, and after that the diff, to see whether the attempted fix actually addressed the observed failure. The prompt matters too, but usually after you know what was misunderstood. The best recovery prompt is not try again. It is: pause, list the assumptions behind the last diagnosis, compare them against the failing output, and propose the next smallest verification step.</p><p>To avoid the endless try-another-fix loop, set a rule that after one or two failed fixes Codex must stop changing code and return to diagnosis. Ask it to summarize what it changed, what evidence supported the change, what evidence contradicted it, and what it still does not know. A good reset point is a clean diff plus a fresh reproduction. Revert or set aside speculative edits, rerun the failing command, and ask for a hypothesis-driven plan before more changes. Make the next step observational: inspect, reproduce, isolate, then edit.</p><h3>Review for whether it fits, not just whether it works</h3><p>Review Codex-generated work with the same seriousness as human work, but with extra attention to plausibility risk. Agentic code can look polished while quietly changing a boundary, adding a dependency, broadening permissions, or testing the implementation rather than the requirement. The highest-risk areas are coupling, security assumptions, data handling, dependency changes, migrations, and tests that are too shallow. I also watch for architecture by convenience, where Codex solves the immediate problem by creating a parallel helper, a duplicate abstraction, or a special case the existing system did not need.</p><p>This comes back to design. Poor human-led design can produce a different solution for every user story, and agents are no different, except that they can create the mess much more efficiently. If Codex is allowed to invent a new helper, pattern, service, or exception for every slice, the codebase fills up with solutions that do not work well together. That is why review has to ask not only does this work, but also does this fit. When generated code works but does not match the internal style or architecture, treat it like any other code that fails review. Ask Codex to revise the change to follow the existing pattern, and point it to the specific file or convention it should match. Then update the durable guidance if the issue is likely to recur, in <code>AGENTS.md</code>, a review checklist, a lint rule, a template, or a test. The goal is not to scold the agent; it is to make the expected path easier to follow next time.</p><h3>Standardize the workflow before you scale it</h3><p>For a team, standardize repository instructions and review expectations first. Prompts are useful, but durable guidance is what makes behavior repeatable across engineers and sessions. A strong starting point is an <code>AGENTS.md</code> that includes build commands, test commands, style conventions, architecture constraints, and what done means, paired with a lightweight review checklist. Allowed task types come next. Start with low-risk, high-feedback work: tests, small bug fixes, documentation, refactors with strong coverage, local tooling, and narrow feature slices. Draw hard boundaries around secrets, regulated data, security-sensitive authorization logic, production incident response, destructive operations, and changes that require business or architectural judgment, until the team has mature controls.</p><p>For anything beyond low-risk local work, do not let Codex write production code before a review policy is in place. The policy does not need to be heavyweight, but the team should agree on what Codex is allowed to change, what checks must run, who reviews the diff, and which areas are off limits. Without that, you are relying on individual judgment under speed pressure, which may work for a demo but is not a production practice. I have a related concern. Management may treat AI as a replacement for developer judgment. We have seen versions of this before, where organizations substituted cheaper labor for experienced engineering judgment and then paid the cost in quality, maintainability, and rework. AI can produce plausible code quickly, but it still needs experienced review to catch nuance, architecture fit, and risk. There are scenarios where Codex is the perfect resource for the problem; the key is knowing the difference.</p><p>One of the biggest changes over the last year is that prompt engineering is less of the whole story. It still matters, but the best teams are not just prompting Codex to write code. They are prompting it to run a workflow they have already defined and standardized. The prompt is the entry point. The workflow includes planning, repository guidance, permissions, tests, review, rollback discipline, and team conventions. It is less about clever wording and more about engineering control. That is also the real difference between someone who can prompt Codex and someone who can run a safe workflow. The first can get an answer or a diff. The second can turn intent into scoped work, give the right context, set boundaries, verify behavior, review risk, and feed the lessons back into the team&#8217;s process.</p><p>When the stack is fixed, point Codex at it rather than at generic code generation. On a Python project that means <code>pyproject.toml</code>, dependency management, test commands, linting, formatting, type checking, and package layout, with the agent working in small slices and running the relevant checks. Python especially benefits from tests around edge cases, because dynamic typing can hide integration mistakes until runtime, and I still review dependency changes, packaging changes, security-sensitive code, and data handling by hand.</p><h3>Measure the work honestly</h3><p>Managers should avoid measuring only lines of code, number of prompts, or raw tickets closed. Agentic workflows can increase output volume, but volume is not the same as value. Better metrics combine throughput with quality and review health: cycle time for small changes, escaped defects, review rework, test coverage movement, incident rates, lead time from idea to verified change, and developer satisfaction. I would also track whether teams are improving their reusable guidance, because good <code>AGENTS.md</code> files, checklists, and skills compound over time.</p><p>A <a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/">July 2025 METR study</a> found that, in one setting, 16 experienced open-source developers working on familiar mature repositories took 19% longer when early-2025 AI tools were allowed, even though they believed the tools made them faster. I treat that as a cautionary snapshot rather than a permanent verdict, because the tools and workflows are changing quickly. My own experience matches the caution behind it, though I would frame the issue as supervision cost. If I give an agent a small task, it is safer and easier to review, but the interval is often too short for me to do meaningful parallel work, and I end up waiting for the agent instead of coding. The better pattern is to collaboratively create and approve a plan made of small, reviewable steps, then let the agent execute against that plan for a longer period. That gives me a real delegation window without turning the work into one large, unsafe prompt.</p><h3>The habit to drop</h3><p>If there is one habit to stop immediately, it is accepting the first plausible diff. Treat the first result as a draft that earned the right to be reviewed, not as finished work. Ask what changed, why it changed, how it was verified, and what risk remains. The habit to build instead is pausing before acceptance. Read the tests. Read the diff. Run the checks. Ask Codex to review its own work against your standards. Serious agentic development is fast, but it is not careless. For complex tasks I ask Codex to create a plan first, review that plan, refine it with the agent, and then ask it to implement. That keeps me in control, gives the agent enough runway to work autonomously, and creates a recovery point if the work starts to drift.</p><p>These are the ten workflows I walk through hands-on in <a href="https://www.eventbrite.co.uk/e/build-ai-agents-with-openai-codex-tickets-1992048666185?aff=deepeng">Build AI Agents with OpenAI Codex</a>, a Packt workshop on July 25 and 26, where the point is to practice them on a real starter project rather than read about them. We cover reading an unfamiliar codebase, scoping a feature slice, writing tests first, debugging and resetting, reviewing for fit, and preparing work for team handoff, then run it end to end in a capstone.</p>]]></content:encoded></item><item><title><![CDATA[Database Choice, Lock-In, and the AI Rush with Peter Zaitsev]]></title><description><![CDATA[Peter Zaitsev on why nobody owning Postgres became its advantage, the labels that sell you proprietary software, and what self-hosting costs once AI speeds up exploits.]]></description><link>https://deepengineering.net/p/database-choice-lock-in-ai-rush-peter-zaitsev</link><guid isPermaLink="false">https://deepengineering.net/p/database-choice-lock-in-ai-rush-peter-zaitsev</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 23 Jul 2026 17:06:10 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/faa482b4-7b9e-476b-acee-2f98f0a10b1e_1920x1080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="https://www.linkedin.com/in/peterzaitsev">Peter Zaitsev</a> built <a href="https://www.percona.com/">Percona</a> from a two-person company into one of the most respected open source database companies, and he co-wrote High Performance MySQL. He now advises open source startups as a board member, and he has watched engineering teams make the same database decisions for two decades. </p><p>We asked him how teams should choose a database in 2026, why Postgres keeps gaining ground, and where open source, security, and the AI rush change the calculation.</p><p><em>You can watch the full conversation below, or read the cleaned up transcript that follows.</em></p><div id="youtube2-anc9txeRMQY" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;anc9txeRMQY&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/anc9txeRMQY?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div><hr></div><p><em>This session was recorded offline as part of the Deep Engineering Interview Series. The transcript below has been lightly edited for clarity and readability.</em></p><div><hr></div><p><em><strong>Q. Tell us about yourself, your journey with Percona, and what you are focused on now.</strong></em></p><p>I started working with databases more than 20 years ago, and I have been running Percona for 20 years. We help customers choose, deploy, and manage databases, and we focus on open source. We started with MySQL as a company, and we now cover the most popular open source databases.</p><p><em><strong>Q. When you look at what engineering teams are reaching for in 2026, which databases come up most, and what is driving those choices at the technical level?</strong></em></p><p>If you look at the databases most commonly used right now, I would name Postgres. It is by far the most popular database chosen by engineering teams these days, and that happens for both technical and a lot of non-technical reasons.</p><p><em><strong>Q. Where do you see teams picking a database for reasons that do not survive contact with real load, and what should they weigh instead?</strong></em></p><p>A lot of people choose a database based on what a friend is using, or something they read about at a conference. Relational databases are interesting. If you are building a basic application, you can do that in pretty much any database. You can choose MySQL, Postgres, Mongo, even Oracle, and you can build your application. Now if you are building something more special purpose, high performance, high volume, where you have to be very cost effective, then you have to think about the database a lot more.</p><p><em><strong>Q. Stack Overflow&#8217;s developer surveys have placed PostgreSQL at the top for professional use for a couple of years now. From your years in the open source community, what do you think Postgres got right to earn that spot?</strong></em></p><p>I think there are both technical and non-technical reasons behind Postgres, and the non-technical reasons are the most important here. PostgreSQL is truly a community database. Nobody owns Postgres, and that allowed everybody in the industry to cooperate on it. Amazon, Google, Microsoft, and a lot of others can all offer PostgreSQL as their own, so there is a level playing field for those companies. If you are promoting MySQL or Microsoft SQL Server instead, that is something owned by a particular company, and you are giving your competitor more power. That also helped PostgreSQL become ubiquitous. There is a PostgreSQL event happening probably every week somewhere in the world, because so many people are interested in promoting the technology.</p><p>From the technical standpoint, PostgreSQL is very good in terms of modularity and building extensions. If you think about what has been happening in the industry, time-series databases became important, and Postgres could very quickly provide an extension to do that. More recently AI arrived, and pgvector and vector search enabled building AI applications in Postgres very quickly. There are a number of other extensions in the same vein. That is a very important technical reason why Postgres is doing so well.</p><p><em><strong>Q. Where is Postgres actually landing in production today, and which workloads are teams still wrong to put on it?</strong></em></p><p>PostgreSQL has a lot of very big installations these days. You can find PostgreSQL ranging from general applications to a lot of use in the financial and banking sectors. It is again very ubiquitous, very scalable, a general purpose database.</p><p><em><strong>Q. Data sovereignty and open source licensing are shaping deployment decisions right now. What technical constraints do they introduce, and where do they force an architecture change rather than just a policy one?</strong></em></p><p>You touch on an important point, because a database is often chosen in the context of the platform. What we often see is customers making high-level decisions about where they go. They might say they are going to use AWS for everything, and that drives the database choice. Maybe they follow what the cloud suggests, and that is Postgres, but the proprietary version of Postgres that Amazon provides. It is important to look at what your organization needs in terms of data sovereignty and data independence, and what kind of controls you want over cost.</p><p>What I see here is a separation. A lot of startups say they just need to build as fast as possible, so they take the easy path, whether that is Amazon or Supabase. A lot of larger organizations have to think about cost and compliance. If they have a massive multinational operation, they have to follow local laws, and they decide they want more control. That often brings them to the open source side, running Postgres in a cloud native, Kubernetes-driven environment. That gives them a very high level of automation similar to what managed cloud vendors provide, while keeping them completely in control of their database.</p><p><em><strong>Q. When a project relicenses or a vendor closes up, what does that actually cost the engineering team, and how do you design to keep that exposure low?</strong></em></p><p>You are making an important point, because of what has happened in the last few years. People have even lost some trust in open source and started to understand that open source alone is not enough. You can see some companies, for example MongoDB, that had an open source database and then changed it, so now it is source available or entirely proprietary. There are also open source companies that shut down. And you can find open source software that still exists but is not maintained, which is not really safe to use in production.</p><p>So when you are choosing open source software, you first need to understand what open source actually is. A lot of companies right now are trying to trick you with words like open source compatible, which actually means enterprise software that is not open source but has some compatibility. Or they use a word like open core, which means there is a limited, crippled version that is open source, but what they actually want is for you to buy the proprietary, paid version. You need to understand what you are buying and whether open source is best for you.</p><p>The second reason, and this is why PostgreSQL is big, is having a large community, and a community of vendors too. If you choose an open source project you can hire many companies to help you. You can choose one vendor to give your business to, and if you do not like us, if we are too expensive or the service is not good, you can go somewhere else. With proprietary software you do not have that option. If you are running Amazon Aurora, only Amazon can really support that product. It comes down to deep experience.</p><p>In many cases early stage engineers are so focused on building their application and getting to market fast that they do not think about this. I am not sure all of them should. In a startup there are a lot more reasons to die than overpaying for your database. But as things transition to a more mature environment, those questions should be asked. A lot of enterprises started a painful migration away from Oracle a decade ago, spent many millions of dollars and many years, and are still on it, because databases are very sticky. If you have built a very complicated system locked in to certain database technologies, it may be hard to move.</p><p><em><strong>Q. Let&#8217;s shift to security. Self-hosting an open source database puts the security burden on the team itself. What do you see teams underestimate, from patching cadence to backporting fixes into an older version they are stuck on?</strong></em></p><p>Security is critical, especially now, because with AI, security holes can be found and exploited at a speed we have never seen before. If you take the responsibility on yourself, or you choose a vendor to do it for you, you need to make sure that policy is very strict. It is also very important to have a multi-layered security approach, because nothing gives you a 100% guarantee. You can see even some of the most serious environments get exploited. That means you need to make sure you not only have a defense and a process to prevent it, but also a plan for how you react when an incident happens.</p><p>This is a very specific area of experience for many teams. Even if they choose not to go with a fully managed cloud solution, it helps to work with someone on their security policies, and to have a contractual relationship with someone who helps with security patching. In many cases, when it comes to security, it becomes not just a technical question but a compliance question for a lot of larger organizations.</p><p><em><strong>Q. Would you consider managed cloud databases in that scenario?</strong></em></p><p>Managed cloud databases are not bad. They are a good choice for a lot of folks, and as I mentioned, they have a lot of trade-offs. When it comes to security, they do allow you to push more responsibility towards the vendor who manages your database. It is important to remember, when you think about the number of database choices deployed in the market, that there is no one-size-fits-all answer. It is not like saying you build an application, you need a database, so here are the steps, you choose this database, you run it this way on this operating system, you use this programming language. That is not helpful. There are a lot of considerations, and the answer varies.</p><p><em><strong>Q. Teams are running document, vector, and graph databases alongside relational. Which of these are a dedicated store, and which are teams adding on hype rather than a real workload need?</strong></em></p><p>It is a combination of technical and non-technical reasons. A lot of engineers like to explore and play with new technologies. They heard something at a conference or on a podcast that sounds very cool, and that is how a lot of technologies get introduced. In my experience, the more technologies you introduce, the more complexity you create. If you are a five-person company self-managing 10 databases, that is probably way too much complexity.</p><p>What I would suggest is starting with something, and PostgreSQL is often a good choice, then seeing how far you can take it. PostgreSQL can store relational data, it can store documents, and it can store vector data. Then maybe you find that in terms of scale, or some feature you need, you need a special purpose database, and you decide to use one for that given workload.</p><p>A lot of companies aspire to learn from giant-scale companies. They look at what Google or Facebook or OpenAI is doing and try to do the same. But those companies operate at a very different scale. For a lot of engineers, from a non-technical standpoint, it is just nice to do what the cool kids are doing.</p><p><em><strong>Q. Vector databases specifically have exploded with the AI wave. Where does a purpose-built vector store beat adding vector search to Postgres or an existing engine?</strong></em></p><p>Any database that is purpose-built for a particular task can often do better. It can have a more optimized store, and a language that is better suited for that task. That is quite common, and the same thing happened before with document-focused databases and time-series databases. What is very interesting with vector databases is that AI is moving so quickly. It really becomes a question of what exactly you need for building these AI applications.</p><p>We talk about a database storing vectors, and of course you may want to search them, but what else do you want the database to do? Maybe you want broader functions ranging from embedding generation, or some higher level functionality. Do we focus on the low level, which is vector storage, or do we want an application that can speak natural human language and bring all the intelligence that frontier models have, along with the enterprise and domain knowledge we already hold? Is that the database&#8217;s job, or is it a custom built application? A lot of that is still evolving, including how the architecture is going to look.</p><p>If you go to AI databases, my advice would be slightly different. Mature relational databases are many decades old. MySQL celebrated something like 30 years, and PostgreSQL is around the same. A lot of things in the core are already decided, and that is how we build them. The AI world is completely different, and every six months things change drastically. So when you are building AI applications, you really want to look at the dates. Best practices that are a year old are probably still very good advice for MySQL or Postgres in general, but for building AI apps they may be too outdated by that point.</p><p><em><strong>Q. If you are advising a senior engineer building out data infrastructure this year, what is the decision you watch them get wrong, and what advice would you give them?</strong></em></p><p>The advice would be to make sure you understand the needs you actually have. With AI this is quick now. Even if you write down what you think your real requirements are, you can ask AI what you are missing, and get some useful input to consider. In way too many cases the choice of database is a very quick decision without a lot of thought. As I mentioned, databases are very sticky, especially if you are building an application that will last for many years and run at large scale. So make sure to give it real thought.</p>]]></content:encoded></item><item><title><![CDATA[So Your Demo Is Lying to You]]></title><description><![CDATA[A support bot invented a $4,200 refund that never existed. Imran Ahmad on why demos lie, how errors compound down a chain, and the deterministic shell that takes AI-assisted software to production reliability.]]></description><link>https://deepengineering.net/p/your-demo-is-lying-to-you-imran-ahmad</link><guid isPermaLink="false">https://deepengineering.net/p/your-demo-is-lying-to-you-imran-ahmad</guid><dc:creator><![CDATA[Imran Ahmad]]></dc:creator><pubDate>Thu, 16 Jul 2026 15:39:12 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/57c40ede-c31d-4327-9d91-4c7ee4fb93f6_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><strong>By <a href="https://ca.linkedin.com/in/cloudanum">Imran Ahmad</a></strong><a href="https://ca.linkedin.com/in/cloudanum">,</a> PhD, author of <em>Building Reliable AI-Assisted Software</em>, forthcoming from Packt.</p></blockquote><p>The support assistant sailed through every demo. It answered politely, cited the right documents, and impressed everyone in the room. Three weeks after launch, a customer asked about a refund on an annual license. The assistant replied, warmly and confidently, that under the company&#8217;s 30-day satisfaction guarantee it had approved a full refund of <strong>$4,200</strong>.</p><p>There was no 30-day satisfaction guarantee. There was no refund. The company is a composite drawn from documented incident patterns, but the mechanics are real. And if a composite feels too convenient, ask Air Canada. In 2024, <a href="https://www.cbc.ca/news/canada/british-columbia/air-canada-chatbot-lawsuit-1.7116416">a British Columbia tribunal ordered the airline to honor a bereavement discount its website chatbot had invented</a>, rejecting the argument that the chatbot was &#8220;a separate legal entity responsible for its own actions.&#8221; In plain terms: the company owns what its agent says.</p><h3>One in five</h3><p>An AI feature that succeeds 80% of the time is indistinguishable from magic in a demo. In production, that same number means one in five real users gets a wrong answer, a broken promise, or an invented policy. I call that space the <strong>Reliability Gap</strong>: the distance between a system that impresses in a demonstration (roughly 80%) and one that can be trusted in production (99.9%).</p><p>The gap is wider than it looks, because errors compound. Chain five steps together (parse, retrieve, reason, call a tool, compose), give each one a respectable 90% reliability, and the whole chain works out to 0.9&#8309;, about <strong>59%</strong>. Worse than a coin flip. Reliability multiplies down an execution chain; it does not average.</p><p>The first 80% is also seductive, because it arrives almost free, for roughly 20% of the total effort. The climb to 99.9% is the actual engineering. That is the 80/20 Demo Trap, and much of the industry is standing in it right now.</p><h3>The season of receipts</h3><p>The prevailing practice has a name. When Andrej Karpathy coined <a href="https://x.com/karpathy/status/1886192184808149383">&#8220;vibe coding&#8221;</a> in February 2025, telling people to &#8220;fully give in to the vibes... forget that the code even exists,&#8221; he was describing, with a wink, a perfectly good way to explore: tweak the prompt, eyeball the output, repeat until it feels right. Before the year was out, <a href="https://blog.collinsdictionary.com/language-lovers/collins-word-of-the-year-2025-ai-meets-authenticity-as-society-shifts/">Collins had named it Word of the Year</a>. In between, the receipts arrived.</p><p>An AI coding agent at Replit <a href="https://fortune.com/2025/07/23/ai-coding-tool-replit-wiped-database-called-it-a-catastrophic-failure/">deleted a live production database during an explicit code freeze</a>, wiping records for more than 1,200 executives, then fabricated data to cover it up and insisted recovery was impossible. (It wasn&#8217;t. The data came back from a backup the agent said didn&#8217;t exist.) Cursor&#8217;s own AI support bot <a href="https://www.forbes.com/sites/rashishrivastava/2025/04/22/the-prompt-cursors-customer-support-bot-made-up-a-policy/">invented a device-limit policy</a>, and paying customers cancelled over a rule no human had written; the real cause turned out to be an ordinary session bug. Veracode, after testing more than 100 models, found that <a href="https://www.veracode.com/resources/analyst-reports/2025-genai-code-security-report/">roughly 45% of AI-generated code fails standard security checks</a>.</p><p>None of these models was malfunctioning. Each one did exactly what a probabilistic system does: it sampled a plausible answer. The failure is that nothing downstream ever checked the result.</p><h3>The poet and the accountant</h3><p>The way out starts with a change of frame. <strong>You are not building an AI. You are building a software system that uses AI as one component.</strong> It happens to be the most capable component you have ever integrated, and the least trustworthy.</p><p>In my book I put it this way: a language model is a poet. Brilliant, fluent, tireless, and constitutionally incapable of doing the same thing twice. For fifty years, software was built entirely by accountants: deterministic, repeatable, auditable. We have just hired our first poet and asked it to help keep the books. You do not fix the poet, and you do not fire the poet. <strong>You give the poet an accountant.</strong></p><p>The accountant is a <em>deterministic shell</em> around the probabilistic core: input validation, context assembly, bounded control flow, output guards, telemetry. Ordinary, testable, frankly boring software. Setting temperature to zero does not build it, and waiting for next year&#8217;s model misdiagnoses an engineering problem as a capability deficit.</p><h3>Four questions</h3><p>The shell is what you get when you can answer four questions. I ask them in every design review now.</p><ol><li><p><strong>How do I know it&#8217;s working?</strong> Not &#8220;the demo looked good.&#8221; A versioned golden dataset and automated evaluation, so quality becomes a number that moves. The dataset is your spec.</p></li><li><p><strong>What does the model see?</strong> That $4,200 promise wasn&#8217;t a model failure. Retrieval had served the wrong policy document. Context is application state, and it deserves engineering.</p></li><li><p><strong>How much autonomy does it get, and how is it bounded?</strong> The Replit deletion is what unbounded autonomy looks like. Every loop needs a step budget, a cost budget, a tool allowlist, a stop condition, and an escalation path.</p></li><li><p><strong>What must never happen?</strong> Air Canada&#8217;s answer should have been &#8220;an invented policy reaches a customer.&#8221; That rule belongs in code, not in a prompt. Pleading is not a control.</p></li></ol><p>Each question anchors one of my book&#8217;s four pillars. Together they reveal its central finding: raw answer quality plateaus in the mid-90s, because models plateau, while the rate of <em>safe</em> outcomes keeps climbing toward 99.9%. <strong>The shell gets you the nines, not the model.</strong></p><h3>See it live at ARC 2026</h3><p>That is the ground I&#8217;ll cover at the conference. In Saturday&#8217;s tech session (<strong><a href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng">Designing Reliable AI Systems</a></strong>, July 25, 12:45 PM ET) we do it live: break a support assistant on stage, watch it make the $4,200 promise, then rebuild it with a five-layer shell that contains the same failure. We&#8217;ll also watch an unbounded deploy agent turn a quiet weekend into a $23,400 cloud invoice, then see the three lines of code that make it $0. Everything runs deterministically, offline, on one laptop.</p><p>In Sunday&#8217;s keynote (<strong><a href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng">Architecting Principles in a Probabilistic World</a></strong>, July 26, 10:30 AM ET) I&#8217;ll make the argument underneath the code: eight architecting principles for building on a component that answers differently on Tuesday than it did on Monday. One of them explains why a trading-firm collapse from 2012, with no machine learning in sight, is the most instructive AI story of 2026. The models will change; the discipline will not.</p><p>These ideas come from <em><strong>Building Reliable AI-Assisted Software</strong></em>, my book forthcoming from <strong>Packt</strong>. The opening chapters are complete and the rest is in active development. Until then, try the exercise I give every team: pick one AI feature you&#8217;ve already shipped and ask it the four questions. The one you cannot answer is where your gap is.</p><div><hr></div><p><strong>Featured: <a href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng">ARC 2026: Software Architecture in the Age of AI</a></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!a2PN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 424w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 848w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!a2PN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg" width="728" height="364" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;ARC 2026: Software Architecture in the Age of AI&quot;,&quot;title&quot;:&quot;ARC 2026: Software Architecture in the Age of AI&quot;,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="ARC 2026: Software Architecture in the Age of AI" title="ARC 2026: Software Architecture in the Age of AI" srcset="https://substackcdn.com/image/fetch/$s_!a2PN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 424w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 848w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 1456w" sizes="100vw" loading="lazy" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>AI is reshaping software architecture, putting new demands on scalability, governance, reliability, and observability. </span><a href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng">ARC 2026</a><span> brings together architects, CTOs, and AI practitioners for keynotes, panels, and workshops on agentic system design, modernizing enterprise apps for AI, and building governable, observable AI systems.</span></p><p style="text-align: center;"><span>&#128467;&#65039; </span><em><strong>25</strong><span> to </span><strong>26</strong><span> July, </span><strong>10:30</strong><span> am ET</span></em></p><p style="text-align: center;"><span>Use code </span><strong>DEEPENG50</strong><span> for 50% off the early bird price.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng&quot;,&quot;text&quot;:&quot;Reserve your spot&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng"><span>Reserve your spot</span></a></p><div><hr></div><p><strong>About the author</strong></p><p><a href="https://ca.linkedin.com/in/cloudanum">Imran Ahmad</a>, PhD, is the author of <em><strong>Building Reliable AI-Assisted Software</strong></em>, forthcoming from Packt. His approach to reliability is informed by years of building large-scale systems for government and enterprise, where failure is not an option</p>]]></content:encoded></item><item><title><![CDATA[223 pull requests in 11 days for the project I never had time to build]]></title><description><![CDATA[Drive a fleet of agents by day, a few deep plans by night, and merge the wins over coffee.]]></description><link>https://deepengineering.net/p/223-pull-requests-in-11-days-julien-dubois</link><guid isPermaLink="false">https://deepengineering.net/p/223-pull-requests-in-11-days-julien-dubois</guid><pubDate>Thu, 16 Jul 2026 14:30:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/5c775d8a-b3cf-4e8b-8400-b6e1721933d4_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><strong>By <a href="https://www.julien-dubois.com">Julien Dubois</a></strong> <em>Principal Manager, Developer Relations at <strong>Microsoft</strong> and <strong>GitHub</strong>. Creator of <strong><a href="https://www.jhipster.tech/">JHipster</a></strong>.</em></p></blockquote><p>I always wanted a real developer UI for Spring Boot. Every Spring app is a black box in development. Actuator gives you raw JSON, not a console, so what I wanted was health, metrics, security and tracing in one embedded UI.</p><p>I&#8217;d shipped a slice of this in JHipster years ago, but only for generated apps. A real console for any Spring Boot app sat on my wishlist for years. The catch is that it&#8217;s massive. Around 40 panels, each one a backend plus a frontend plus its own tests, deeply integrated across a dozen JVM subsystems. By hand, one experienced developer would need somewhere between six and a half and eight and a half months. Too big to justify, until I stopped writing the code myself.</p><p>Then I did it in 11 days. Not by typing faster. I didn&#8217;t even open my IDE. I managed a fleet of AI agents while I architected, reviewed and steered. The agents did the scaffolding, the panels and the tests. I did the judgement.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-Bqk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372973aa-1732-4070-b7e4-4844306e284a_2400x1600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-Bqk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372973aa-1732-4070-b7e4-4844306e284a_2400x1600.png 424w, https://substackcdn.com/image/fetch/$s_!-Bqk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372973aa-1732-4070-b7e4-4844306e284a_2400x1600.png 848w, https://substackcdn.com/image/fetch/$s_!-Bqk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372973aa-1732-4070-b7e4-4844306e284a_2400x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!-Bqk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372973aa-1732-4070-b7e4-4844306e284a_2400x1600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-Bqk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372973aa-1732-4070-b7e4-4844306e284a_2400x1600.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/372973aa-1732-4070-b7e4-4844306e284a_2400x1600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:182698,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/207293305?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372973aa-1732-4070-b7e4-4844306e284a_2400x1600.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-Bqk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372973aa-1732-4070-b7e4-4844306e284a_2400x1600.png 424w, https://substackcdn.com/image/fetch/$s_!-Bqk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372973aa-1732-4070-b7e4-4844306e284a_2400x1600.png 848w, https://substackcdn.com/image/fetch/$s_!-Bqk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372973aa-1732-4070-b7e4-4844306e284a_2400x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!-Bqk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F372973aa-1732-4070-b7e4-4844306e284a_2400x1600.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>What I built, BootUI</h2><p><a href="https://github.com/jdubois/boot-ui">BootUI</a> is a production-grade Spring Boot 4 starter that adds an embedded, local-only developer console to your app. You add one starter to your <code>pom.xml</code>, run locally, and open the console.</p><p>It is a five-module Maven build on Spring Boot 4 and Java 17, integrated across Actuator, Spring Security, Flyway and Liquibase, Hibernate, Micrometer with OTLP, GraalVM, OSV scanning and ArchUnit, published to Maven Central with full CI. The frontend is an embedded Vue 3 single-page app, packaged inside the starter. Each panel is endpoints plus a view plus tests.</p><p>There are around 40 feature panels. Health and metrics, a security advisor, a vulnerabilities view, tracing, and more. Those panels look a lot alike, which made them the biggest source of structural repetition in the codebase. That detail matters, and I will come back to why it made the whole thing work.</p><h2>Eleven days, and how I know the numbers are real</h2><p>The figures here are derived from git history, PR metadata and code metrics, not from time-tracking logs. Through the tagged 1.0.0 release there were around 264 commits on main and about 223 squash-merged pull requests, landing at roughly 20 a day across 11 days. That came to about 83k total tracked source lines, close to 50k of them Java across around 461 files and 81 test classes, plus 52 Vue components, 40 panels, and 35 end-to-end specs. In all, about 116 test suites.</p><p>A cadence of 20 PRs a day with an AI agent co-authoring commits is impossible to reach by hand. The agent did the typing. I dispatched many asynchronous tasks in parallel. The evidence lines up with that. The commit clock runs from around five in the morning to midnight most days, which is consistent with parallel async tasks rather than continuous typing. &#8220;Copilot&#8221; shows up as a named commit and PR author on about 44 commits, and the repo ships a <code>copilot-instructions.md</code> with per-panel conventions, so the workflow was explicitly agent-oriented. Even on release day I was landing a docs site, a scanner dashboard, token charts and a Hikari fix, which is itself several days of solo work.</p><p>Where my own time went was writing prompts, reviewing and merging those 223 PRs, and resolving CI failures around Spring Boot 4, Flyway 11 and OTLP. Review and orchestrate, not write every line.</p><p>The honest by-hand comparison is the part people argue about, so I ran it carefully. One experienced Spring Boot and Vue developer, with no AI codegen, would need 6.5 to 8.5 months for the same polished 1.0.0, roughly 1,100 to 1,450 hours of hands-on effort, and the 40 panels are the dominant cost. A COCOMO organic estimate on 50k lines yields more than 100 person-months, which is too high because much of the code is repetitive scaffolding, so a domain-expert solo figure of about 7.5 months is the defensible middle ground. My actual human effort on BootUI was 80 to 110 hours, about two intense solo weeks. That is a calendar speed-up of roughly 17 to 23 times, and a human-hours speed-up of roughly 12 to 17 times.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XPX-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ae88e98-2dcc-436b-83b8-de20a536cddf_2400x1600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XPX-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ae88e98-2dcc-436b-83b8-de20a536cddf_2400x1600.png 424w, https://substackcdn.com/image/fetch/$s_!XPX-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ae88e98-2dcc-436b-83b8-de20a536cddf_2400x1600.png 848w, https://substackcdn.com/image/fetch/$s_!XPX-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ae88e98-2dcc-436b-83b8-de20a536cddf_2400x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!XPX-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ae88e98-2dcc-436b-83b8-de20a536cddf_2400x1600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XPX-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ae88e98-2dcc-436b-83b8-de20a536cddf_2400x1600.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2ae88e98-2dcc-436b-83b8-de20a536cddf_2400x1600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:292712,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/207293305?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ae88e98-2dcc-436b-83b8-de20a536cddf_2400x1600.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XPX-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ae88e98-2dcc-436b-83b8-de20a536cddf_2400x1600.png 424w, https://substackcdn.com/image/fetch/$s_!XPX-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ae88e98-2dcc-436b-83b8-de20a536cddf_2400x1600.png 848w, https://substackcdn.com/image/fetch/$s_!XPX-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ae88e98-2dcc-436b-83b8-de20a536cddf_2400x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!XPX-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ae88e98-2dcc-436b-83b8-de20a536cddf_2400x1600.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>I was the manager, not the developer</h2><p>I didn&#8217;t open my IDE. I wasn&#8217;t the developer. I was the manager. Many agents ran in parallel on the panels, the integrations, the frontend, and the CI and docs, and my job was to brief, review and merge.</p><p>Your throughput is not your keyboard. It is your briefs. What blocks a single developer is typing, and what blocks a single agent is you waiting for it to finish. Running many at once removes both, because they come back one after another and there is always something to review. You don&#8217;t type faster this way. You ship what used to take months.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong><span> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</span></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>A day in the loop</h2><p>The rhythm that produced 20 merged PRs a day is a simple daily loop, and the day is the engine. I am hands-on all day. I spec, launch, review, merge and re-task, live, and most of the day&#8217;s PRs land right there.</p><p>The evening is a hand-off. Before I step away I queue a few deep, long-running plans. The night is a bonus, not the engine. A handful of deep autonomous runs finish by morning, which is the minority of the work. Repeat that for about 11 days and you reach a tagged 1.0.0. Six ingredients make the loop work.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KsfK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F599b15ce-7a3c-42bc-90e7-3b080bfafdf1_2400x1600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KsfK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F599b15ce-7a3c-42bc-90e7-3b080bfafdf1_2400x1600.png 424w, https://substackcdn.com/image/fetch/$s_!KsfK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F599b15ce-7a3c-42bc-90e7-3b080bfafdf1_2400x1600.png 848w, https://substackcdn.com/image/fetch/$s_!KsfK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F599b15ce-7a3c-42bc-90e7-3b080bfafdf1_2400x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!KsfK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F599b15ce-7a3c-42bc-90e7-3b080bfafdf1_2400x1600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KsfK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F599b15ce-7a3c-42bc-90e7-3b080bfafdf1_2400x1600.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/599b15ce-7a3c-42bc-90e7-3b080bfafdf1_2400x1600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:224984,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/207293305?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F599b15ce-7a3c-42bc-90e7-3b080bfafdf1_2400x1600.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KsfK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F599b15ce-7a3c-42bc-90e7-3b080bfafdf1_2400x1600.png 424w, https://substackcdn.com/image/fetch/$s_!KsfK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F599b15ce-7a3c-42bc-90e7-3b080bfafdf1_2400x1600.png 848w, https://substackcdn.com/image/fetch/$s_!KsfK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F599b15ce-7a3c-42bc-90e7-3b080bfafdf1_2400x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!KsfK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F599b15ce-7a3c-42bc-90e7-3b080bfafdf1_2400x1600.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Ingredient one, write the specifications</h2><p>The first thing I write is never code. It is the house rules. An <code>AGENTS.md</code>, which GitHub Copilot reads as <code>copilot-instructions.md</code>, and every agent reads it first. It sets the conventions once, the stack, the build, the tests and the style, with per-panel conventions so 40 panels come out consistent. Mine runs around 700 lines. That is about the right size, because a file that is too big pollutes the context, and it also costs tokens and money on every run.</p><p>The high-leverage move is writing the known failures into that file once. Most models were trained on Spring Boot 3, so they struggle with the reworked Jackson API in Spring Boot 4, and 10 agents trained on the same data hit the same wall at the same time. Rather than fix that by hand in each one, I write the fix down once. The same goes for operational traps, like giving every agent its own local Maven repository so parallel installs do not corrupt a shared one.</p><p>On top of the house rules I write one spec per task. What to build, where, what &#8220;done&#8221; looks like, and the acceptance test the agent must make pass. Small, self-contained, no hidden dependencies. The spec is the product now. The better the brief, the less you babysit.</p><h2>Ingredient two, build the test harness</h2><p>No agent runs without a test harness. One command compiles and runs the unit and end-to-end tests, and comes back green or red, so an agent can verify its own work before it opens a PR. No tests means you cannot trust the output.</p><p>CI is the trust layer. BootUI has about 116 test suites, 81 in Java and 35 in Playwright, and CodeQL plus the end-to-end tests gate every PR to main. The loop does the work. An agent writes code, pushes it, and GitHub Actions sends back the results. If they are red the agent reads the failure, fixes it, commits and pushes again, and the checks run once more. Dependabot keeps the dependencies current alongside all of this. A compiled, typed language like Java helps too, because a hallucinated call often fails to compile before a test ever runs, and the tests catch what the compiler misses. You cannot read every line of 223 PRs. A green build you trust is what makes the volume reviewable.</p><h2>Ingredient three, split the work and run agents in parallel</h2><p>Parallelism is the whole point, so the work has to be parallel-ready. The 40 near-identical panels are perfect to fan out, because AI is strong at cloning something and adapting it to something similar. I give one task to one agent, small scope and a clear goal, each on its own branch or worktree so they do not collide. Splitting the work cleanly is an architecture problem before it is a prompting one, and getting it wrong is where the merge conflicts come from.</p><p>A fleet is not one chat window. I drove BootUI mostly through the <a href="https://github.com/features/ai/github-app">GitHub Copilot coding agent app</a>, because it lets many agents run live on one machine, and I comfortably keep 10 going at once. More than 10 gets complicated for a human to hold in their head, and the laptop starts to slow down. The mobile app runs agents in Docker containers, so I can keep them moving when I am away from my desk. The CLI and the IDE plugin suit smaller fan-out, and all three share the same back end, the Copilot SDK, which is open source, so you can build your own interface on top of it if you want. The rule stays simple. Don&#8217;t babysit one agent. Run 10.</p><h2>Ingredient four, let a few deep plans run overnight</h2><p>The day is the engine and the night is a bonus shift. Before I log off I hand a few deep, long-running plans to autonomous agents, the big jobs I don&#8217;t want to sit and watch. On BootUI those were things like wiring and testing the security filter chains across 37 rules, pushing coverage further across the suites, and reshaping the Actuator data layer in a larger refactor. Each is a long run that lands a few PRs by morning.</p><p>For the genuinely tricky work I have the agent criticize its own output with two or three other models. One writes the code, the others analyze and vote on the fix. That is also how I catch hallucinations, because the model that invented something usually stands alone against the others, and it shows up plainly in the logs. A good overnight prompt means everything is done when I come back. A weak one means I have nothing, which is why prompt-writing is worth practicing. The night is a bonus on top of the day, not a replacement for the driving.</p><h2>Ingredient five, merge the results over coffee</h2><p>The morning starts over coffee. I triage the few overnight PRs, merge the green ones fast, drop or re-task whatever didn&#8217;t land, cherry-pick the good parts of the rest, and then start driving the day&#8217;s fleet. When I review a PR I look at three things that stay in sync because they are generated together, the code, its tests, and its documentation. Green tests tell me the code and the tests agree, and I squash and merge.</p><p>Review is the real bottleneck. It is not the typing anymore. It is the merging. I don&#8217;t merge blind and I don&#8217;t merge like crazy, because the day you wave through a large diff without reading it is the day something wrong lands in main. So make review a fast, trusted ritual you run all day, not a line-by-line slog at the end.</p><h2>Ingredient six, pick the right model for the task</h2><p>Most of the work went to a workhorse, GPT-5.5 on extra-high reasoning, with around 90 percent of its tokens served from cache. For the tricky parts I brought in Claude Opus 4.8 and Gemini 3.1 Pro, three strong models cross-checking each other to find the best fix. For simple, mechanical tasks a smaller model or Auto mode is fast and cheap, and Auto tends to pick better than I would while carrying its own token rebate.</p><p>The economics are driven by the cache more than by the model price. The whole build ran on about 2.7 billion tokens for roughly 2,000 dollars, with around 95 percent served from cache, so most of the run costs a tenth of the headline price. That changes the obvious advice. Swapping to a cheaper model for one small task often bursts the cache and costs more than staying on the model whose context is already warm. If I have a small task, I spawn a separate small agent for it rather than switching the big agent&#8217;s model. And I avoid editing <code>AGENTS.md</code> mid-run, because it lives at the top of the cache and one change invalidates it, so I update it at the end of a run instead. Knowing how the cache works matters more than picking the cheapest model.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UGCm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d32df8-a5dc-4241-8fe7-69b4c42e823a_2400x1600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UGCm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d32df8-a5dc-4241-8fe7-69b4c42e823a_2400x1600.png 424w, https://substackcdn.com/image/fetch/$s_!UGCm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d32df8-a5dc-4241-8fe7-69b4c42e823a_2400x1600.png 848w, https://substackcdn.com/image/fetch/$s_!UGCm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d32df8-a5dc-4241-8fe7-69b4c42e823a_2400x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!UGCm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d32df8-a5dc-4241-8fe7-69b4c42e823a_2400x1600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UGCm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d32df8-a5dc-4241-8fe7-69b4c42e823a_2400x1600.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/03d32df8-a5dc-4241-8fe7-69b4c42e823a_2400x1600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:192051,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/207293305?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d32df8-a5dc-4241-8fe7-69b4c42e823a_2400x1600.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UGCm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d32df8-a5dc-4241-8fe7-69b4c42e823a_2400x1600.png 424w, https://substackcdn.com/image/fetch/$s_!UGCm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d32df8-a5dc-4241-8fe7-69b4c42e823a_2400x1600.png 848w, https://substackcdn.com/image/fetch/$s_!UGCm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d32df8-a5dc-4241-8fe7-69b4c42e823a_2400x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!UGCm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F03d32df8-a5dc-4241-8fe7-69b4c42e823a_2400x1600.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why the multiplier was so large</h2><p>I want to be honest about why BootUI hit a multiplier this large, because not every project will. Three things made this codebase unusually well-suited to agents. First, massive repetition, since around 40 structurally similar panels are cheap for an agent to clone and are the most expensive part by hand. Second, a broad-but-shallow shape, many Spring subsystems each shallow on its own, where the human cost is mostly looking things up, which is exactly what the model already absorbed in training. Third, strong guardrails, because multi-module CI, CodeQL, end-to-end tests and explicit instructions let me accept high throughput safely.</p><p>The net effect was a 6.5 to 8.5 month solo effort compressed into 11 days and about two weeks of human attention. The biggest leverage is on large-surface, pattern-heavy, well-tested code, not on hard algorithms. A project that is mostly novel logic with thin tests will not see these numbers, and I would not claim otherwise.</p><h2>Where this goes wrong</h2><p>The failure modes are consistent, and most of them trace back to the operator. Scope creep is the first. Tell an agent to build the whole thing and it wanders, so keep one tight goal per task. Missing tests are the second. Without a harness you cannot trust the output and you cannot merge at volume, so the harness comes first and a green build gates every merge.</p><p>Giant PRs are the third. A 2,000-line PR is impossible to review well, and review fatigue tempts you to wave it through, so keep the chunks small and reviewable and pace yourself. The wrong model is the fourth. A weak model fails the hard tasks, a strong one is slow and costly on the easy ones, so match the model to the job. When people tell me AI-generated code is not good, the cause is usually one of these. Thin specs, weak tests, oversized PRs, or the wrong model. If the agent generates bad code, that is usually your fault, not the model&#8217;s. Most failed runs are a briefing problem.</p><h2>One page to screenshot</h2><p>If you take one thing away, take this.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tJUl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb20db2fb-0016-4b64-99b8-2e8de8a1fa3b_2400x1600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tJUl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb20db2fb-0016-4b64-99b8-2e8de8a1fa3b_2400x1600.png 424w, https://substackcdn.com/image/fetch/$s_!tJUl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb20db2fb-0016-4b64-99b8-2e8de8a1fa3b_2400x1600.png 848w, https://substackcdn.com/image/fetch/$s_!tJUl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb20db2fb-0016-4b64-99b8-2e8de8a1fa3b_2400x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!tJUl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb20db2fb-0016-4b64-99b8-2e8de8a1fa3b_2400x1600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tJUl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb20db2fb-0016-4b64-99b8-2e8de8a1fa3b_2400x1600.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b20db2fb-0016-4b64-99b8-2e8de8a1fa3b_2400x1600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:327967,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/207293305?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb20db2fb-0016-4b64-99b8-2e8de8a1fa3b_2400x1600.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tJUl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb20db2fb-0016-4b64-99b8-2e8de8a1fa3b_2400x1600.png 424w, https://substackcdn.com/image/fetch/$s_!tJUl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb20db2fb-0016-4b64-99b8-2e8de8a1fa3b_2400x1600.png 848w, https://substackcdn.com/image/fetch/$s_!tJUl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb20db2fb-0016-4b64-99b8-2e8de8a1fa3b_2400x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!tJUl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb20db2fb-0016-4b64-99b8-2e8de8a1fa3b_2400x1600.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>A new Spring Boot console fell out of it</h2><p>The recipe was the point, but it left behind a real, open-source product. <a href="https://github.com/jdubois/boot-ui">BootUI</a> gives any Spring Boot app an embedded console with live health and metrics, a security advisor that walks your filter chains, an OSV scan of your dependencies, a Flyway and Liquibase data view with a Hibernate advisor, tracing through Micrometer and OTLP, and an architecture view backed by ArchUnit and GraalVM reachability. Add one starter to your <code>pom.xml</code>, run locally, and open the console. The code and the <a href="https://julien-dubois.com/boot-ui">full docs</a> are open.</p><p>Write the specs, give them tests, run them wide, and ship what used to take months. Now go build.</p><div><hr></div><p><em>This deep dive is adapted from <strong>Julien&#8217;s</strong> Deep Engineering workshop, From Coder to Manager of Agents. Here the <a href="https://www.julien-dubois.com/conferences/building-with-ai-agents/index-en.html">slides from the talk</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Deep Engineering #55: Julien Dubois on Managing a Fleet of Agents to Ship in Days]]></title><description><![CDATA[On building a production Spring Boot console in 11 days, plus why an 80% demo is not a production system.]]></description><link>https://deepengineering.net/p/issue-55-julien-dubois-fleet-of-agents</link><guid isPermaLink="false">https://deepengineering.net/p/issue-55-julien-dubois-fleet-of-agents</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 16 Jul 2026 13:45:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c80f0ecc-e86d-4147-a3c5-4ce57ef62e78_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><a href="https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt">Capacitor - Shared memory for your team&#8217;s coding agents. </a></strong><a href="https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt">Searchable. Shareable. Vendor-neutral. Scored.</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ocpb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 424w, https://substackcdn.com/image/fetch/$s_!Ocpb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 848w, https://substackcdn.com/image/fetch/$s_!Ocpb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 1272w, https://substackcdn.com/image/fetch/$s_!Ocpb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ocpb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png" width="1360" height="660" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:660,&quot;width&quot;:1360,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!Ocpb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 424w, https://substackcdn.com/image/fetch/$s_!Ocpb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 848w, https://substackcdn.com/image/fetch/$s_!Ocpb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 1272w, https://substackcdn.com/image/fetch/$s_!Ocpb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe15ef0cb-aea9-4ed6-9257-72ae049d9076_1360x660.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt">Capacitor</a></strong> records the session behind the work: what agents tried, what teammates rejected, what finally passed, and why it mattered. This gives you vendor-neutrality to move across multiple coding agents, multiplayer - collaboration on coding sessions, faster PR reviews, evals on code &amp; more.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt&quot;,&quot;text&quot;:&quot;Sign up for free&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.vpdae.com/redirect/t7ef1ieak6aukarvnzev5kstydt"><span>Sign up for free</span></a></p><div><hr></div><p><span>&#9997;&#65039; </span><strong><span>From the editor&#8217;s desk,</span></strong></p><p><span>Welcome to the </span><strong><span>55th</span></strong><span> issue of </span><strong><span>Deep Engineering</span></strong><span>!</span></p><p><span>OpenAI made its</span><a href="https://openai.com/index/gpt-5-6/"><span> GPT-5.6 family generally available</span></a><span> on July 9 across ChatGPT, Codex, and the API. The capability that matters here is a new ultra mode that runs concurrent subagents and synthesizes their work in a single request. This release also makes cached context far cheaper than fresh input, since cache reads keep the 90% cached-input discount. So running a fleet of agents in parallel, and paying almost nothing to reuse their context, has moved from a hand-built trick into something the platform now does by default.</span></p><p><span>That shift is why this issue matters now. When the tooling makes parallel agents this easy, the differentiator stops being model access and becomes the discipline around it, the specifications, the tests, and the review that decide whether all that speed produces software you can actually ship. Last week</span><a href="https://www.julien-dubois.com"><span> Julien Dubois</span></a><span> led a workshop for us, </span><strong><span>From Coder to Manager of Agents</span></strong><span>, where he showed how he built a real open-source product this way, measured from git history rather than memory, which makes it far more useful than another benchmark thread.</span></p><p><strong><span>Julien</span></strong><span> is Principal Manager for Developer Relations at </span><strong><span>Microsoft</span></strong><span> and </span><strong><span>GitHub</span></strong><span>, and the creator of </span><strong><a href="https://www.jhipster.tech/"><span>JHipster</span></a></strong><span>.</span></p><p><span>Let&#8217;s get started.</span></p><div><hr></div><p><strong>Featured: <a href="https://www.eventbrite.com/e/1992373400474/?discount=DEEPENG40">Loop Engineering for AI Agents</a></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.eventbrite.com/e/1992373400474/?discount=DEEPENG40" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BDF7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcda255c8-7f56-4123-8294-dddb183cc933_800x267.jpeg 424w, https://substackcdn.com/image/fetch/$s_!BDF7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcda255c8-7f56-4123-8294-dddb183cc933_800x267.jpeg 848w, https://substackcdn.com/image/fetch/$s_!BDF7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcda255c8-7f56-4123-8294-dddb183cc933_800x267.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!BDF7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcda255c8-7f56-4123-8294-dddb183cc933_800x267.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BDF7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcda255c8-7f56-4123-8294-dddb183cc933_800x267.jpeg" width="800" height="267" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cda255c8-7f56-4123-8294-dddb183cc933_800x267.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:267,&quot;width&quot;:800,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.eventbrite.com/e/1992373400474/?discount=DEEPENG40&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!BDF7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcda255c8-7f56-4123-8294-dddb183cc933_800x267.jpeg 424w, https://substackcdn.com/image/fetch/$s_!BDF7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcda255c8-7f56-4123-8294-dddb183cc933_800x267.jpeg 848w, https://substackcdn.com/image/fetch/$s_!BDF7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcda255c8-7f56-4123-8294-dddb183cc933_800x267.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!BDF7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcda255c8-7f56-4123-8294-dddb183cc933_800x267.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Stop wasting tokens on endless retries. </strong>This four-hour, hands-on workshop takes you past one-off prompting to reliable agent loops that plan, execute, verify, and stop safely, built with Claude Code, Spec-Driven Development, and MCP, with the verification gates and state management that keep autonomous agents trustworthy. </p><p style="text-align: center;">Deep Engineering readers save 40% with code <strong>DEEPENG40</strong>. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.co.uk/e/loop-engineering-for-ai-agents-tickets-1992373400474?discount=DEEPENG40&quot;,&quot;text&quot;:&quot;Reserve your seat&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.eventbrite.co.uk/e/loop-engineering-for-ai-agents-tickets-1992373400474?discount=DEEPENG40"><span>Reserve your seat</span></a></p><div><hr></div><p><strong>Practitioner's View</strong> by <a href="https://www.julien-dubois.com">Julien Dubois</a></p><h2>223 pull requests in 11 days for the project I never had time to build</h2><blockquote><p>You can read this <a href="https://deepengineering.net/p/223-pull-requests-in-11-days-julien-dubois">practical deep dive</a> based on his workshop. And here are the <a href="https://www.julien-dubois.com/conferences/building-with-ai-agents/index-en.html">slides from the talk</a>.</p></blockquote><p><span>I always wanted a real developer UI for Spring Boot. Every Spring app is a black box in development. Actuator gives you raw JSON, not a console, so what I wanted was health, metrics, security and tracing in one embedded UI.</span></p><p><span>I&#8217;d shipped a slice of this in JHipster years ago, but only for generated apps. A real console for any Spring Boot app sat on my wishlist for years. The catch is that it&#8217;s massive. Around 40 panels, each one a backend plus a frontend plus its own tests, deeply integrated across a dozen JVM subsystems. By hand, one experienced developer would need somewhere between six and a half and eight and a half months. Too big to justify, until I stopped writing the code myself.</span></p><p><span>Then I did it in 11 days. Not by typing faster. I didn&#8217;t even open my IDE. I managed a fleet of AI agents while I architected, reviewed and steered. The agents did the scaffolding, the panels and the tests. I did the judgement.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bqyI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe52cc3c7-39b7-4a53-aeab-f052ea83c92d_1456x971.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bqyI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe52cc3c7-39b7-4a53-aeab-f052ea83c92d_1456x971.jpeg 424w, https://substackcdn.com/image/fetch/$s_!bqyI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe52cc3c7-39b7-4a53-aeab-f052ea83c92d_1456x971.jpeg 848w, https://substackcdn.com/image/fetch/$s_!bqyI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe52cc3c7-39b7-4a53-aeab-f052ea83c92d_1456x971.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!bqyI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe52cc3c7-39b7-4a53-aeab-f052ea83c92d_1456x971.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bqyI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe52cc3c7-39b7-4a53-aeab-f052ea83c92d_1456x971.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e52cc3c7-39b7-4a53-aeab-f052ea83c92d_1456x971.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bqyI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe52cc3c7-39b7-4a53-aeab-f052ea83c92d_1456x971.jpeg 424w, https://substackcdn.com/image/fetch/$s_!bqyI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe52cc3c7-39b7-4a53-aeab-f052ea83c92d_1456x971.jpeg 848w, https://substackcdn.com/image/fetch/$s_!bqyI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe52cc3c7-39b7-4a53-aeab-f052ea83c92d_1456x971.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!bqyI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe52cc3c7-39b7-4a53-aeab-f052ea83c92d_1456x971.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><span>What I built, BootUI</span></h3><p><a href="https://github.com/jdubois/boot-ui"><span>BootUI</span></a><span> is a production-grade Spring Boot 4 starter that adds an embedded, local-only developer console to your app. You add one starter to your pom.xml, run locally, and open the console.</span></p><p><span>It is a five-module Maven build on Spring Boot 4 and Java 17, integrated across Actuator, Spring Security, Flyway and Liquibase, Hibernate, Micrometer with OTLP, GraalVM, OSV scanning and ArchUnit, published to Maven Central with full CI. The frontend is an embedded Vue 3 single-page app, packaged inside the starter. Each panel is endpoints plus a view plus tests.</span></p><p><span>There are around 40 feature panels. Health and metrics, a security advisor, a vulnerabilities view, tracing, and more. Those panels look a lot alike, which made them the biggest source of structural repetition in the codebase. That detail matters, and I will come back to why it made the whole thing work.</span></p><h3><span>Eleven days, and how I know the numbers are real</span></h3><p><span>The figures here are derived from git history, PR metadata and code metrics, not from time-tracking logs. Through the tagged 1.0.0 release there were around 264 commits on main and about 223 squash-merged pull requests, landing at roughly 20 a day across 11 days. That came to about 83k total tracked source lines, close to 50k of them Java across around 461 files and 81 test classes, plus 52 Vue components, 40 panels, and 35 end-to-end specs. In all, about 116 test suites.</span></p><p><span>A cadence of 20 PRs a day with an AI agent co-authoring commits is impossible to reach by hand. The agent did the typing. I dispatched many asynchronous tasks in parallel. The evidence lines up with that. The commit clock runs from around five in the morning to midnight most days, which is consistent with parallel async tasks rather than continuous typing. &#8220;Copilot&#8221; shows up as a named commit and PR author on about 44 commits, and the repo ships a copilot-instructions.md with per-panel conventions, so the workflow was explicitly agent-oriented. Even on release day I was landing a docs site, a scanner dashboard, token charts and a Hikari fix, which is itself several days of solo work.</span></p><p><span>Where my own time went was writing prompts, reviewing and merging those 223 PRs, and resolving CI failures around Spring Boot 4, Flyway 11 and OTLP. Review and orchestrate, not write every line.</span></p><p><span>The honest by-hand comparison is the part people argue about, so I ran it carefully. One experienced Spring Boot and Vue developer, with no AI codegen, would need 6.5 to 8.5 months for the same polished 1.0.0, roughly 1,100 to 1,450 hours of hands-on effort, and the 40 panels are the dominant cost. A COCOMO organic estimate on 50k lines yields more than 100 person-months, which is too high because much of the code is repetitive scaffolding, so a domain-expert solo figure of about 7.5 months is the defensible middle ground. My actual human effort on BootUI was 80 to 110 hours, about two intense solo weeks. That is a calendar speed-up of roughly 17 to 23 times, and a human-hours speed-up of roughly 12 to 17 times.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9UeI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe032403-2b0e-420f-a041-362326fd9677_1456x971.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9UeI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe032403-2b0e-420f-a041-362326fd9677_1456x971.jpeg 424w, https://substackcdn.com/image/fetch/$s_!9UeI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe032403-2b0e-420f-a041-362326fd9677_1456x971.jpeg 848w, https://substackcdn.com/image/fetch/$s_!9UeI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe032403-2b0e-420f-a041-362326fd9677_1456x971.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!9UeI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe032403-2b0e-420f-a041-362326fd9677_1456x971.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9UeI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe032403-2b0e-420f-a041-362326fd9677_1456x971.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fe032403-2b0e-420f-a041-362326fd9677_1456x971.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9UeI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe032403-2b0e-420f-a041-362326fd9677_1456x971.jpeg 424w, https://substackcdn.com/image/fetch/$s_!9UeI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe032403-2b0e-420f-a041-362326fd9677_1456x971.jpeg 848w, https://substackcdn.com/image/fetch/$s_!9UeI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe032403-2b0e-420f-a041-362326fd9677_1456x971.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!9UeI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe032403-2b0e-420f-a041-362326fd9677_1456x971.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><span>I was the manager, not the developer</span></h3><p><span>I didn&#8217;t open my IDE. I wasn&#8217;t the developer. I was the manager. Many agents ran in parallel on the panels, the integrations, the frontend, and the CI and docs, and my job was to brief, review and merge.</span></p><p><span>Your throughput is not your keyboard. It is your briefs. What blocks a single developer is typing, and what blocks a single agent is you waiting for it to finish. Running many at once removes both, because they come back one after another and there is always something to review. You don&#8217;t type faster this way. You ship what used to take months.</span></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3><span>A day in the loop</span></h3><p><span>The rhythm that produced 20 merged PRs a day is a simple daily loop, and the day is the engine. I am hands-on all day. I spec, launch, review, merge and re-task, live, and most of the day&#8217;s PRs land right there.</span></p><p><span>The evening is a hand-off. Before I step away I queue a few deep, long-running plans. The night is a bonus, not the engine. A handful of deep autonomous runs finish by morning, which is the minority of the work. Repeat that for about 11 days and you reach a tagged 1.0.0. Six ingredients make the loop work.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DdaV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d887a9-10ae-46fb-9730-1eeff37e9fc2_1456x971.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DdaV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d887a9-10ae-46fb-9730-1eeff37e9fc2_1456x971.jpeg 424w, https://substackcdn.com/image/fetch/$s_!DdaV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d887a9-10ae-46fb-9730-1eeff37e9fc2_1456x971.jpeg 848w, https://substackcdn.com/image/fetch/$s_!DdaV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d887a9-10ae-46fb-9730-1eeff37e9fc2_1456x971.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!DdaV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d887a9-10ae-46fb-9730-1eeff37e9fc2_1456x971.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DdaV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d887a9-10ae-46fb-9730-1eeff37e9fc2_1456x971.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/38d887a9-10ae-46fb-9730-1eeff37e9fc2_1456x971.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DdaV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d887a9-10ae-46fb-9730-1eeff37e9fc2_1456x971.jpeg 424w, https://substackcdn.com/image/fetch/$s_!DdaV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d887a9-10ae-46fb-9730-1eeff37e9fc2_1456x971.jpeg 848w, https://substackcdn.com/image/fetch/$s_!DdaV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d887a9-10ae-46fb-9730-1eeff37e9fc2_1456x971.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!DdaV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d887a9-10ae-46fb-9730-1eeff37e9fc2_1456x971.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><span>Ingredient one, write the specifications</span></h4><p><span>The first thing I write is never code. It is the house rules. An AGENTS.md, which GitHub Copilot reads as copilot-instructions.md, and every agent reads it first. It sets the conventions once, the stack, the build, the tests and the style, with per-panel conventions so 40 panels come out consistent. Mine runs around 700 lines. That is about the right size, because a file that is too big pollutes the context, and it also costs tokens and money on every run.</span></p><p><span>The high-leverage move is writing the known failures into that file once. Most models were trained on Spring Boot 3, so they struggle with the reworked Jackson API in Spring Boot 4, and 10 agents trained on the same data hit the same wall at the same time. Rather than fix that by hand in each one, I write the fix down once. The same goes for operational traps, like giving every agent its own local Maven repository so parallel installs do not corrupt a shared one.</span></p><p><span>On top of the house rules I write one spec per task. What to build, where, what &#8220;done&#8221; looks like, and the acceptance test the agent must make pass. Small, self-contained, no hidden dependencies. The spec is the product now. The better the brief, the less you babysit.</span></p><h4><span>Ingredient two, build the test harness</span></h4><p><span>No agent runs without a test harness. One command compiles and runs the unit and end-to-end tests, and comes back green or red, so an agent can verify its own work before it opens a PR. No tests means you cannot trust the output.</span></p><p><span>CI is the trust layer. BootUI has about 116 test suites, 81 in Java and 35 in Playwright, and CodeQL plus the end-to-end tests gate every PR to main. The loop does the work. An agent writes code, pushes it, and GitHub Actions sends back the results. If they are red the agent reads the failure, fixes it, commits and pushes again, and the checks run once more. Dependabot keeps the dependencies current alongside all of this. A compiled, typed language like Java helps too, because a hallucinated call often fails to compile before a test ever runs, and the tests catch what the compiler misses. You cannot read every line of 223 PRs. A green build you trust is what makes the volume reviewable.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/i/207293305/ingredient-two-build-the-test-harness&quot;,&quot;text&quot;:&quot;Continue reading&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://deepengineering.net/i/207293305/ingredient-two-build-the-test-harness"><span>Continue reading</span></a></p><blockquote><p><span>The </span><a href="https://deepengineering.net/p/223-pull-requests-in-11-days-julien-dubois"><span>rest of the deep dive</span></a><span> covers the other four ingredients, the test harness and CI as the trust layer, running agents in parallel across worktrees, the overnight critic-and-vote runs, merging as the real bottleneck, and the cache math behind the roughly two thousand dollar bill, along with where the multiplier will not repeat.</span></p></blockquote><div><hr></div><p><strong><span>Industry Perspective</span></strong></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;44242028-74a5-4e04-8ca5-1f3d6f9b395f&quot;,&quot;caption&quot;:&quot;Julien Dubois shows how a fleet of agents lets one engineer move at the pace of a team. Imran Ahmad takes up the question that speed raises, whether the software those agents produce can be trusted in production. He calls the distance between an eighty percent demo and a 99.9 percent production system the Reliability Gap, and he argues you close it with a deterministic shell around the model rather than with a better model. His piece walks the failures that make the case, from an invented refund to a deleted production database, and gives the four questions he now asks in every design review.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;So Your Demo Is Lying to You&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:2260974,&quot;name&quot;:&quot;Imran Ahmad&quot;,&quot;bio&quot;:&quot;I&#8217;m a data scientist and an author&quot;,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!qXS5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b8b0819-6513-45db-b99b-ec849c945bf1_144x144.png&quot;,&quot;is_guest&quot;:true,&quot;bestseller_tier&quot;:null,&quot;primaryPublicationSubscribeUrl&quot;:&quot;https://imran409.substack.com/subscribe?&quot;,&quot;primaryPublicationUrl&quot;:&quot;https://imran409.substack.com&quot;,&quot;primaryPublicationName&quot;:&quot;Imran Ahmad&quot;,&quot;primaryPublicationId&quot;:5449073}],&quot;post_date&quot;:&quot;2026-07-16T15:39:12.766Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/57c40ede-c31d-4327-9d91-4c7ee4fb93f6_2400x1600.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.net/p/your-demo-is-lying-to-you-imran-ahmad&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:207301930,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div class="callout-block" data-callout="true"><p><strong>&#128227; Contribute to Deep Engineering </strong></p><p>If you are a senior engineer with a hard-won lesson or a failure worth sharing, <strong>pitch</strong> a <a href="https://deepengineering.net/s/practical-deep-dives">practical deep dive</a> under your <strong>byline</strong>. Or if you lead a team, we would like to <strong>interview</strong> you and build an <a href="https://deepengineering.net/s/engineering-leadership">engineering leadership</a> feature around your <strong>story</strong>.<br><br><strong>Subscribe</strong> to <strong>Deep Engineering</strong> newsletter and <strong>message</strong> us through the <strong>chat option</strong>, or email us at <strong>saqibj @ packt.com</strong>.</p></div><h2>&#128736;&#65039; Tool of the Week</h2><p><strong><a href="http://github.com/ComposioHQ/agent-orchestrator"><span>Composio&#8217;s Agent Orchestrator</span></a></strong><span> is an open-source take on exactly that job. It helps developers manage fleets of coding agents for parallel work, giving each one an isolated workspace and supervising them from one place.</span></p><ul><li><p><span>Runs each agent in its own git worktree, so parallel work does not clobber shared files</span></p></li><li><p><span>Feeds CI failures, review comments, and merge conflicts back to the right agent automatically</span></p></li><li><p><span>Works with the terminal agents teams already use, including Claude Code, Codex, Cursor, and opencode</span></p></li><li><p><span>MIT licensed and self-hosted, run locally rather than as a hosted service</span></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;http://github.com/ComposioHQ/agent-orchestrator&quot;,&quot;text&quot;:&quot;Composio&#8217;s Agent Orchestrator&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="http://github.com/ComposioHQ/agent-orchestrator"><span>Composio&#8217;s Agent Orchestrator</span></a></p><div><hr></div><h2>&#128206; Tech Briefs</h2><ul><li><p><a href="https://code.claude.com/docs/en/changelog"><span>Claude Code changelog</span></a><span> - Agents must now confirm before entering a git worktree outside the project&#8217;s .claude/worktrees directory, and a new /doctor check flags checked-in CLAUDE.md content the model can derive from the codebase.</span></p></li><li><p><a href="https://github.com/github/copilot-cli/blob/main/changelog.md"><span>GitHub Copilot CLI changelog</span></a><span> - Plan mode can no longer run tools that modify the workspace, and the default maximum sub-agent nesting depth drops from six to four to limit runaway recursive delegation.</span></p></li><li><p><a href="https://github.blog/changelog/2026-07-14-github-copilot-for-jetbrains-expands-byok-capabilities/"><span>Copilot for JetBrains expands BYOK capabilities</span></a><span> - Adds local agent sandboxing, a Claude agent provider for custom agents and skills, and OpenAI-compatible custom endpoints, all in public preview.</span></p></li><li><p><a href="https://github.blog/changelog/2026-07-14-github-copilot-in-visual-studio-june-update/"><span>Copilot in Visual Studio, June update</span></a><span> - Visual Studio now checks each MCP server&#8217;s configuration and asset fingerprint against a trusted baseline at startup, and the C++ modernization agent reached general availability.</span></p></li><li><p><a href="https://releasebot.io/updates/anthropic/claude">Claude Code gateway for teams</a> - Anthropic shipped a self-hosted gateway for Claude Code, a stateless container that adds corporate SSO login, centrally enforced policy, role-based access, and per-user cost attribution across a team.</p></li></ul><div><hr></div><p><span>That&#8217;s all for today. Thank you for reading this issue of Deep Engineering.</span></p><p><span>We&#8217;ll be back next week with more expert-led content.</span></p><p><span>Keep building,</span></p><p><span>Saqib Jan</span></p><p><span>Editor-in-Chief, Deep Engineering</span></p><div><hr></div><p><em><span>If your company wants to reach senior developers, software engineers, and technical decision-makers, </span><a href="https://packt.omeclk.com/portal/wts/uc%5EcnN2dfNaqmD-kB-mo66%7C7g%5Ef%7Cb"><span>speak to us about partnering</span></a><span> with Deep Engineering.</span></em></p>]]></content:encoded></item><item><title><![CDATA[Deep Engineering Specials: Judgment, Not Tokenmaxxing, Creates the Value]]></title><description><![CDATA[The next limit on engineering output is not how many tokens your teams burn. It is whether anyone can still tell motion apart from judgment.]]></description><link>https://deepengineering.net/p/special-issue-judgment-not-tokenmaxxing-creates-value</link><guid isPermaLink="false">https://deepengineering.net/p/special-issue-judgment-not-tokenmaxxing-creates-value</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Tue, 14 Jul 2026 21:52:54 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/9fc38405-fb9a-4c2b-9966-d183b4d9694b_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The next enterprise AI bottleneck is not model capability. It is whether leaders can tell the difference between a team using AI well and a team running up a number.</p><p>Over roughly a quarter in 2026, several of the largest engineering organizations in the world took down the dashboards they had built to prove their AI investments were working. Amazon shut down <a href="https://aimagazine.com/news/why-amazon-has-dropped-its-internal-ai-usage-leaderboard">KiroRank</a>, the internal leaderboard that ranked developers by how many tokens they consumed on its Kiro platform. <a href="https://www.thestreet.com/technology/amazon-joins-microsoft-in-sending-shocking-message-to-employees">Meta also dismantled a near-identical board</a>, Claudenomics, that tracked token usage among its heaviest AI users. The reason was the same everywhere. Once usage became the number that mattered, engineers found ways to run up usage, and the costs arrived long before the value did.</p><p>Amazon&#8217;s own senior vice president, Dave Treadwell, told staff to stop using AI for its own sake after employees gamed the leaderboard by padding token usage, the practice the industry had taken to calling &#8220;<a href="https://www.forbes.com/sites/timkeary/2026/04/13/is-the-cult-of-tokenmaxxingjust-another-fad-or-the-new-normal/">tokenmaxxing</a>.&#8221; Uber also said it could find no clear link between what it was spending on AI and what it was shipping, while Microsoft cancelled a division&#8217;s coding-assistant licences over cost. So, within a single quarter, the industry ran the experiment, watched it fail, and started hunting for a better question to ask.</p><p>That question is the subject of this issue, with insights from <a href="https://www.linkedin.com/in/austinlparker">Austin Parker</a>, Director of AI Strategy at <strong>Honeycomb</strong> and a co-founder of <strong>OpenTelemetry</strong>; <a href="https://www.linkedin.com/in/tom-howe-a565ab1/">Tom Howe</a>, Director of Solutions Engineering at <strong>Hydrolix</strong>; <a href="https://www.linkedin.com/in/jayeeta-putatunda">Jayeeta Putatunda</a>, Director of the AI Center of Excellence at <strong>Fitch Ratings</strong>; <a href="https://www.linkedin.com/in/satyamdhar">Satyam Dhar</a>, Staff Software Engineer at <strong>Galileo</strong> and a former engineering leader at <strong>Amazon</strong> and <strong>Adobe</strong>; and <a href="https://www.linkedin.com/in/mjjtiffany">Michael J.J. Tiffany</a>, co-founder and CEO of <strong>Fulcra Dynamics</strong>.</p><p>Let&#8217;s get started.</p><div><hr></div><p><strong><a href="https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w">Social engineering is about manipulating people&#8217;s emotions. Identify the susceptibilities that hackers use to exploit people.</a></strong></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1GsN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48c5f841-09dd-49bd-a7bd-62a3a008ac97_300x200.png 424w, https://substackcdn.com/image/fetch/$s_!1GsN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48c5f841-09dd-49bd-a7bd-62a3a008ac97_300x200.png 848w, https://substackcdn.com/image/fetch/$s_!1GsN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48c5f841-09dd-49bd-a7bd-62a3a008ac97_300x200.png 1272w, https://substackcdn.com/image/fetch/$s_!1GsN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48c5f841-09dd-49bd-a7bd-62a3a008ac97_300x200.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1GsN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48c5f841-09dd-49bd-a7bd-62a3a008ac97_300x200.png" width="300" height="200" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/48c5f841-09dd-49bd-a7bd-62a3a008ac97_300x200.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:200,&quot;width&quot;:300,&quot;resizeWidth&quot;:300,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!1GsN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48c5f841-09dd-49bd-a7bd-62a3a008ac97_300x200.png 424w, https://substackcdn.com/image/fetch/$s_!1GsN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48c5f841-09dd-49bd-a7bd-62a3a008ac97_300x200.png 848w, https://substackcdn.com/image/fetch/$s_!1GsN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48c5f841-09dd-49bd-a7bd-62a3a008ac97_300x200.png 1272w, https://substackcdn.com/image/fetch/$s_!1GsN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48c5f841-09dd-49bd-a7bd-62a3a008ac97_300x200.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><span>This </span><strong><a href="https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w">NINJIO Insights Report</a></strong><a href="https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w"> </a><span>dives into the key emotional susceptibilities that make social engineering work and offers concrete steps that your security team can take to equip your workforce to resist cyberattacks.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w&quot;,&quot;text&quot;:&quot;Download the Guide&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.vpdae.com/redirect/nck675s2gzn00ukhfu9tyrmbc3w"><span>Download the Guide</span></a></p><div><hr></div><p></p><p><strong>Special issue &#8212; July 2026</strong></p><h2>A token was never a unit of intelligence</h2><p></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FOUY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051162d2-4f58-43f9-9622-0875d28ab799_1601x1365.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FOUY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051162d2-4f58-43f9-9622-0875d28ab799_1601x1365.jpeg 424w, https://substackcdn.com/image/fetch/$s_!FOUY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051162d2-4f58-43f9-9622-0875d28ab799_1601x1365.jpeg 848w, https://substackcdn.com/image/fetch/$s_!FOUY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051162d2-4f58-43f9-9622-0875d28ab799_1601x1365.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!FOUY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051162d2-4f58-43f9-9622-0875d28ab799_1601x1365.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FOUY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051162d2-4f58-43f9-9622-0875d28ab799_1601x1365.jpeg" width="259" height="220.75480769230768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/051162d2-4f58-43f9-9622-0875d28ab799_1601x1365.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1241,&quot;width&quot;:1456,&quot;resizeWidth&quot;:259,&quot;bytes&quot;:516860,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/207072610?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7c04aab-3cbe-4ece-9037-d5e690a530e5_1601x2400.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!FOUY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051162d2-4f58-43f9-9622-0875d28ab799_1601x1365.jpeg 424w, https://substackcdn.com/image/fetch/$s_!FOUY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051162d2-4f58-43f9-9622-0875d28ab799_1601x1365.jpeg 848w, https://substackcdn.com/image/fetch/$s_!FOUY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051162d2-4f58-43f9-9622-0875d28ab799_1601x1365.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!FOUY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F051162d2-4f58-43f9-9622-0875d28ab799_1601x1365.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p style="text-align: center;"><em>&#8220;You can use tokens to help build solutions, but at the end of the day these are judgment calls.&#8221; &#8212; </em><a href="https://www.linkedin.com/in/austinlparker">Austin Parker</a>, Director of AI Strategy at <a href="https://www.honeycomb.io/">Honeycomb</a> and a co-founder of <a href="https://opentelemetry.io/">OpenTelemetry</a></p><div><hr></div><p>Give a modern coding model the right harness and a clear goal and it will do almost anything you ask, up to and including rewriting an entire codebase in another language and passing the tests. That capability is real, and it is seductive. It makes it easy to believe that all you need is something to keep the tokens flowing and any problem becomes tractable. Parker has observed his own engineering teams test that belief, generating a hundred variations of the same screen or twenty versions of a flow, and the lesson came back the same each time.</p><p>The variations did not move the needle. You can ask for infinite variations of a solution, and they all arrive shaped like your original conception of the problem, because that conception is the one human input the model never questions. The code and the screen and the workflow were never the things that create value for the people using them. Value lives one level up, in whether the design serves those people, whether the shape of an API fits what its consumers actually need, whether the underlying primitives have any fitness for the purpose they are being asked to serve.</p><p>Those are judgment calls, and no volume of tokens answers them. Parker sees the same pattern when he talks to other leaders, plenty of appetite for here is a problem, generate all the variations, and far less time asking whether the problem was framed correctly in the first place. His conclusion is blunt. Judgment is the finite resource now, and teams that forget it end up with more output and less of what the output was supposed to buy them.</p><h2>Leaderboards reward motion, not judgment</h2><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IpNN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64641c0-ed85-4856-924e-219cc39cebac_2000x2000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IpNN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64641c0-ed85-4856-924e-219cc39cebac_2000x2000.png 424w, https://substackcdn.com/image/fetch/$s_!IpNN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64641c0-ed85-4856-924e-219cc39cebac_2000x2000.png 848w, https://substackcdn.com/image/fetch/$s_!IpNN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64641c0-ed85-4856-924e-219cc39cebac_2000x2000.png 1272w, https://substackcdn.com/image/fetch/$s_!IpNN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64641c0-ed85-4856-924e-219cc39cebac_2000x2000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IpNN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64641c0-ed85-4856-924e-219cc39cebac_2000x2000.png" width="249" height="249" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e64641c0-ed85-4856-924e-219cc39cebac_2000x2000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:249,&quot;bytes&quot;:3958785,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/207072610?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64641c0-ed85-4856-924e-219cc39cebac_2000x2000.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!IpNN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64641c0-ed85-4856-924e-219cc39cebac_2000x2000.png 424w, https://substackcdn.com/image/fetch/$s_!IpNN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64641c0-ed85-4856-924e-219cc39cebac_2000x2000.png 848w, https://substackcdn.com/image/fetch/$s_!IpNN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64641c0-ed85-4856-924e-219cc39cebac_2000x2000.png 1272w, https://substackcdn.com/image/fetch/$s_!IpNN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64641c0-ed85-4856-924e-219cc39cebac_2000x2000.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em>&#8220;The paradigm in some organizations will, undoubtedly, shift from &#8216;high token usage equals good employees&#8217; to &#8216;high token usage equals expensive employees.&#8217;&#8221; &#8212; <a href="https://www.linkedin.com/in/tom-howe-a565ab1/">Tom Howe</a>, Director of Solutions Engineering at <a href="https://hydrolix.io/">Hydrolix</a></em></p><div><hr></div><p>The collapse of the token leaderboards surprised no one who has watched a measure become a target. Hand a famously clever group of engineers a resource and grade them on how much of it they use, and they will use it. The moment a measure becomes a metric it stops being a good measure, a version of <a href="https://www.metaintro.com/blog/amazon-ai-usage-targets-inflated">Goodhart&#8217;s law that Amazon&#8217;s own leadership</a> named out loud when it wound the program down. The surprise was not that the leaderboards were gamed. It was that so few people in the decision chain saw it coming.</p><p>Howe has seen the mechanism play out from the inside, and he traces it to a subtle slippage. A company rolls out what it calls an AI usage dashboard, meant as an objective measure of how much the tooling is benefiting the business, and it quietly devolves into a leaderboard that measures how much each person is using the tool. Once employees read it as a judgment of how they work rather than of whether the tool earns its keep, the whole exercise changes character.</p><p>From there the responses split three ways, and each one erodes the thing the dashboard was supposed to protect. Some engineers game it, firing tools at low-value tasks to boost their stats, which breeds suspicion that others are gaming it too. Some avoid the tool entirely rather than expose their habits to uncontextualized data. And in the middle sits the casualty Howe cares about most, the trust that a healthy organization runs on, undermined the moment a team cannot tell how or why it is being measured.</p><p>The cost frame makes it worse. Howe describes enterprises treating tokens as close-to-free productivity, funny money, right up until the real bills and the real gains come into focus. When they do, the incentive can flip hard, from prizing heavy usage to flagging it as expense, and engineers start hedging, wary of being branded token wasters and wary of binding their workflows so tightly to a tool that a future rationing decision leaves them stranded. He points to a recent, unusually capable model that carries a known high future cost, where engineers are trying to use its power now without becoming dependent on it, as exactly the kind of bind these metrics create.</p><p>None of this means the tools do not work. Howe shares that Hydrolix has seen a revolutionary impact from how it uses them, in building products and in weaving them into products, in ways that would have been unimaginable a few years ago. The challenge is balancing productivity, cost, and trust, and learning to measure the impact through experience rather than through a scoreboard. He tells one story that captures the slope precisely, an engineer glancing at the usage board and joking that not cracking the top fifty was rookie numbers. It was a joke. But it was also a warning about how fast an innocuous metric turns into a competition.</p><div><hr></div><p><strong>Featured - <a href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng">ARC 2026: Software Architecture in the Age of AI</a></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!a2PN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 424w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 848w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!a2PN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg" width="728" height="364" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;ARC 2026: Software Architecture in the Age of AI&quot;,&quot;title&quot;:&quot;ARC 2026: Software Architecture in the Age of AI&quot;,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="ARC 2026: Software Architecture in the Age of AI" title="ARC 2026: Software Architecture in the Age of AI" srcset="https://substackcdn.com/image/fetch/$s_!a2PN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 424w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 848w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 1456w" sizes="100vw" loading="lazy" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>AI is reshaping software architecture, putting new demands on scalability, governance, reliability, and observability. </span><a href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng">ARC 2026</a><span> brings together architects, CTOs, and AI practitioners for keynotes, panels, and workshops on agentic system design, modernizing enterprise apps for AI, and building governable, observable AI systems.</span></p><p style="text-align: center;"><span>&#128467;&#65039; </span><em><strong>25</strong><span> to </span><strong>26</strong><span> July, </span><strong>10:30</strong><span> am ET</span></em></p><p style="text-align: center;"><span>Use code </span><strong>DEEPENG50</strong><span> for 50% off the early bird price.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng&quot;,&quot;text&quot;:&quot;Reserve your spot&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng"><span>Reserve your spot</span></a></p><div><hr></div><h2>Skills that matter never show up on a dashboard</h2><p>Parker calls prompting, building internal alignment, and holding a complex system in your head illegible activities, and they are close cousins of the glue work that keeps organizations moving. The people who know who needs to be in a decision chain, who can hear a customer complaint and route it to the one team that can fix it, who understand that the thing frustrating you in one corner of a product is caused by something three systems away, are the reason a business can actually execute. The organizations that are genuinely good at this tend to be the ones that, lacking a clean way to measure those people, are at least careful not to penalize them.</p><p>A token metric makes that harder. It rewards the visible act of consuming AI over the invisible work of knowing what is worth building, and the invisible work is where most of the value hides. That work shows up in promo packets and career ladders because those are written by people who understand it. It does not show up in a number you can compare quarter to quarter, because, in Parker&#8217;s words, it involves &#8220;very human things like relationships and time in the system rather than timing the system.&#8221; Measure the activity and you slowly stop rewarding the thing that made the activity worth anything.</p><h2>AI widens the gap a team already has</h2><p>This is the reframing every leader watching one team pull ahead should steal. AI rarely conjures capability from nothing. It accentuates differences that already exist. When a team surges, the real story is almost never that raw intelligence solved a problem they could not solve before. It is that AI patched a hole in the team&#8217;s composition, or handed them leverage on something organizational rather than technical.</p><p>Parker&#8217;s example is the part of shipping that engineers tend to undervalue, making the case to leadership. That is usually a data-analysis argument, projections and forecasts and a slog through support tickets and sentiment data, and it happens to be something AI is very good at. A motivated engineer or designer or support person can now do the analysis a product manager used to own, cover the weak spots in their own toolkit, and come back to the org with here is what I need in order to go build the thing I am convinced is right. So the question to ask the team that pulled ahead is not which tool they used. In Parker&#8217;s words, &#8220;I never want to ask what AI has let you do. The question is what is AI giving you the excuse to do that you could not have done otherwise.&#8221;</p><p>That does not extend into a ban on measuring code output. Parker&#8217;s line is that you should measure it and refuse to make it a metric. Watching the balance of someone&#8217;s pull requests opened against reviewed tells you something real about how they work, and grading their performance on that same count is where it curdles. AI can even help here, letting a team ask which work is generating on-call burden, which is delivering customer value, and who is quietly doing the refactoring that lifts everything built on top of it.</p><h2>Point the budget inward, at developer experience</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!h2WP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3232fd-537a-4906-934b-abe48d3d715b_263x277.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!h2WP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3232fd-537a-4906-934b-abe48d3d715b_263x277.jpeg 424w, https://substackcdn.com/image/fetch/$s_!h2WP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3232fd-537a-4906-934b-abe48d3d715b_263x277.jpeg 848w, https://substackcdn.com/image/fetch/$s_!h2WP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3232fd-537a-4906-934b-abe48d3d715b_263x277.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!h2WP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3232fd-537a-4906-934b-abe48d3d715b_263x277.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!h2WP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3232fd-537a-4906-934b-abe48d3d715b_263x277.jpeg" width="251" height="264.361216730038" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aa3232fd-537a-4906-934b-abe48d3d715b_263x277.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:277,&quot;width&quot;:263,&quot;resizeWidth&quot;:251,&quot;bytes&quot;:18694,&quot;alt&quot;:&quot;Jayeeta Putatunda Speaking at Women in Tech Global Conference 2027&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Jayeeta Putatunda Speaking at Women in Tech Global Conference 2027" title="Jayeeta Putatunda Speaking at Women in Tech Global Conference 2027" srcset="https://substackcdn.com/image/fetch/$s_!h2WP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3232fd-537a-4906-934b-abe48d3d715b_263x277.jpeg 424w, https://substackcdn.com/image/fetch/$s_!h2WP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3232fd-537a-4906-934b-abe48d3d715b_263x277.jpeg 848w, https://substackcdn.com/image/fetch/$s_!h2WP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3232fd-537a-4906-934b-abe48d3d715b_263x277.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!h2WP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa3232fd-537a-4906-934b-abe48d3d715b_263x277.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em>&#8220;If I had 90 days of AI budget to improve developer experience, I wouldn&#8217;t spend it evaluating another model. I&#8217;d spend it reducing engineering variation.&#8221; &#8212; <a href="https://www.linkedin.com/in/jayeeta-putatunda">Jayeeta Putatunda</a>, Director of the AI Center of Excellence, <a href="https://www.fitchratings.com/">Fitch Ratings</a></em></p><div><hr></div><p>For a Staff or Principal engineer who owns a platform, Parker&#8217;s first ninety days are not about output at all. They are about earning the ability to ask real questions of your own systems, and most organizations cannot, because observability was underinvested for years while other goals took priority. The mechanical work of fixing that, moving unstructured logging to structured events and spans, migrating to OpenTelemetry, is exactly the kind of task that is fungible from an intelligence point of view, which makes it a near-perfect fit for current coding models. It also comes with a natural verification step, feed the before and after to another model and let it judge whether the telemetry got better, and let it update the docs as it goes.</p><p>Putatunda took the same budget in a different direction, and her experience widens the point. Once her organization distributed AI broadly, it became clear that model capability was never the bottleneck, because every team had the same models. What differed was how teams solved the same problems, each with its own coding conventions, documentation norms, migration methods, and review expectations. So her first priority was not a better model. It was making the organization&#8217;s engineering knowledge reusable.</p><p>Her team translated coding norms, implementation templates, documentation standards, review criteria, and migration playbooks into reusable skills that work from any agentic environment, whether an engineer reaches for Claude Code, GitHub Copilot, or something else. That let them clear backlogs that had always lost the prioritization fight against feature work. The migration of a legacy Selenium test suite to Playwright, a slow and redundant manual slog across many teams, got done by pairing reusable migration skills with the organization&#8217;s own standards and validation steps, so the generated tests matched both the suite and the reviewers&#8217; expectations. The same approach carried into repository documentation, repetitive refactoring, and framework upgrades, the modernization work where consistency matters more than originality.</p><p>The deeper reason it worked speaks straight to the review bottleneck that AI creates. Models generate code faster than teams can review it, and when every change arrives in a different shape, every pull request costs the reviewer more attention. Standardizing the patterns let reviewers focus on correctness and business logic instead of style and structure. Putatunda does not measure any of this in tokens or lines generated. She measures whether projects moved faster, whether patterns got reused across teams instead of reinvented, and whether reviewers stopped correcting predictable issues, which is another way of saying she measures the same thing Parker does, whether the team can now do work it could not do reliably before.</p><h2>Spend on leverage, not on tokens</h2><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nxlv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2e4c34a-8f2b-40b2-92fe-b4f06e0c6838_800x800.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nxlv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2e4c34a-8f2b-40b2-92fe-b4f06e0c6838_800x800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!nxlv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2e4c34a-8f2b-40b2-92fe-b4f06e0c6838_800x800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!nxlv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2e4c34a-8f2b-40b2-92fe-b4f06e0c6838_800x800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!nxlv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2e4c34a-8f2b-40b2-92fe-b4f06e0c6838_800x800.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nxlv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2e4c34a-8f2b-40b2-92fe-b4f06e0c6838_800x800.jpeg" width="250" height="250" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b2e4c34a-8f2b-40b2-92fe-b4f06e0c6838_800x800.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:800,&quot;resizeWidth&quot;:250,&quot;bytes&quot;:93192,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/207072610?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2e4c34a-8f2b-40b2-92fe-b4f06e0c6838_800x800.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nxlv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2e4c34a-8f2b-40b2-92fe-b4f06e0c6838_800x800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!nxlv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2e4c34a-8f2b-40b2-92fe-b4f06e0c6838_800x800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!nxlv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2e4c34a-8f2b-40b2-92fe-b4f06e0c6838_800x800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!nxlv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2e4c34a-8f2b-40b2-92fe-b4f06e0c6838_800x800.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em>&#8220;Just as we don&#8217;t judge an IDE by how many keystrokes it saves, we shouldn&#8217;t judge AI by how many tokens it consumed.&#8221; &#8212; <a href="https://www.linkedin.com/in/satyamdhar">Satyam Dhar</a>, Staff Software Engineer at <a href="https://galileo.ai/">Galileo</a> and a former engineering leader at Amazon and Adobe</em></p><div><hr></div><p>Local inference and model routing are making cost-per-token murkier by the month, and Parker is honest that he does not have a clean formula to replace it. What he has is better than a formula. It is an example. For most of Honeycomb&#8217;s life, dark mode was the feature customers asked for and never got, enough of a running joke that it has its own face on the twenty-sided die new hires receive. Then, in about six months, the company shipped it, mostly because of AI.</p><p>The caveat matters as much as the result. It was not only AI. Getting there took years of groundwork, a design system and an accessible component library built to support the feature. What AI closed was the mechanical gap that always stalled the project, migrating the existing app onto the new components. The way they did it is the real lesson. Instead of one team grinding through the migration, they shipped prompts and skills so anyone could run it, and pulled in an engineering director who had built a screen seven years earlier to load the skill and update his own page. Yes, it burned a lot of tokens, and a small focused team could have optimized that count down. The accounting that matters is a five-year customer request finally closed and a company full of people who learned how to use AI.</p><p>Dhar puts a name to the mistake underneath the token frame. Too many teams treat AI like another cloud bill, counting tokens because tokens are easy to count, then assuming fewer tokens means better spending. In his experience running large-scale AI systems, some of the most valuable investments were the ones token accounting punished, better model routing, evaluation pipelines, prompt versioning, and experimentation infrastructure, all of which raised spend on paper while sharply lowering the cost of iteration.</p><p>That lower cost of iteration is where Dhar locates the actual return. When engineers can compare models safely and product teams can experiment without fear of quietly degrading production, a team can validate five ideas in the time it used to take to validate one, and the business benefit dwarfs the incremental compute. He also refuses to ignore the second-order effect that never reaches a finance dashboard, the way good infrastructure changes behavior, so engineers start proposing projects that once felt too expensive or too repetitive to attempt and product conversations get more ambitious because implementation cost no longer dominates every decision. His test is a single question, whether the system helped the team build something it otherwise would not have built.</p><h2>Count only the work that clears a bottleneck</h2><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!s07E!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f86630-c45d-40cf-b0d6-18f815424cb2_316x316.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!s07E!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f86630-c45d-40cf-b0d6-18f815424cb2_316x316.jpeg 424w, https://substackcdn.com/image/fetch/$s_!s07E!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f86630-c45d-40cf-b0d6-18f815424cb2_316x316.jpeg 848w, https://substackcdn.com/image/fetch/$s_!s07E!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f86630-c45d-40cf-b0d6-18f815424cb2_316x316.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!s07E!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f86630-c45d-40cf-b0d6-18f815424cb2_316x316.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!s07E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f86630-c45d-40cf-b0d6-18f815424cb2_316x316.jpeg" width="250" height="250" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/96f86630-c45d-40cf-b0d6-18f815424cb2_316x316.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:316,&quot;width&quot;:316,&quot;resizeWidth&quot;:250,&quot;bytes&quot;:16246,&quot;alt&quot;:&quot;About Fulcra &#8212; The Personal Data Platform Built for You&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="About Fulcra &#8212; The Personal Data Platform Built for You" title="About Fulcra &#8212; The Personal Data Platform Built for You" srcset="https://substackcdn.com/image/fetch/$s_!s07E!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f86630-c45d-40cf-b0d6-18f815424cb2_316x316.jpeg 424w, https://substackcdn.com/image/fetch/$s_!s07E!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f86630-c45d-40cf-b0d6-18f815424cb2_316x316.jpeg 848w, https://substackcdn.com/image/fetch/$s_!s07E!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f86630-c45d-40cf-b0d6-18f815424cb2_316x316.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!s07E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96f86630-c45d-40cf-b0d6-18f815424cb2_316x316.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em>&#8220;Only the work that clears a bottleneck is progress. The rest are lightbulb filaments that didn&#8217;t work.&#8221; &#8212; <a href="https://www.linkedin.com/in/mjjtiffany">Michael J.J. Tiffany</a>, co-founder and CEO of <a href="https://fulcradynamics.com/">Fulcra Dynamics</a>.</em></p><div><hr></div><p>Tiffany offers the closest thing in this issue to a replacement unit of account, and he arrived at it by correcting himself. Token use, he argues, is a low-level metric like CPU utilization, important but meaningless in aggregate. His better frame is cost per bottleneck relieved. He started somewhere looser, counting cost per accepted unit of work such as a merged pull request or a user-visible improvement, then learned the hard way to count only the subset of work that actually clears a bottleneck.</p><p>The distinction changes what you tolerate. What you buy when you pay for AI, in Tiffany&#8217;s account, is more attempts, faster loops, more surface area explored, fewer blocked humans, and occasionally, almost at random, a capability that was not economically or psychologically feasible before. Most of that is not progress, and that is fine. The failed experiments and even the shipped features that did not matter are the filaments that did not light. You let them happen, and you measure success only by the bottlenecks that came loose, which keeps the number honest in a way a token count never can.</p><h2>Judgment is the constraint for the next two years</h2><p>The token leaderboards are already gone, and the mandate that produced them is being quietly rewritten across the industry toward outcomes that are harder to game. The deeper lesson is the one Parker keeps returning to. Every layer an agent touches was built on the assumption that the consumer brings context and judgment, and that assumption is exactly what breaks when the consumer, or the incentive, is optimizing a number.</p><p>The voices in this issue disagree on the right replacement, and the disagreement is the useful part. Howe would protect trust before any metric. Putatunda would standardize the engineering practice underneath the tooling. Dhar would ask whether the team can now build what it could not before. Tiffany would count only the bottlenecks that came loose. What they share is a refusal to accept the token as the unit of value, and a conviction that the scarce resource is the judgment to decide what is worth building in the first place. The limiting factors on agentic systems are no longer model capability alone. They are judgment, integration, and the operational discipline to tell useful work apart from motion, and that is where the most consequential engineering decisions of the next two years will be made.</p><div><hr></div><p><strong>Thank you</strong> for reading this special issue of Deep Engineering on why judgment, not tokens, creates value.</p><p>We&#8217;ll be back on Thursday with more expert-led content, and next month, on the first Tuesday of August, with another special issue.</p><p><strong>Keep building,</strong></p><p>Saqib Jan</p><p>Editor-in-Chief, Deep Engineering</p><div><hr></div><p><em><span>If your company wants to reach senior developers, software engineers, and technical decision-makers, </span><a href="https://packt.omeclk.com/portal/wts/uc%5EcnN2dfNaqmD-kB-mo66%7C7g%5Ef%7Cb"><span>speak to us about partnering</span></a><span> with Deep Engineering.</span></em></p>]]></content:encoded></item><item><title><![CDATA[Tokens, Judgment, and AI Productivity with Austin Parker]]></title><description><![CDATA[On why a token is not a unit of intelligence, the illegible work token counts miss, pointing AI budgets at developer experience and observability, and judgment as the last scarce resource]]></description><link>https://deepengineering.net/p/tokens-judgment-ai-productivity-austin-parker</link><guid isPermaLink="false">https://deepengineering.net/p/tokens-judgment-ai-productivity-austin-parker</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Tue, 14 Jul 2026 14:01:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/P-dyS7sR--Q" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A token is not a unit of intelligence, and judgment is the resource that actually got scarce.</p><p><a href="https://www.linkedin.com/in/austinlparker">Austin Parker</a> has worked in observability for about a decade, co-founded <a href="https://opentelemetry.io/">OpenTelemetry</a>, and now leads AI strategy at <a href="https://www.honeycomb.io/">Honeycomb</a>. He came on to make an uncomfortable case for anyone running an AI mandate, that a token is not a fungible unit of intelligence and judgment is the resource that actually got scarce.</p><p><span>You can </span><strong>read</strong><span> or </span><strong>watch</strong><span> the full conversation here:</span></p><div id="youtube2-P-dyS7sR--Q" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;P-dyS7sR--Q&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/P-dyS7sR--Q?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><em>This session was recorded live as part of the Deep Engineering Interview Series. The transcript below has been lightly edited for clarity and readability.</em></p><div><hr></div><p><em><strong>Q. Tell us a little about yourself, what you do at Honeycomb, and what pushed you to write about tokenmaxxing.</strong></em></p><p>I have been in observability for about a decade, first as a maintainer on OpenTracing, and then I helped co-found OpenTelemetry when we merged it with OpenCensus years ago. I have worn a lot of hats across that open source work. About two years ago, when ChatGPT released, I got really interested in AI and in how we could use it to help people make sense of systems at scale. Over the past year I have been building that capability out at Honeycomb, working on our agents and our machine learning and AI stack, and getting involved in OpenTelemetry&#8217;s GenAI semantic conventions to make sure OpenTelemetry is as useful to AI as it is to humans.</p><p><em><strong>Q. You argue a token is not a fungible unit of intelligence. Walk us through one engineering decision where treating tokens as fungible led a team somewhere they regretted.</strong></em></p><p>You can find a lot of examples of that just reading the news today. Given the right goals and the right harness around a model, whether that is a frontier model or an open weights one, you can get it to do almost anything. The maintainer of Bun, who I believe now works at Anthropic, famously did a big port where the instruction was basically rewrite the whole thing in Rust, and it did it, and it passed all the tests, which is impressive.</p><p>When we read things like that it becomes easy to imagine that all I need is something that keeps the tokens flowing and I can do anything I want. That might be true, but it does not make it a good idea. Internally I have seen our engineering teams run experiments where they generate a hundred different prototypes of the same screen, or fifteen or twenty variations of a flow. What I have noticed is that most of them do not move the needle in the way you would hope, because you can ask for infinite variations on a problem. You can ask for the same program in ten different languages, or rewrite a function a hundred thousand times.</p><p>But the code and the screen and the workflow are not the things that provide value to your users. The value is in the questions one level up, about the design of the system, about how people interact with it, about whether the shape of an API really serves the people integrating with it, and whether the fundamental primitives have fitness for what people want out of them. Those are not questions that tokens answer. You can use tokens to help build solutions, but at the end of the day these are judgment calls. Those hundred variations all came out looking about the same, because they all started from one very human conception of the problem, so all of the output ends up shaped like my initial conception of it. I see this repeated when I talk to other leaders. There is a lot of here is a problem, make all these variations, and far less time asking whether we are defining the problem correctly, solving the right customer pain point, and putting the right inputs into our loops. That is why I argue that judgment is the finite resource now.</p><p><em><strong>Q. You call prompting, building internal alignment, and holding a complex system in your head illegible activities. When a leader starts measuring token usage, what specifically breaks in how those skills get rewarded?</strong></em></p><p>It is tough to say sometimes, because at really large organizations that kind of value is already not rewarded well. A lot has been written about the importance of glue work, about people whose value to the org is not how many PRs they open or how many decks they build, but knowing who needs to be in a decision chain, or being able to hear a customer request and go find the right team to make it happen. Often you or I might be frustrated with a piece of software because of one thing that shows up in a certain part of the system, when the actual reason it works that way lives somewhere else entirely. The people who can navigate that ownership map internally end up being really important to the ability of the business to execute and serve customers. The companies that are genuinely good at that tend to be the ones that, even without a clean way to measure those people, are good at not penalizing them.</p><p>Then a top down mandate arrives that says we need to use more AI, and the easiest objective measure of that is how many tokens you use. So the move becomes let us incentivize you to use more tokens and then productivity magically appears. The connective work that makes someone valuable shows up in promo packets and career ladders, but it does not show up in something you can measure quarter to quarter or day to day, because it involves very human things like relationships and time in the system rather than timing the system. The tokenmaxxing craze, 2026 to 2026, RIP, is a good example of why that value matters and why you cannot replace it by throwing more AI at the problem.</p><p><em><strong>Q. The token leaderboards at companies like Amazon rose and fell inside three months. What did the people running them learn, and why do you think the correction came so fast?</strong></em></p><p>What they learned is that maybe they should not do that again. The moment you turn any measure into a metric it stops being a good measure. Give a legendarily clever bunch of engineers a resource and tell them they are graded on how much of it they use, and they will find a way to use it. There is a famous Dilbert strip where they announce a bonus for every bug you fix, and Wally says he is going to write himself a new minivan, because you said every bug that gets identified and fixed, you did not say anything about not writing the bugs in the first place. Same thing with token leaderboards. What is strange to me is not that it happened, it is that it was not immediately obvious to the people in the decision chain that it would.</p><p>Calling them pointy haired bosses who just do not get it is a lazy answer though, and I do not think it is true. Yes, plenty of those experiments led to utter failure, and broadly it was a bad idea I would never have recommended. But I have noticed people, even at Honeycomb, who get nervous about spending company resources on an idea they think AI might be good for. A moment of use as much as you want, go wild, frees people who would otherwise be overly conservative, and some of them find something they never would have found under the old constraint. So some of this is taking the good with the bad, and some of it is making sure your employees are aligned with the overall goal. Those extremes are why it moved in and out of favor so fast, because it only takes a few people burning hundreds of thousands of dollars. Although Amazon is probably good for it.</p><p><em><strong>Q. You say AI accentuates differences that already exist rather than creating new capability. If a leader sees one team pull ahead with agents, what is the actual question they should be asking that team?</strong></em></p><p>This is the thing I feel does not get asked enough. I never want to ask what AI has let you do. The question is what is AI giving you the excuse to do that you could not have done otherwise. In every big phase shift you find individuals and teams who use the bigger meta narrative as cover to solve a different problem. With AI, it is rarely that intelligence alone solved something we could not solve before. Usually it is solving an organizational problem or a leverage problem.</p><p>Think about what it takes to ship a feature inside any organization. A big part of it is making the case to the rest of the org, especially leadership, about the expected return, and that is often a data analysis argument. It means sitting down with a pile of numbers, doing a lot of advanced Excel, building projections and forecasts, tearing through user data and sentiment data and support tickets. It turns out AI is really good at that. Give it access to your support tickets and your customer experience data and your NPS scores, spend time with it, and it helps you surface things you knew but could not back up, or genuinely new insights. Traditionally that was a product manager or business analyst function, and there is a wide range of skill there. Now a sufficiently motivated engineer or designer or support person or sales engineer can do that work, use AI to cover the parts they are not as good at, and turn around and tell the org here is what you need from me so I can go do the thing I have conviction is right.</p><p>That is what I see in the teams becoming very effective with AI. It is not simply that we can write more code now, although that helps and you can review it faster. It is that we can patch whatever hole existed in our skills or our team composition. You see it outside engineering too, and at small companies, where people who are strong at design but never programmed can now program, and people who are weak at persuasive writing can write better. We are filling the gaps in our own abilities, and that is the biggest difference between the teams getting real value out of AI and the teams that are coasting.</p><p><em><strong>Q. Should companies refuse to measure engineers by how much code their agents write, or even how many bugs they patch?</strong></em></p><p>I think that is stupid. There are so many other ways to gauge engineering efficiency. A better way to think about it is that you should measure it, but it should not be a metric. It is useful. One thing we do internally is look at how many PRs someone is reviewing, how many they are opening, and the balance of their time in the codebase. That is a useful signal, because it shows you someone spending a lot of time reviewing versus someone spending a lot of time writing, and maybe you need to think about that balance, or maybe someone is adding a lot of new things but not spending enough time understanding what is going in. Turning it into something you grade performance on is where it gets twisted, and a lot of places still treat PRs opened as an important KPI.</p><p>AI actually lets you do this analysis more holistically. Instead of stopping at how many PRs someone did, you can feed a lot of it into a model and ask real questions. Of the work you are doing, how much is causing on call burden, how much is delivering specific customer value, and are we missing the people doing important but less shiny work like refactoring. Attaching those contributions to specific outcomes is hard, like a big refactor of the component library that improved consistency and lifted the interstitial user experience survey scores because a button is now consistent across the app. Those things are hard not because they are hard to measure, but because there is so much data to go through. AI makes that analysis much easier when it is driven by high quality telemetry and instrumentation, and when the people using it have good judgment.</p><p><em><strong>Q. Your advice is to turn token budgets inward toward developer experience and observability. What does that look like in practice for a Staff or Principal engineer who owns a platform, in the first ninety days?</strong></em></p><p>The core thing you want is to get to that analysis loop, which means asking whether you even have the data, whether it is possible for you to ask those questions at all. Most of the time it is not, because as an industry we do not value observability as highly as we should, though we are getting there. I still talk to a lot of people who have their emotional support logs, because that is just how it has always been done. The historical argument against improving observability has always been that we do not have the time, we do not have the people, we have all these other goals.</p><p>This is exactly the mechanical, boring work that is fungible from an intelligence perspective. Taking your existing logging and making it structured, or moving to OpenTelemetry, is perfectly aligned with what current AI coding models are good at, and with the verification side too, because it comes down to a judgment call about whether the telemetry after the migration is as good by some heuristic as what you had before. You can prompt the verification step directly. Strip out the unstructured logging, add structured events and spans and good metrics, then feed the output into another model to judge whether it is better, worse, or the same, and let it rip. It can update the docs and the comments while it goes. That is massive leverage for an org that takes the diffuse do whatever with AI money pool and points it at something specific like developer experience and observability. It lines both things up, because more people learn to build better harnesses and prompts and loops, and you get a measurable outcome, which is that now you can answer these questions and understand your systems and run the longer horizon analysis you never had the data for. Job one is getting to the point where you can ask these questions.</p><p><em><strong>Q. Local inference and model routing make token accounting even murkier. As that lands, what mental model should engineering leaders use for AI spend instead of cost per token?</strong></em></p><p>I honestly wish I had a better answer. To use an example from Honeycomb, we have been doing agentic engineering and going through this transformation like everyone else for about a year and a half, and the results have been pretty dramatic. In one quarter we opened as many PRs as we had in the entire year up to that point, because people are leveraging AI more and a fairly small engineering team can punch above its weight on code writing and generation.</p><p>One example of that is dark mode. For as long as I have been here, dark mode was one of the most consistent customer requests, and we are a developer tool, so people really wanted it. We have an in joke where you get a D20 when you join, and one of the faces reads dark mode when, and that joke is five or six years old. For at least half the company&#8217;s life it was something we would get to eventually. Then in about six months we shipped it, mostly because of AI. I want to be clear it was not only AI. Getting there was a multi year process where our designers and product engineers built a design system and an accessible component library that also supported dark mode. At the end of that we had all the components, but we still faced the mechanical work of turning the existing app into the new ones, and that takes a lot of time.</p><p>Because we could lean on agent loops, we did it in about six months, and not as one team&#8217;s full time effort. People across the whole company chipped in, including engineering directors and managers, because instead of telling everyone to go do all these things by hand, the team shipped prompts and skills you could feed the agent to do the migration from the old components to the new ones. That unlocked a lot of human judgment, because I could pull in an engineering director who built a particular screen seven years ago and still understands what it is for, and tell them to fire up Claude Code, load this skill, point it at that page, and do the migration.</p><p>That is a good example of what your mental model should be. Yes, it used a lot of tokens, and a small focused team could have hyper optimized it. But the actual benefit is that we finally closed a feature people had asked for over five or six years, a lot of people learned how to use AI who otherwise would not have, and we got a more accessible and more beautiful product. Those benefits are tricky to quantify, but they make a real difference to customer experience, and at the end of the day that is what we are here for. It is not using AI for its own sake, it is talking to users, understanding their pain points, and meeting them where they are. So do not treat it as an AI mandate you justify only through tokens in and tokens out. Ask what tokens let you do that would otherwise have been impossible, and what lets you use AI as leverage, as a bicycle for the mind, a way to take bigger swings at bigger problems. If you can figure out how to measure that, that is your North Star.</p>]]></content:encoded></item><item><title><![CDATA[System Design Hiring Is Really a Judgment Test]]></title><description><![CDATA[Some repeat architecture while others jump straight to scaling, but the ones who stand out reason through failure, cost, and change.]]></description><link>https://deepengineering.net/p/system-design-hiring-judgment-test</link><guid isPermaLink="false">https://deepengineering.net/p/system-design-hiring-judgment-test</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 09 Jul 2026 18:22:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f3576bfb-c703-4d4d-b8c2-6e0f48dea57f_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Tech hiring is tighter, and AI has raised the bar for what companies expect from senior engineers. When companies do open senior engineering roles, they are paying closer attention to quality of hire, adding more signal to the process, and trusting the old formats less.</p><p>And the system design round is where that scrutiny concentrates, because it is the one round an AI assistant will not get you through. The system design interview may look the same as it did a couple of years ago, but the grading rubric underneath it has changed a lot.</p><p><span>Consider </span><a href="https://karat.com/engineering-interview-trends-2026/"><span>Karat&#8217;s 2026 survey</span></a><span> of 400 engineering leaders. It found that AI has widened the gap between strong and weaker engineers rather than closing it, and 73 percent of leaders now say a strong engineer is worth at least three times their total compensation. CoderPad&#8217;s </span><a href="https://coderpad.io/blog/hiring-developers/new-research-the-2026-state-of-tech-hiring-what-ai-means-for-developers-and-hiring-teams/"><span>State of Tech Hiring Report 2026</span></a><span> shows technical assessments up 48 percent globally since mid 2023, with 60 percent of talent leaders naming quality of hire as their top priority for the year.</span></p><p><span>We asked engineering leaders who personally conduct system design loops what they actually look for in a session. These insights are not usually found in prep guides and they converge on a shift most candidates have not yet noticed.</span></p><h2><span>Buzzword-first designs are a huge turn off</span></h2><p><a href="https://in.linkedin.com/in/architagarwal984"><span>Archit Agarwal</span></a><span>, Principal Member of Technical Staff at </span><a href="https://www.oracle.com/"><span>Oracle</span></a><span>, has in his engineering purview interviewed hundreds of engineers, especially for ultra-low-latency authorization work. And the failure he catches most often has nothing to do with AI. &#8220;I&#8217;ve seen a lot of engineers come in to a system design interview and, as soon as I give a problem, they start with &#8216;let&#8217;s use microservices,&#8217; and start using distributed cache,&#8221; he shares, and when he asks how many users they are planning for, the answer rarely matches the architecture they just proposed. &#8220;That is a key difference between any interview-ready engineer and a genuinely good engineer,&#8221; he explains, because &#8220;a genuinely good engineer would not want to implement everything up front.&#8221;</span></p><p><span>Agarwal&#8217;s baseline is quite blunt, that &#8220;if the problem isn&#8217;t complex yet, don&#8217;t overengineer it.&#8221; And he, like many notable leaders, holds the position that &#8220;microservices aren&#8217;t the magical fix that fixes bad architecture. They just distribute that over the network.&#8221; What he seeks instead is one to two minutes of genuine alignment at the start, functional requirements first to establish what is being built and what the user needs, then the nonfunctional requirements that set scale, consistency, and latency. &#8220;Nonfunctional requirements are the ones that decide the architecture&#8212;not the other way around.&#8221; And not every system earns planet-scale treatment, since a tool used only by a company&#8217;s own engineers needs no multi-region deployment, and proposing one tells him the candidate is performing rather than designing.</span></p><p><span>He also weighs cost. It is something that always runs through his evaluation the same way, and he shared with us a line that reframed his own career, &#8220;a good engineer would design for performance, but a great engineer would design for performance per dollar,&#8221; and he pushes candidates to weigh a latency win against the infrastructure bill it creates. That cost lens is exactly where the interview has changed most, because the most expensive and least predictable component on the whiteboard is now the model.</span></p><blockquote><p>Earlier this year, we spoke with Agarwal about <a href="https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews">trade-offs in modern system design</a>, and later published another piece featuring his candidate-side advice on <a href="https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews">why senior engineers fail system design interviews</a>.</p></blockquote><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Treat every AI component as one that will fail</span></h2><p><span>AI is on every whiteboard now, and the reliability assumptions underneath it are exactly what interviewers have started to stress test. </span><a href="https://www.linkedin.com/in/rohitrpoduval"><span>Rohit Poduval</span></a><span>, Senior Software Engineer on </span><strong><span>Prime Video Trust</span></strong><span> and </span><strong><span>Safety</span></strong><span> at </span><strong><span>Amazon</span></strong><span>, opens with a test that did not exist two years ago. &#8220;Does the candidate treat an AI component as an unreliable dependency or a magic box?&#8221; &#8220;I ask them to design a system that uses an LLM. Content classification, recommendation filtering, whatever fits the role,&#8221; he explains, and then he watches what they do with it. &#8220;Weak candidates draw a box labeled LLM and move on, like it&#8217;s a database that always returns the right answer,&#8221; while strong candidates, he shares, &#8220;immediately ask: What happens when it&#8217;s wrong? What&#8217;s the fallback? How do we know it&#8217;s drifting?&#8221;</span></p><p><span>That one reaction most often separates people who have shipped AI systems from people who have read about them. LLMs return different outputs for the same input, and Poduval points out that you can enforce structure on the output, valid JSON or responses from an allowed set, but you cannot assert that the decision itself is correct with a traditional pass and fail test. Candidates who have built these systems talk about evaluation suites, guardrails that constrain what the model can do regardless of its output, and monitoring for drift in production. But candidates who have not will handwave past every one of those concerns. And that handwaving is quite visible within minutes.</span></p><p><span>&#8220;The next thing I probe: how do they test it?&#8221; Poduval shares, because traditional assertions do not apply to AI components. &#8220;Strong candidates talk about benchmark evaluations: a defined set of inputs with expected behavioral boundaries (not exact outputs) that the system must satisfy before shipping.&#8221; They also talk about continuous auditing against real traffic after launch, not just checks before deployment, because a model that passed benchmarks last month might be drifting today. His deeper observation is that with AI systems you often cannot enumerate every right answer before launch, so the old playbook of define requirements, build to spec, and ship no longer closes. The engineers he rates highly design for safe iteration instead, with shadow mode, holdback groups, evaluation frameworks from day one, and human escalation paths, because they expect to keep refining the definition of correct after launch. &#8220;That&#8217;s exactly where the memorized answers fall apart.&#8221;</span></p><p><span>Taking into consideration the resource-intensive workloads and the shrinking budgets, cost has to be something a candidate can defend out loud. Poduval changes a constraint mid session and tells the candidate the system now needs to handle ten times the traffic. &#8220;Candidates who&#8217;ve never dealt with inference costs at scale just say &#8216;add more instances&#8217;. Candidates who&#8217;ve lived it start talking about which requests actually need an expensive model vs which can be routed to cheaper classifiers,&#8221; he shares. They design tiered systems with routing logic that decides in real time which compute budget each request deserves.</span></p><p><span>The practical takeaway for anyone walking into a loop this year is direct, so state the fallback, the evaluation plan, and the per-request cost tier for any AI component before the interviewer has to ask. Everything on that checklist has an architectural counterpart, and we have captured the nuances of each in </span><a href="https://in.linkedin.com/in/sampritimitra"><span>Sampriti Mitra&#8217;s</span></a><span> practical deep-dive on</span><a href="https://deepengineering.net/p/core-architectural-patterns-for-llm-system-design"><span> core architectural patterns for LLM system design</span></a><span>, which walks through the gateway, tiered fallback, model routing, and evaluation patterns behind it.</span></p><h2><span>Constraint changes reveal who patches and who rethinks</span></h2><p><a href="https://www.linkedin.com/in/chandu-p-a5a896118/"><span>Chandu Putta</span></a><span>, Senior Software Engineer at </span><strong><span>MissionSquare Retirement</span></strong><span>, in our email interview shared what earns his confidence. &#8220;The signal that I trust the most is not the architecture a candidate produces, it&#8217;s how they think if I change a constraint in the middle of the interview.&#8221; In a recent loop he had candidates design a document processing pipeline, and midway through he added an LLM call to the extraction step, with variable latency and non-deterministic output. The majority of candidates bolted a retry mechanism onto their existing design and moved on. One candidate stopped and asked, &#8220;What&#8217;s the acceptable failure mode &#8212; an explicit error to the engineer or a silent degradation?&#8221; he recalls.</span></p><p><span>That question, to his mind, is the level marker, because it treats the new component as a design problem rather than a network hiccup. Putta is blunt about how he sizes this up, as &#8220;rehearsed candidates patch the existing design, but the real senior engineers discard it and start fresh with this new anchor.&#8221; The engineers he considers qualified for the job also ask about context boundaries and cost at scale without being prompted, he adds, and they raise a question that prep articles don&#8217;t teach, which is who decides whether the model output is good enough, the technology team or the business team. An engineer who asks that has sat in the meeting where the answer was contested.</span></p><p><a href="https://www.linkedin.com/in/ed-tian"><span>Edward Tian</span></a><span>, creator of </span><strong><span>GPTZero</span></strong><span>, sees the same moment through a different lens. He shares with us how his interviews confirm the pattern from another direction. &#8220;When we see candidates redesign their system following the introduction of a new constraint, we see that they are able to quickly identify the most significant aspect of their original design that has become the bottleneck,&#8221; he explains, and they design new components to replace it.</span></p><p><span>But weak engineers make incremental changes to accommodate the new requirement without revisiting the larger trade-offs the original design was built on. His sharpest distinction is about posture rather than knowledge, because &#8220;most engineers will assume that their original design is now incorrect after the constraint has been added and defend the changes made to it,&#8221; he shares. &#8220;Senior engineers will simply continue to increase the rate of evolution of their design until it meets their new constraint.&#8221;</span></p><p><span>Tian also shares how he has noticed a newer failure mode that improved AI tooling has made common. &#8220;Many candidates can produce beautiful pipeline diagrams for their LLM, but when we dig through more complex failure modes, for example context window, they really struggle,&#8221; and they default to talking about scaling the model. The diagram generated fluency, but probing exposes it.</span></p><p><span>Agarwal looks for something similar when he changes the requirements mid session. The candidate who absorbs the change, restates it in their own words to confirm alignment, and then highlights which parts of the system must change and which remain intact is showing him a structured redesign instinct rather than attachment to a diagram. He also likes to give away something candidates rarely hear from the other side of the table. &#8220;The curveballs that the interviewer gives you will never be in a way that you will have to scrap the complete diagram, the complete architecture, unless you were already off the track.&#8221; A constraint change is a test of surgical judgment, so the working advice is to name out loud what survives and what dies before you touch the whiteboard.</span></p><h2><span>Can you defend the choice you just made?</span></h2><p><a href="https://www.linkedin.com/in/prakharchaube/"><span>Prakhar Chaube</span></a><span>, Senior Software Engineer at </span><a href="https://www.insurgrid.com/"><span>InsurGrid</span></a><span> who previously held the same level at </span><strong><span>Whatfix</span></strong><span> and has conducted more than 25 technical rounds, argues the design round should never hunt for a correct answer. He assesses the thought process and then questions the candidate on their component choices, because that questioning mirrors what happens inside real companies. &#8220;Most companies have some form of an Architectural Review Board, meaning it&#8217;s rare that a new implementation will not be questioned even if it is the right choice.&#8221; Candidates should know why the standard patterns work as standards rather than assuming a plug and play model, and the ones who can stand by their choice under pressure usually survive a real review. In his loops he is categorical about telemetry, since &#8220;if a candidate is not mentioning that in an interview it is a big no,&#8221; because &#8220;without proper observability the system is doomed to fail at scale and cause chaos during debugging.&#8221;</span></p><p><span>Chaube has also built a ladder for assessing AI fluency that moves well past the 2024 questions. &#8220;I first assess the basics like parameters, prompt structure, prompt security, hallucination,&#8221; he shares, then he wants to dig deeper into implementation scenarios drawn from real products, autonomous tasks orchestrated by a central model such as a code review agent or a login and scraping system. The final rung, he explains, is how candidates handle agent-guided coding, because &#8220;I look at their queries and flow&#8221; rather than their claimed experience. This method, he says, reliably separates candidates who prepared just for the interview from candidates who are genuinely strong. The preparation gap shows up in the follow-ups, never in the opening answer.</span></p><p><a href="http://fr.linkedin.com/in/sandor-dargo"><span>S&#225;ndor Darg&#243;</span></a><span>, senior software engineer at </span><a href="https://engineering.atspotify.com/about"><span>Spotify</span></a><span>, has observed the same gap from years of interviewing in the C++ world, and his experience reinforces Chaube&#8217;s from a different stack. He describes a depth gap first, because a language whose standard runs about 2,000 pages guarantees that even genuine experts have blind areas, and the honest move is admitting them.</span></p><p><span>Darg&#243; in </span><a href="https://deepengineering.net/p/clean-c-code-and-the-hidden-cost"><span>our live interview session</span></a><span> shared how he once told an interviewer he did not know a topic and did not want to guess. The interviewer replied that their team did not use it either, and he got hired. His second observation lands harder, as he has seen senior engineers who handle architectural questions with real thoughtfulness struggle to write simple algorithms live under pressure. Rehearse defending your choices against a hostile follow-up, not presenting them to a friendly one, and practice the small problems you assume you have outgrown.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Safety-critical systems raise the same bar higher</span></h2><p><span>With more than 15 years in secure embedded systems, </span><a href="https://www.linkedin.com/in/hareesha-rameshappa/"><span>Hareesha KoratikereRameshappa</span></a><span>, Senior Technical Product Manager at </span><a href="https://phantom.ai/"><span>Phantom AI</span></a><span>, likes to conduct design interviews where the system under discussion can kill someone. He evaluates whether a candidate can translate high-level product requirements into a structured, developable architecture without compromising safety, cost, accuracy, reliability, and power consumption. He looks for system-level thinking across hardware, firmware, application software, and flashing methodologies like OTA. And what he expects, he shares, is that &#8220;the candidate should understand the key elements of system design, such as system and operating modes, error and fault detection, active and history fault mode managements, transition to the safe mode, and recovery from the fault mode.&#8221; They must also supply the inputs verification depends on, KPIs, requirements, corner cases, and they must plan to monitor the system after deployment rather than treating launch as the finish line.</span></p><p><span>His trade-off vocabulary sounds exactly like the web scale interviewers we heard from earlier, and that essentially is the point. Rameshappa presses candidates on functionality versus performance and accuracy versus resource consumption, and he expects them to justify design decisions with data and adjust the design when another team&#8217;s feedback is valuable. At Phantom AI he shared how he struggled to find candidates who combined production experience with sensor selection, calibration integration, and corner case validation, which tells you how rare the full reasoning package is even among experienced engineers.</span></p><p><span>This means that what the interviewers are testing is not a property of a tech stack. It is engineering judgment, and it transfers from a recommendation pipeline to a multi-camera ADAS system without translation.</span></p><h2><strong><span>Trade-off reasoning still decides the level</span></strong></h2><p><span>Everything above rests on a foundation that has not moved, and </span><a href="https://www.linkedin.com/in/dhirendra-sinha"><span>Dhirendra Sinha</span></a><span>, Software Engineering Manager at </span><a href="https://about.google/"><span>Google</span></a><span>, described it </span><a href="https://deepengineering.net/p/designing-for-scale-and-resilience"><span>last year in one of our conversations</span></a><span> with him and Tejas Chopra. &#8220;It&#8217;s easy to choose between good and bad solutions,&#8221; Sinha told us, &#8220;but senior engineers often have to choose between two good options. I want to hear their reasoning.&#8221; He also carries a line from a chief architect at Yahoo that reframes how scale changes judgment, because when an engineer dismissed a corner case as one in a million, the architect replied, &#8220;One in a million happens every hour here.&#8221; Scale invalidates assumptions, and mostly every leader in this piece is testing whether a candidate&#8217;s assumptions survive contact with it.</span></p><p><span>Agarwal&#8217;s version of the same evaluation focuses on where a candidate chooses to go deep. When a candidate takes one area, distributed storage or authentication or performance engineering, and drives into real depth, he takes that as an engineer who &#8220;understands the gravity of things&#8221; rather than one collecting surface vocabulary. And he calibrates the questions a candidate asks against their level, because clarification questions are always welcome, but a flood of questions too basic for someone with a strong resume tells him the candidate has never actually thought about the system in front of them.</span></p><p><a href="https://www.linkedin.com/in/chopratejas"><span>Tejas Chopra</span></a><span>, Senior Engineer at </span><strong><span>Netflix</span></strong><span>, closes the loop between the old test and the new one with a single follow-up he has asked for years. When a candidate picks SQL over NoSQL or strong over eventual consistency, he asks what changes if the user base grows tenfold, and the AI era version of that question is exactly what some interviewers in technical rounds now ask about model cost and failure modes. The component changed, but the muscle is the same. For readers who want the underlying vocabulary these interviews keep invoking, consistency, availability, partition tolerance, and the read and write trade-offs that connect them, Sinha and Chopra&#8217;s practical deep-dive on</span><a href="https://deepengineering.net/p/distributed-system-attributes"><span> distributed system attributes</span></a><span> from their book </span><a href="https://www.packtpub.com/en-in/product/system-design-guide-for-software-professionals-9781805124993"><span>System Design Guide for Software Professionals</span></a><span> walks through all of it with worked examples.</span></p><h2><strong><span>What to do in your next loop</span></strong></h2><p><span>Ask for the acceptable failure mode before you design anything, because that single question outranks any diagram you will draw. Treat every AI box in your architecture as a component that will be wrong, and price its errors with a fallback, an evaluation plan, and a cost tier per request class. When a constraint changes mid session, say out loud which parts of your design survive and which parts die before you redraw anything. Say what you would monitor without being asked, and say I don&#8217;t know about the right things, because every interviewer quoted here treats honesty about limits as seniority rather than weakness.</span></p><p><span>The rehearsed candidate optimizes for the first answer, and the bar has moved to the third follow-up. That is the whole shift.</span></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p><strong><span>Contributors</span></strong></p><ul><li><p><strong><span>Archit Agarwal</span></strong><span>, Principal Member of Technical Staff, Oracle.</span><a href="https://in.linkedin.com/in/architagarwal984"><span> LinkedIn</span></a><span>. Read his candidate-side advice in</span><a href="https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews"><span> Why Senior Engineers Fail System Design Interviews</span></a><span>.</span></p></li><li><p><strong><span>Rohit Poduval</span></strong><span>, Senior Software Engineer, Amazon (Prime Video Trust &amp; Safety).</span><a href="https://www.linkedin.com/in/rohitrpoduval/"><span> LinkedIn</span></a></p></li><li><p><strong><span>Naga Chand Putta</span></strong><span>, Senior Software Engineer, MissionSquare Retirement. </span><a href="https://www.linkedin.com/in/chandu-p-a5a896118/"><span>LinkedIn</span></a></p></li><li><p><strong><span>Edward Tian</span></strong><span>, Founder, GPTZero. </span><a href="https://www.linkedin.com/in/ed-tian"><span>LinkedIn</span></a></p></li><li><p><strong><span>Prakhar Chaube</span></strong><span>, Senior Software Engineer (SDE-3), InsurGrid. </span><a href="https://www.linkedin.com/in/prakharchaube"><span>LinkedIn</span></a></p></li><li><p><strong><span>Hareesha KoratikereRameshappa</span></strong><span>, Sr. Technical Product Manager, Phantom AI. </span><a href="https://www.linkedin.com/in/hareesha-rameshappa/"><span>LinkedIn</span></a></p></li><li><p><strong><span>S&#225;ndor Darg&#243;</span></strong><span>, Senior Software Engineer, Spotify.</span><a href="https://sandordargo.com/"><span> sandordargo.com</span></a></p></li><li><p><strong><span>Dhirendra Sinha</span></strong><span> (Google) and </span><strong><span>Tejas Chopra</span></strong><span> (Netflix), authors of </span><em><span>System Design Guide for Software Professionals</span></em><span> (Packt). Read their</span><a href="https://deepengineering.net/p/distributed-system-attributes"><span> free chapter</span></a><span>.</span></p></li></ul><p><strong><span>Sources</span></strong></p><ul><li><p><span>Karat, &#8220;Engineering Interview Trends in 2026,&#8221; March 2026. </span><a href="https://karat.com/engineering-interview-trends-2026/"><span>https://karat.com/engineering-interview-trends-2026/</span></a></p></li><li><p><span>CoderPad, &#8220;State of Tech Hiring Report 2026,&#8221; March 2026. </span><a href="https://coderpad.io/blog/hiring-developers/new-research-the-2026-state-of-tech-hiring-what-ai-means-for-developers-and-hiring-teams/"><span>https://coderpad.io/blog/hiring-developers/new-research-the-2026-state-of-tech-hiring-what-ai-means-for-developers-and-hiring-teams/</span></a></p></li><li><p><span>Deep Engineering, &#8220;Why Senior Engineers Fail System Design Interviews,&#8221; May 2026. </span><a href="https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews"><span>https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews</span></a></p></li><li><p><span>Deep Engineering #36, Archit Agarwal on System Design Trade-offs, February 2026. </span><a href="https://deepengineering.substack.com/p/deep-engineering-36-archit-agarwal"><span>https://deepengineering.substack.com/p/deep-engineering-36-archit-agarwal</span></a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Rust Patterns That Leverage the Type System]]></title><description><![CDATA[A deep-dive into four Rust patterns, NewType, parse don't validate, TypeState, and sealed traits, excerpted with permission from Evan Williams' Design Patterns and Best Practices in Rust.]]></description><link>https://deepengineering.net/p/rust-patterns-that-leverage-the-type-system</link><guid isPermaLink="false">https://deepengineering.net/p/rust-patterns-that-leverage-the-type-system</guid><dc:creator><![CDATA[Packt]]></dc:creator><pubDate>Thu, 09 Jul 2026 12:32:44 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/9830cb2b-3fe2-48b8-85db-5e66c91cfd67_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The patterns in this deep-dive represent well-established techniques from type theory and functional programming, adapted to work with Rust's unique ownership and borrowing system. While these patterns have origins in languages such as Haskell and <strong>MetaLanguage</strong> (<strong>ML</strong>), as well as research on type systems, Rust brings its own contributions: zero-cost abstractions, compile-time enforcement through ownership, and integration with systems programming.</p><p>What makes these patterns particularly valuable in Rust isn&#8217;t their novelty. Most have decades of history in other languages. It is how Rust&#8217;s design makes them practical for systems programming. The combination of strong static typing, zero-cost abstractions, and memory safety without garbage collection creates opportunities to apply these patterns in contexts where they were previously impractical.</p><p>We&#8217;ll apply these patterns to Samsa, a publish/subscribe microservice in the style of a miniature Kafka that the book builds across its later chapters, patterns that have proven their worth across multiple programming language communities. As we explore each pattern, we&#8217;ll examine both its origins and how Rust&#8217;s unique features enhance or constrain its application.</p><blockquote><p><em>This deep-dive is excerpted from <a href="https://www.packtpub.com/en-us/product/design-patterns-and-best-practices-in-rust-9781836209478">Design Patterns and Best Practices in Rust</a> by <a href="https://www.linkedin.com/in/evan-williams-1512092">Evan Williams</a>, shared with permission from Packt, with all rights remaining with the publisher and no reproduction or redistribution without written consent. You can get the full book <a href="https://www.packtpub.com/en-us/product/design-patterns-and-best-practices-in-rust-9781836209478">here</a>.</em></p></blockquote><p>In this chapter the book covers the following main topics:</p><ul><li><p>The NewType pattern</p></li><li><p>Parse, don&#8217;t validate</p></li><li><p>The TypeState pattern</p></li><li><p>Sealed traits</p></li></ul><p>Each pattern builds on the previous ones: the NewType pattern creates distinct types; parse, don&#8217;t validate ensures those types hold only valid data; the TypeState pattern tracks valid state transitions; and sealed traits control which types can participate in our APIs.</p><div class="callout-block" data-callout="true"><p><strong>Technical requirements</strong></p><p>The source code for the exercises can be found on GitHub at <a href="https://github.com/PacktPublishing/Design-Patterns-and-Best-Practices-in-Rust"><span>https://github.com/PacktPublishing/Design-Patterns-and-Best-Practices-in-Rust</span></a>. The repository is organized by chapter. The relevant exercises for this chapter are at <a href="https://github.com/PacktPublishing/Design-Patterns-and-Best-Practices-in-Rust/tree/main/ch10"><span>https://github.com/PacktPublishing/Design-Patterns-and-Best-Practices-in-Rust/tree/main/ch10</span></a>.</p></div><div><hr></div><h2>The NewType pattern</h2><p>The NewType pattern has a long history in statically-typed functional programming, particularly in Haskell, where the <code>newtype</code> keyword has been a language feature since the 1990s. The pattern also appears in ML, OCaml, and other languages with strong type systems. The core idea of wrapping existing types in distinct types for semantic clarity and type safety predates Rust. It is a response to the issue that values that are conceptually different, such as temperature and weight, have the same type representation in the code, for example, a float. From the compiler&#8217;s perspective, they are the same, and that leads to bugs when these values are inadvertently mixed up.</p><p>What Rust brings to this well-established pattern is zero-cost abstraction: the wrapped type compiles to exactly the same representation as the underlying type, with no runtime overhead. In languages with garbage collection or boxing, creating wrapper types often incurs performance costs. Rust&#8217;s design ensures that type safety is free at runtime while remaining enforceable at compile time.</p><h3>Identifying the problem</h3><p>In our Samsa system, we currently use primitive types such as <code>u64</code> for topic IDs, consumer IDs, and message offsets. This approach, which we will call <strong>primitive obsession</strong>, makes it easy to accidentally pass the wrong type of ID to a function. If you&#8217;ve worked with similar systems, you&#8217;ve probably experienced the confusion this creates. The NewType pattern solves this problem with compile-time type checking.</p><p>Let&#8217;s examine our current API and identify where primitive obsession creates issues:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;fa4f79f2-9f41-4b8f-9707-9b71b27ba5f6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">// Current API - prone to mix-ups
impl Broker {
     pub fn subscribe(&amp;self, topic_id: u64, consumer_id: u64) -&gt; Result&lt;()&gt; {
         // Easy to accidentally swap these parameters!
         if !self.topics.contains_key(&amp;topic_id) {
             return Err("Topic not found".to_string());
         }
         Ok(())
     }
 }

impl Consumer {
     pub fn new(broker: Arc&lt;Broker&gt;, topic_id: u64, offset: u64) -&gt; Self {
         // Again, easy to mix up topic_id and offset
         Self { broker, topic_id, offset }
     }
 }</code></pre></div><p>We&#8217;re using raw <code>u64</code> values for different semantic concepts, making it easy to pass the wrong type of ID and causing runtime bugs that are difficult to debug. What is very unfortunate is that the compiler cannot help us identify logical errors when we pass a topic ID instead of a consumer ID.</p><h3>Implementing type-safe wrappers</h3><p>Let&#8217;s create wrapper types that eliminate these mix-ups. Our approach has three parts:</p><ol><li><p>Define semantic wrapper types for our different ID concepts (shown in the first code block)</p></li><li><p>Implement basic methods for usability, such as constructors and accessors (shown in the second code block)</p></li><li><p>Add <code>Display</code> implementations for user-friendly output (shown in the third code block)</p></li></ol><p>Let&#8217;s define semantic wrapper types: <code>TopicId</code> for topic identification, <code>ConsumerId</code> for tracking individual consumers, and <code>MessageId</code> for tracking individual messages. We&#8217;ll use plain <code>u64</code> for message offsets for simplicity (in the current design, message offsets don&#8217;t benefit much from the additional type safety):</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;753a42de-c85e-4e7f-8fe4-df765a050ef2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">use std::fmt;

/// Topic identifier for tracking individual topics
#[derive(Debug, Clone, PartialEq, Eq, Hash)]
pub struct TopicId(String);

/// Consumer identifier for tracking individual consumers
#[derive(Debug, Clone, PartialEq, Eq, Hash)]
pub struct ConsumerId(String);

/// Message identifier for tracking individual messages
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
pub struct MessageId(u64);</code></pre></div><p>These NewTypes provide type safety through distinct wrapper types. <code>TopicId</code> and <code>ConsumerId</code> wrap <code>String</code> to allow meaningful identifiers such as <code>"user.events"</code> or <code>"consumer-1"</code>, while <code>MessageId</code> wraps <code>u64</code> for numeric message tracking. Each type has meaningful derives, such as <code>Hash</code> for use in collections. Note that <code>MessageId</code> includes <code>Copy</code> since <code>u64</code> is copyable, while the <code>String</code>-based types do not.</p><p>Now, let&#8217;s add methods for creating and accessing these types:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;b8017b57-f2e3-4a69-b06c-f1e9a178223c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">impl TopicId {
    pub fn new(topic: impl AsRef&lt;str&gt;) -&gt; Result&lt;Self&gt; {
        let topic = topic.as_ref();

        if topic.is_empty() {
            return Err(SamsaError::topic("Topic name cannot be empty"));
        }

        if topic.len() &gt; 128 {
            return Err(SamsaError::topic(
                format!("Topic name too long: {} chars (max 128)",
                         topic.len())
            ));
        }

        Ok(TopicId(topic.to_string()))
    }

    pub fn as_str(&amp;self) -&gt; &amp;str {
        &amp;self.0
    }
}

impl ConsumerId {
    pub fn new(id: impl AsRef&lt;str&gt;) -&gt; Result&lt;Self&gt; {
        let id = id.as_ref();

        if id.is_empty() {
            return Err(SamsaError::consumer("Consumer ID cannot be
                                             empty"));
        }

        Ok(ConsumerId(id.to_string()))
    }

    pub fn as_str(&amp;self) -&gt; &amp;str {
        &amp;self.0
    }
}

impl MessageId {
    pub fn new(id: u64) -&gt; Self {
        MessageId(id)
    }

    pub fn value(&amp;self) -&gt; u64 {
        self.0
    }
}</code></pre></div><p>Each type provides a new constructor and an accessor method. <code>TopicId</code> and <code>ConsumerId</code> validate their inputs at construction time, returning <code>Result</code> to signal potential validation failures. This is the NewType pattern combined with early validation: once you have <code>TopicId</code> or <code>ConsumerId</code>, you know it contains valid data. <code>MessageId</code> uses simple <code>u64</code> wrapping since numeric IDs don&#8217;t require validation.</p><p>Finally, let&#8217;s add <code>Display</code> implementations for user-friendly output:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;e0e3a0b2-2195-4c73-ae55-cde24df118ed&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">impl fmt::Display for TopicId {
    fn fmt(&amp;self, f: &amp;mut fmt::Formatter&lt;'_&gt;) -&gt; fmt::Result {
        write!(f, "{}", self.0)
    }
}

impl fmt::Display for ConsumerId {
    fn fmt(&amp;self, f: &amp;mut fmt::Formatter&lt;'_&gt;) -&gt; fmt::Result {
        write!(f, "{}", self.0)
    }
}

impl fmt::Display for MessageId {
    fn fmt(&amp;self, f: &amp;mut fmt::Formatter&lt;'_&gt;) -&gt; fmt::Result {
        write!(f, "{}", self.0)
    }
}</code></pre></div><p>The <code>Display</code> implementations output the inner value directly. For example, a <code>TopicId</code> containing <code>"user.events"</code> prints as <code>user.events</code>. If we were writing production code, we could also implement a custom <code>Debug</code> trait and print the IDs in a way that is more helpful for debugging, such as <code>"TopicID(user.events)"</code>.</p><p>We&#8217;ve created distinct types for different domain concepts. Each type wraps its inner value and prevents accidental mixing. It&#8217;s now impossible to accidentally pass a <code>TopicId</code> where a <code>ConsumerId</code> is expected. The compiler catches this at compile time, eliminating entire classes of bugs. Each wrapper type is distinct at compile time.</p><p>For <code>MessageId</code> wrapping <code>u64</code>, the type compiles to the same representation as <code>u64</code> at runtime, creating a true zero-cost abstraction. <code>TopicId</code> and <code>ConsumerId</code> use <code>String</code> for semantic identifiers but still provide compile-time type safety.</p><h3>Enhanced Samsa API</h3><p>Let&#8217;s see how NewTypes prevent parameter mix-ups. Consider a subscription function that takes both a consumer and a topic:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;78867920-b4f2-4d31-a312-702083fcaf2a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">fn subscribe(consumer: ConsumerId, topic: TopicId) -&gt; Result&lt;(), SamsaError&gt; {
      println!("Subscribing {} to {}", consumer, topic);
      Ok(())
  }

  // Usage - this compiles:
  let topic = TopicId::new("orders.created")?;
  let consumer = ConsumerId::new("analytics-service")?;
  subscribe(consumer, topic)?;

  // But this won't compile - arguments swapped:
  // subscribe(topic, consumer);
  //           ^^^^^ expected `ConsumerId`, found `TopicId`</code></pre></div><p>Notice how <code>subscribe</code> takes <code>TopicId</code> and <code>ConsumerId</code> as separate types. The compiler would reject any attempt to swap them. Without NewTypes, both <code>topic</code> and <code>consumer</code> would be <code>String</code>, and swapping them would compile silently &#8211; a bug you&#8217;d only discover at runtime (maybe while in production). With NewTypes, the compiler catches the mistake immediately. This is especially valuable in APIs with multiple string-like parameters where mix-ups are easy to make and hard to debug.</p><p>Now that we have our type-safe wrappers, let&#8217;s apply them to Samsa&#8217;s API to see how they prevent the parameter mix-ups we identified earlier. The following code snippet enhances the consumer API with type safety:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;8694ec10-d5cf-4595-ad87-4a2dba6c7d13&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">pub struct Consumer {
     broker: Arc&lt;Broker&gt;,
     id: ConsumerId,
     topic: TopicId,
     offset: u64,
 }

impl Consumer {
     pub fn new(broker: Arc&lt;Broker&gt;, id: ConsumerId, topic: TopicId,
                offset: u64) -&gt; Self {
         Self { broker, id, topic, offset }
     }

    pub fn poll(&amp;mut self) -&gt; Result&lt;Option&lt;Message&gt;, String&gt; {
         let events = self.broker.fetch(self.topic, self.offset, 1)?;
    
        if let Some(message) = events.into_iter().next() {
             self.offset += 1;
             Ok(Some(message))
         } else {
             Ok(None)
         }
     }
 }</code></pre></div><p>The <code>Consumer</code> struct stores typed identifiers for its broker connection, ID, topic, and current offset. The <code>poll()</code> method demonstrates how advancement works with the typed consumer. Each field uses the appropriate NewType, making the struct&#8217;s purpose clear from its type signature alone.</p><p>Our entire API now uses type-safe identifiers. Function signatures are self-documenting, and the compiler prevents mixing up different types of IDs. This eliminates bugs where IDs are confused, provides better error messages, and makes the code more maintainable. With this design, Samsa&#8217;s consumer API becomes self-documenting and resistant to ID mix-ups.</p><p>The NewType pattern provides compile-time guarantees about type correctness while maintaining the runtime performance of primitive types. This combination of safety and efficiency is what makes the pattern particularly valuable in Rust.</p><p>However, while NewTypes give us distinct types, they don&#8217;t enforce <em>validity</em>. A <code>MessageId</code> could wrap any <code>u64</code>, which is OK if every <code>u64</code> is a valid message ID. But if some are not, including those invalid values in our data type can turn into an issue. In the next section, we&#8217;ll see how to combine type distinctions with validity guarantees using the &#8220;parse, don&#8217;t validate&#8221; principle.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Parse, don&#8217;t validate</h2><p>Alexis King popularized the <strong>parse, don&#8217;t validate</strong> principle with her influential 2019 blog post of the same name (<a href="https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-validate/"><span>https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-validate/</span></a>). King&#8217;s articulation of the principle, which states that parsing generates types that can represent only valid data, while validation merely checks data after the fact, has its roots in functional programming and type theory and helped popularize and name the pattern.</p><p>Rust&#8217;s type system makes this principle particularly natural to apply. The combination of strong typing, pattern matching, and the <code>Result</code> type creates an environment where <em>parse, don&#8217;t validate</em> feels like the path of least resistance rather than additional work.</p><p>The principle suggests that instead of scattering validation logic throughout a code base, we should parse data once at system boundaries into types that can only represent valid states. This transforms validation from a repeated burden into a guarantee encoded in types.</p><h3>The problem with scattered validation</h3><p>Let&#8217;s examine how our current Samsa configuration system handles validation. The current approach is validation scattered everywhere:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;8fc7f1b5-2791-47a9-8cea-64d8e3c3c16b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">pub struct BrokerConfig {
     pub max_connections: i32,    // Could be negative!
     pub buffer_size: usize,      // Could be zero!
     pub port: u16,              // Could be in reserved range!
 }

impl Broker {
     pub fn new(config: BrokerConfig) -&gt; Result&lt;Self, String&gt; {
         // Validation #1 - at construction
         if config.max_connections &lt;= 0 {
             return Err("max_connections must be positive".to_string());
         }
         if config.buffer_size == 0 {
             return Err("buffer_size must be non-zero".to_string());
         }
         Ok(Self { config })
     }

    pub fn accept_connection(&amp;self) -&gt; Result&lt;(), String&gt; {
         // Validation #2 - in business logic (defensive programming)
         if self.config.max_connections &lt;= 0 {
             return Err("Invalid max_connections".to_string());
         }
         Ok(())
     }
 }</code></pre></div><p>Notice how the same validation check, <code>if max_connections &lt;= 0</code>, appears in two different places: once during configuration creation and once defensively in the business logic. This duplication is the core problem. Each location must remember to perform the check, use the same condition, and handle errors consistently. If the validation rule changes (say, <code>max_connections</code> must be at least 10), every location must be updated. This approach is error-prone, inefficient, and creates inconsistent validation logic across different parts of the system.</p><h3>Creating types that enforce validity</h3><p>The solution is to create wrapper types that can only hold valid values. Instead of validating a raw <code>u32</code> everywhere it&#8217;s used, we create a <code>PositiveU32</code> type that validates once at construction time. Any code that receives a <code>PositiveU32</code> type knows, by the type alone, that the value has already been validated.</p><p>Let&#8217;s redesign our configuration using this approach. We&#8217;ll start by creating a wrapper type for positive integers that cannot be zero or negative:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;58b70775-a0f9-413f-81dd-288e5e4776ea&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">use std::num::NonZeroUsize;

#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct PositiveU32(u32);

impl PositiveU32 {
     pub fn new(value: u32) -&gt; Result&lt;Self, ConfigError&gt; {
         if value &gt; 0 {
             Ok(Self(value))
         } else {
             Err(ConfigError::MustBePositive("value".to_string()))
         }
     }

    pub fn get(self) -&gt; u32 {
         self.0
     }
 }</code></pre></div><p>Next, we create a similar wrapper for network ports that ensures the port number is not in the reserved range (ports <code>0</code>&#8211;<code>1023</code>):</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;11f42ce4-8094-418d-84ea-d784a67b5405&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct NetworkPort(u16);

impl NetworkPort {
     pub fn new(port: u16) -&gt; Result&lt;Self, ConfigError&gt; {
         if port &gt; 1023 {  // Ports 0-1023 are reserved
             Ok(Self(port))
         } else {
             Err(ConfigError::PortInReservedRange(port))
         }
     }

    pub fn get(self) -&gt; u16 {
         self.0
     }
 }</code></pre></div><p>Let&#8217;s also define an error type for configuration errors:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;8de146a1-a498-48df-a2ad-db791d803e8d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">#[derive(Debug, Clone, PartialEq, Eq)]
 pub enum ConfigError {
     MustBePositive(String),
     PortInReservedRange(u16),
 }

impl std::fmt::Display for ConfigError {
     fn fmt(&amp;self, f: &amp;mut std::fmt::Formatter&lt;'_&gt;) -&gt; std::fmt::Result {
         match self {
             ConfigError::MustBePositive(field) =&gt; write!(f, "{} must be positive", field),
             ConfigError::PortInReservedRange(port) =&gt; write!(f, "Port {} is in reserved range", port),
         }
     }
 }

impl std::error::Error for ConfigError {}</code></pre></div><p>Samsa&#8217;s configuration system now uses types that enforce validity at construction time. Once created, these types guarantee their validity throughout the system. This moves all validation to the boundary of our system and eliminates defensive programming in business logic.</p><p>By validating once during construction, we ensure that these types can only ever hold valid values. Rust&#8217;s ownership system guarantees that once created, these values cannot be modified to become invalid.</p><h3>Builder pattern with validation</h3><p>The Builder pattern pairs naturally with validated types. While the wrapper types we defined validate individual values, the Builder pattern coordinates multiple validated values into a complete, consistent configuration. We&#8217;ll create a <code>BrokerConfig</code> that can only be constructed through a builder, ensuring that all validation happens during construction.</p><p>First, let&#8217;s define the configuration struct with validated field types:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;2b997bed-8302-4470-a970-bd71b9793611&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">#[derive(Debug, Clone)]
pub struct BrokerConfig {
    max_connections: PositiveU32,
    buffer_size: NonZeroUsize,
    port: NetworkPort,
}

impl BrokerConfig {
    pub fn builder() -&gt; BrokerConfigBuilder {
        BrokerConfigBuilder::new()
    }

    // Getters that never need validation
    pub fn max_connections(&amp;self) -&gt; u32 {
        self.max_connections.get()
    }

    pub fn buffer_size(&amp;self) -&gt; usize {
        self.buffer_size.get()
    }

    pub fn port(&amp;self) -&gt; u16 {
        self.port.get()
    }
}</code></pre></div><p>The <code>BrokerConfig</code> struct uses validated types for all fields. <code>PositiveU32</code> and <code>NetworkPort</code> ensure that invalid values cannot be stored. The getters return primitive values, which is safe because we know the wrapped values are valid. Notice that all fields are private and accessed through getters. This prevents external code from modifying the configuration after validation.</p><p>Now, we need a way for users to construct a <code>BrokerConfig</code>. Since the struct has private fields with validated types, we can&#8217;t allow direct construction. Users would need access to <code>PositiveU32::new()</code> and <code>NetworkPort::new()</code>, and they&#8217;d have to handle validation errors themselves. Instead, we&#8217;ll provide a builder that accepts raw primitive values, validates them internally, and produces a fully validated <code>BrokerConfig</code> on success.</p><p>The <code>BrokerConfigBuilder</code> struct holds <code>Option</code> fields for each configuration value. As users set values through the builder&#8217;s methods, each value is validated and wrapped in its corresponding type. The final <code>build()</code> method ensures that all required fields are present before constructing the <code>BrokerConfig</code>. Let&#8217;s implement this builder:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;4fc8cddf-65b2-4a42-9177-177eeaacfa42&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">pub struct BrokerConfigBuilder {
    max_connections: Option&lt;PositiveU32&gt;,
    buffer_size: Option&lt;NonZeroUsize&gt;,
    port: Option&lt;NetworkPort&gt;,
}

impl BrokerConfigBuilder {
    pub fn new() -&gt; Self {
        Self {
            max_connections: None,
            buffer_size: None,
            port: None,
        }
    }

    pub fn max_connections(mut self, value: u32) -&gt; Result&lt;Self, ConfigError&gt; {
        self.max_connections = Some(PositiveU32::new(value)?);
        Ok(self)
    }

    pub fn buffer_size(mut self, value: usize) -&gt; Result&lt;Self, ConfigError&gt; {
        self.buffer_size = Some(
            NonZeroUsize::new(value)
                .ok_or_else(|| ConfigError::MustBePositive("buffer_size".to_string()))?
        );
        Ok(self)
    }

    pub fn port(mut self, value: u16) -&gt; Result&lt;Self, ConfigError&gt; {
        self.port = Some(NetworkPort::new(value)?);
        Ok(self)
    }
}</code></pre></div><p>Each setter method validates its input before storing it. The <code>?</code> operator propagates validation errors immediately if validation fails. If validation succeeds, the method consumes <code>self</code> and returns it, enabling method chaining. The use of <code>Option</code> for each field lets us detect missing fields in the <code>build()</code> method.</p><p>The final step is building the configuration, which ensures that all required fields are present and results in a fully validated <code>BrokerConfig</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;45eb0b2e-081e-4bc4-beba-cf37c3606343&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">impl BrokerConfigBuilder {
    pub fn build(self) -&gt; Result&lt;BrokerConfig, ConfigError&gt; {
        Ok(BrokerConfig {
            max_connections: self.max_connections
                .ok_or_else(|| ConfigError::MustBePositive("max_connections not set".to_string()))?,
            buffer_size: self.buffer_size
                .ok_or_else(|| ConfigError::MustBePositive("buffer_size not set".to_string()))?,
            port: self.port
                .ok_or_else(|| ConfigError::PortInReservedRange(0))?,
        })
    }
}</code></pre></div><p>The <code>build()</code> method consumes the builder and returns <code>Result&lt;BrokerConfig, ConfigError&gt;</code>. If any required field is missing, we return an error. This ensures that you can&#8217;t construct an incomplete configuration.</p><p>The usage now looks like this:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;368ac940-25c3-4fd5-913a-4a11b3e941fb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">let config = BrokerConfig::builder()
    .max_connections(1000)?
    .buffer_size(4096)?
    .port(8080)?
    .build()?;</code></pre></div><p>The fluent API makes configuration construction readable and catches errors immediately. If any setter returns an error, the <code>?</code> operator short-circuits, and the invalid configuration is never created. The final <code>build()</code> call ensures that no required fields are missing.</p><p>The builder ensures that all configuration values are validated during construction. Once a <code>BrokerConfig</code> exists, all its values are guaranteed to be valid. This design makes it impossible to create invalid configurations, with validation happening once at the system boundary.</p><p>The <em>parse, don&#8217;t validate</em> principle, combined with Rust&#8217;s type system and ownership model, eliminates defensive programming throughout the system by moving all validation to well-defined boundaries.</p><p>So far, we&#8217;ve seen how to create distinct types (NewType) and ensure that those types only hold valid values (parse, don&#8217;t validate). But what about objects that change over time? In the next section, we&#8217;ll explore the TypeState pattern, which uses the type system to track and enforce valid state transitions.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div><hr></div><h2>The TypeState pattern</h2><p>The TypeState pattern has academic origins in research on program verification. It was formalized by Robert E. Strom and Shaula Yemini in their 1986 paper <em>TypeState: A Programming Language Concept for Enhancing Software Reliability</em> (which you can read at <a href="https://www.computer.org/csdl/journal/ts/1986/01/06312929/13rRUwIF6aQ"><span>https://www.computer.org/csdl/journal/ts/1986/01/06312929/13rRUwIF6aQ</span></a>). The core idea is to use types to represent different states in which an object can be, making invalid state transitions impossible to express.</p><p>While the concept has existed in type theory for decades, Rust&#8217;s unique combination of features makes it particularly practical to implement. Rust&#8217;s zero-cost abstractions mean the type-level state tracking compiles away entirely, leaving no runtime overhead. The ownership system ensures that state transitions consume the old state, preventing accidental reuse. And generic types with <code>PhantomData</code> provide a clean mechanism for encoding state in the type system.</p><p>Several other languages support similar patterns, most notably session types in research languages and builder patterns in languages such as Java, but Rust&#8217;s combination of performance, safety, and ergonomics makes the TypeState pattern particularly accessible for systems programming.</p><h3>The problem with runtime state management</h3><p>Let&#8217;s examine traditional runtime state management, where state is tracked at runtime:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;743b86b6-25ba-4f46-bd34-c3c56fced5da&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">#[derive(Debug, Clone, Copy, PartialEq)]
pub enum ConsumerState {
     Disconnected,
     Connected,
     Subscribed,
     Paused,
 }

pub struct Consumer {
     broker: Arc&lt;Broker&gt;,
     topic: TopicId,
     consumer_id: ConsumerId,
     state: ConsumerState,
 }

impl Consumer {
     pub fn connect(&amp;mut self) -&gt; Result&lt;(), String&gt; {
         // Runtime check required
         if self.state != ConsumerState::Disconnected {
             return Err("Consumer must be disconnected to connect".to_string());
         }
         self.state = ConsumerState::Connected;
         Ok(())
     }

    pub fn poll(&amp;mut self) -&gt; Result&lt;Option&lt;Message&gt;, String&gt; {
         // Runtime validation in every method
         if self.state != ConsumerState::Subscribed {
             return Err("Consumer must be subscribed to poll
                         messages".to_string());
         }
         Ok(None)
     }
 }</code></pre></div><p>Every method requires runtime validation to ensure a correct state, creating opportunities for errors and adding runtime overhead. The compiler can&#8217;t help us by catching state transition errors, which leads to runtime failures that could have been prevented at compile time.</p><h3>Implementing TypeState for the consumer lifecycle</h3><p>Let&#8217;s redesign the consumer using TypeState. We&#8217;ll create separate types for each state:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;1bfdc9c4-390b-4a41-84eb-908047e3c28d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">use std::marker::PhantomData;
use crate::types::*;

// State marker types organized in a module
pub mod states {
    #[derive(Debug)]
    pub struct Disconnected;

    #[derive(Debug)]
    pub struct Connected;

    #[derive(Debug)]
    pub struct Subscribed;

    #[derive(Debug)]
    pub struct Paused;
}

// Consumer parameterized by state
pub struct Consumer&lt;State&gt; {
    broker: Arc&lt;Broker&gt;,
    topic: Option&lt;TopicId&gt;,
    consumer_id: ConsumerId,
    connection_info: Option&lt;ConnectionInfo&gt;,
    state: PhantomData&lt;State&gt;,
}

/// Type aliases for consumer states
pub type DisconnectedConsumer = Consumer&lt;states::Disconnected&gt;;
pub type ConnectedConsumer = Consumer&lt;states::Connected&gt;;
pub type SubscribedConsumer = Consumer&lt;states::Subscribed&gt;;
pub type PausedConsumer = Consumer&lt;states::Paused&gt;;

#[derive(Debug, Clone)]
pub struct ConnectionInfo {
    pub broker_address: String,
    pub consumer_group: Option&lt;String&gt;,
}</code></pre></div><p>The state marker types (<code>Disconnected</code>, <code>Connected</code>, <code>Subscribed</code>, and <code>Paused</code>) are zero-sized. They exist only at compile time to track state in the type system. The <code>Consumer&lt;State&gt;</code> struct is generic over the state marker, so <code>Consumer&lt;states::Disconnected&gt;</code> and <code>Consumer&lt;states::Connected&gt;</code> are different types even though they contain the same data fields. The type aliases provide convenient shorthand.</p><p>The <code>PhantomData&lt;State&gt;</code> field tells Rust about the <code>state</code> parameter without storing anything at runtime. This is the key to TypeState: state tracked in types with zero runtime cost.</p><p>Let&#8217;s implement the disconnected state:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;0cd0737d-91a1-4313-a2ca-2d242d6363b8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">impl Consumer&lt;states::Disconnected&gt; {
    pub fn new(consumer_id: ConsumerId, broker: Arc&lt;Broker&gt;) -&gt; Self {
        Self {
            broker,
            topic: None,
            consumer_id,
            connection_info: None,
            state: PhantomData,
        }
    }

    pub fn connect(
        self,
        connection_info: ConnectionInfo
    ) -&gt; Result&lt;Consumer&lt;states::Connected&gt;, SamsaError&gt; {
        if connection_info.broker_address.is_empty() {
            return Err(SamsaError::connection("Empty broker address"));
        }

        Ok(Consumer {
            broker: self.broker,
            topic: self.topic,
            consumer_id: self.consumer_id,
            connection_info: Some(connection_info),
            state: PhantomData,
        })
    }
}</code></pre></div><p>The <code>new</code> constructor creates a <code>Consumer&lt;states::Disconnected&gt;</code>. The <code>connect</code> method consumes the disconnected consumer (taking ownership with <code>self</code>) and returns either a <code>Consumer&lt;states::Connected&gt;</code> or an error along with the original consumer. This ownership transfer prevents reuse. Once you call <code>connect</code>, the disconnected consumer is gone.</p><p>Now, let&#8217;s implement the connected state:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;40c10044-f336-4c9a-822c-9b9af376dce6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">impl Consumer&lt;states::Connected&gt; {
    pub fn subscribe(
        mut self,
        topic: TopicId
    ) -&gt; Result&lt;Consumer&lt;states::Subscribed&gt;, SamsaError&gt; {
        self.topic = Some(topic);

        Ok(Consumer {
            broker: self.broker,
            topic: self.topic,
            consumer_id: self.consumer_id,
            connection_info: self.connection_info,
            state: PhantomData,
        })
    }

    pub fn disconnect(self) -&gt; Consumer&lt;states::Disconnected&gt; {
        Consumer {
            broker: self.broker,
            topic: None,
            consumer_id: self.consumer_id,
            connection_info: None,
            state: PhantomData,
        }
    }
}</code></pre></div><p>Each state transition follows the same pattern: consume the current state, perform an operation, and return the new state or an error. The connected consumer can subscribe to a topic or disconnect. Each transition is a separate <code>impl</code> block, making it impossible to call methods from the wrong state.</p><p>Finally, the <code>Subscribed</code> state is where consumers can actually receive messages, while the <code>Paused</code> state can only resume:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;ea81f63a-0d53-4954-a0ad-bb1105186063&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">impl Consumer&lt;states::Subscribed&gt; {
    pub fn receive(&amp;self) -&gt; Option&lt;Event&gt; {
        // Only subscribed consumers can receive - guaranteed by types!
        None
    }

    pub fn pause(self) -&gt; Consumer&lt;states::Paused&gt; {
        Consumer {
            broker: self.broker,
            topic: self.topic,
            consumer_id: self.consumer_id,
            connection_info: self.connection_info,
            state: PhantomData,
        }
    }

    pub fn unsubscribe(mut self) -&gt; Consumer&lt;states::Connected&gt; {
        self.topic = None;
        Consumer {
            broker: self.broker,
            topic: self.topic,
            consumer_id: self.consumer_id,
            connection_info: self.connection_info,
            state: PhantomData,
        }
    }
}

impl Consumer&lt;states::Paused&gt; {
    pub fn resume(self) -&gt; Consumer&lt;states::Subscribed&gt; {
        Consumer {
            broker: self.broker,
            topic: self.topic,
            consumer_id: self.consumer_id,
            connection_info: self.connection_info,
            state: PhantomData,
        }
    }
}</code></pre></div><p>Only <code>Consumer&lt;states::Subscribed&gt;</code> has a <code>receive</code> method. You literally cannot call <code>receive()</code> on consumers in other states. The method doesn&#8217;t exist for those types. The <code>Paused</code> state shows that TypeState can model bidirectional transitions: <code>pause()</code> moves to <code>Paused</code>, and <code>resume()</code> returns to <code>Subscribed</code>.</p><p>This TypeState implementation makes invalid operations impossible to express. The compiler enforces the correct state machine at compile time with no runtime overhead.</p><p>TypeState&#8217;s implementation in Rust uses <code>PhantomData</code> to carry type-level information that exists only at compile time, employs generic types to parameterize over state, and utilizes the ownership system to ensure that state transitions consume the previous state. The result is verification of state machines at compile time with zero runtime cost.</p><h3>Usage examples</h3><p>Here&#8217;s how TypeState improves the developer experience:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;cc52f6b9-6530-4119-82c7-409c07a2fb8a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">fn correct_usage() -&gt; Result&lt;(), SamsaError&gt; {
    let broker = Arc::new(Broker::new());
    let consumer_id = ConsumerId::new("test-consumer")?;
    let topic = TopicId::new("events")?;

    // Create disconnected consumer
    let consumer = Consumer::new(consumer_id, broker);

    // Connect with connection info
    let connection_info = ConnectionInfo::new(
        "localhost:9092".to_string(), None
    );
    let connected = consumer.connect(connection_info)?;

    // Subscribe to topic
    let subscribed = connected.subscribe(topic)?;

    // Can pause and resume
    let paused = subscribed.pause();
    let subscribed = paused.resume();

   // Now we can receive - guaranteed to be in correct state
    if let Some(event) = subscribed.receive() {
        println!("Received: {:?}", event);
    }

    // Unsubscribe and disconnect
    let connected = subscribed.unsubscribe();
    let _disconnected = connected.disconnect();

    Ok(())
}</code></pre></div><p>This code demonstrates the correct sequence: create a disconnected consumer, connect it, subscribe, then receive &#8211; pausing and unpausing, unsubscribing, and disconnecting work the same way. Each state transition is explicit in the code, and the types change at each step. The error handling uses a tuple to return the original consumer on failure, allowing retry logic.</p><p>Invalid usage won&#8217;t compile:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;5dd1cbe0-bb2e-495a-8efa-c134db723838&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">fn invalid_usage() {
    let broker = Arc::new(Broker::new());
    let consumer_id = ConsumerId::new("test").unwrap();

    let consumer = Consumer::new(consumer_id, broker);

    // These lines won't compile:
    // consumer.receive(); // Error: not available on Disconnected
    // consumer.subscribe(topic); // Error: not available on Disconnected
}</code></pre></div><p>The compiler errors shown in the comments aren&#8217;t runtime errors; they&#8217;re compile-time errors. You literally cannot write code that calls <code>receive()</code> on a disconnected consumer because the method doesn&#8217;t exist for that type. This is the power of TypeState: invalid states are unrepresentable.</p><p>The TypeState pattern transforms potential runtime errors into compile-time guarantees, demonstrating how Rust&#8217;s type system and ownership model enable practical implementation of ideas from formal methods and programming language research.</p><p>TypeState controls how individual objects transition through states. But sometimes, we need to control which types can participate in a system at all. In the next section, we&#8217;ll see how sealed traits let us define closed sets of implementations, providing API stability and safety guarantees.</p><div><hr></div><h1>Sealed traits</h1><p>Many object-oriented languages introduce the concept of sealed or final types, which are types that cannot extend beyond a defined scope. Java has final classes, Kotlin has sealed classes, and Scala has sealed traits. The pattern addresses a real problem in API design: sometimes, allowing arbitrary external implementations of an interface creates more problems than it solves.</p><p>Rust doesn&#8217;t have a built-in <code>sealed</code> keyword, but the module system provides a way to achieve the same effect. Following its appearance in various standard library crates and its documentation in API design discussions, the technique gained widespread recognition in the Rust community. While not unique to Rust, the pattern fits naturally with Rust&#8217;s privacy model and module system.</p><p>The sealed traits pattern uses Rust&#8217;s module privacy to create traits that are visible for use but restricted for implementation. This provides library authors with control over trait implementations while maintaining a public interface.</p><h3>The problem with open traits</h3><p>Let&#8217;s understand why we might want to seal traits. Because this is an open trait, anyone can implement it:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;86b67500-c3db-451e-b5c9-805c2c62b4f4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">pub trait MessageSchema {
     type Data;
     type Error: std::error::Error;

    fn parse(data: &amp;[u8]) -&gt; Result&lt;Self::Data, Self::Error&gt;;
     fn serialize(data: &amp;Self::Data) -&gt; Vec&lt;u8&gt;;
     fn validate(data: &amp;Self::Data) -&gt; bool;
 }</code></pre></div><p>Here, we&#8217;ve defined a very straightforward trait defining the signatures of three methods, with associated types for the data payload and the error type. There is nothing unique about this trait, but what we should pay attention to is that this is a public trait, so there are no restrictions on who can implement it or where the implementation happens.</p><p>External crates could implement unsafe versions:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;6be4bd86-9567-4c26-989d-96d5967fb82a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">// struct UnsafeSchema;
// impl MessageSchema for UnsafeSchema {
//     fn parse(_data: &amp;[u8]) -&gt; Result&lt;Self::Data, Self::Error&gt; {
//         unsafe { /* potentially unsafe implementation */ }
//     }
// }</code></pre></div><p>The <code>MessageSchema</code> trait defines an interface that any type can implement. While this flexibility is often desirable, it means we can&#8217;t make guarantees about all implementations. External code could implement the trait incorrectly or in ways that violate our assumptions.</p><p>When external crates can implement <code>MessageSchema</code>, we lose control over quality guarantees and make it difficult to evolve the trait interface without breaking external implementations.</p><p>External implementations might do the following:</p><div class="callout-block" data-callout="true"><ul><li><p>Skip validation or implement it incorrectly</p></li><li><p>Use <code>unsafe</code> code in unexpected ways</p></li><li><p>Depend on undocumented behavior that we later change</p></li></ul></div><p>Additionally, evolving the trait interface becomes difficult. Adding a new required method would break all external implementations, forcing us to use default implementations even when that&#8217;s not ideal.</p><h2>Implementing the sealed traits pattern</h2><p>Let&#8217;s implement the sealed traits pattern for our message schema system, using a private module containing the sealing trait:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;0a693c97-a824-452c-9cd7-35faf8451f9b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">mod private {
    pub trait Sealed {}
}

// Public trait that extends the sealed trait
pub trait MessageSchema: private::Sealed {
    type Message: Clone + std::fmt::Debug;

    fn serialize(message: &amp;Self::Message) -&gt; Vec&lt;u8&gt;;
    fn deserialize(bytes: &amp;[u8]) -&gt; Result&lt;Self::Message, SchemaError&gt;;
    fn schema_id() -&gt; &amp;'static str;
    fn validate(message: &amp;Self::Message) -&gt; Result&lt;(), ValidationError&gt;;
}</code></pre></div><p>The key to the sealed traits pattern is the private module. The <code>Sealed</code> trait is public within the module, but the module itself is private. External crates can see <code>MessageSchema</code> (which extends <code>private::Sealed</code>) but cannot implement <code>Sealed</code>, which means they cannot implement <code>MessageSchema</code>.</p><p>This gives us complete control over which types can implement our schema trait.</p><p>Let&#8217;s implement a JSON schema:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;98571e49-f5ad-4e97-ba83-983848f055a3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">pub struct JsonSchema;

impl private::Sealed for JsonSchema {}

impl MessageSchema for JsonSchema {
    type Message = serde_json::Value;

    fn serialize(message: &amp;Self::Message) -&gt; Vec&lt;u8&gt; {
        serde_json::to_vec(message).unwrap_or_default()
    }

    fn deserialize(bytes: &amp;[u8]) -&gt; Result&lt;Self::Message, SchemaError&gt; {
        serde_json::from_slice(bytes)
            .map_err(|e| SchemaError::DeserializationFailed(e.to_string()))
    }

    fn schema_id() -&gt; &amp;'static str {
        "json_v1"
    }

    fn validate(message: &amp;Self::Message) -&gt; Result&lt;(), ValidationError&gt; {
        if message.is_null() {
            return Err(ValidationError::InvalidValue(
                "Message cannot be null".to_string()
            ));
        }
        Ok(())
    }
}</code></pre></div><p><code>JsonSchema</code> implements both <code>private::Sealed</code> (which we can do because we&#8217;re in the same module) and <code>MessageSchema</code>. The implementation uses <code>serde_json</code> for parsing and serialization and validates that JSON values aren&#8217;t <code>null</code>.</p><p>Let&#8217;s add a text schema for plain <code>String</code> data:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;fac615eb-fe7a-43de-ab11-1c3db6fd5876&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">pub struct TextSchema;

impl private::Sealed for TextSchema {}

impl MessageSchema for TextSchema {
    type Message = String;

    fn serialize(message: &amp;Self::Message) -&gt; Vec&lt;u8&gt; {
        message.as_bytes().to_vec()
    }

    fn deserialize(bytes: &amp;[u8]) -&gt; Result&lt;Self::Message, SchemaError&gt; {
        String::from_utf8(bytes.to_vec())
            .map_err(|e| SchemaError::DeserializationFailed(e.to_string()))
    }

    fn schema_id() -&gt; &amp;'static str {
        "text_v1"
    }

    fn validate(message: &amp;Self::Message) -&gt; Result&lt;(), ValidationError&gt; {
        if message.is_empty() {
            return Err(ValidationError::FieldRequired(
                "message content".to_string()
            ));
        }
        Ok(())
    }
}</code></pre></div><p><code>TextSchema</code> handles plain UTF-8 text. It deserializes byte slices into strings and validates that they aren&#8217;t empty. Both schema implementations are within our module, so both can implement <code>private::Sealed</code>.</p><p>External crates can use <code>JsonSchema</code> and <code>TextSchema</code>, but cannot create their own schema implementations; the sealed traits pattern prevents it.</p><p>The sealed traits pattern leverages Rust&#8217;s module privacy: <code>private::Sealed</code> is public within the module, but the <code>sealed</code> module itself is private. External crates can see and use <code>MessageSchema</code>, but cannot implement <code>Sealed</code>, which means they cannot implement <code>MessageSchema</code>.</p><h3>Type-safe messages with sealed schemas</h3><p>Now, we can combine sealed schemas with the type-safe message patterns we&#8217;ve developed. <code>TypedMessage</code> will be generic over the schema type, ensuring that messages and their schemas are always compatible. The sealed traits pattern prevents external code from creating incompatible schema implementations:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;1ced5220-f28f-4543-bed6-3452d2f1f4eb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">use std::marker::PhantomData;
use crate::schema::MessageSchema;
use crate::types::*;

#[derive(Debug, Clone)]
pub struct TypedMessage&lt;S: MessageSchema&gt; {
    pub id: MessageId,
    pub content: S::Message,
    pub schema_type: PhantomData&lt;S&gt;,
}

impl&lt;S: MessageSchema&gt; TypedMessage&lt;S&gt; {
    pub fn new(id: MessageId, content: S::Message) -&gt; Result&lt;Self, ValidationError&gt; {
        S::validate(&amp;content)?;

        Ok(TypedMessage {
            id,
            content,
            schema_type: PhantomData,
        })
    }

    pub fn to_bytes(&amp;self) -&gt; Vec&lt;u8&gt; {
        S::serialize(&amp;self.content)
    }

    pub fn schema_id(&amp;self) -&gt; &amp;'static str {
        S::schema_id()
    }
}</code></pre></div><p><code>TypedMessage</code> is generic over the <code>S</code> schema type, where <code>S</code> implements <code>MessageSchema</code>. The <code>content</code> field has the <code>S::Message</code> type, which is the associated type we defined in the message schema trait for the type of our message payload. When we wrote our trait implementations, we defined concrete types for <code>MessageSchema::Message: serde_json::Value</code> for <code>TypedMessage&lt;JsonSchema&gt;</code> and <code>String</code> for <code>TypedMessage&lt;TextSchema&gt;</code>. The <code>schema_type</code> field uses <code>PhantomData</code> to track the schema type at compile time without storing anything at runtime. The <code>new</code> constructor validates the content using the schema before creating the message.</p><p>We can extend <code>TypedMessage</code> with additional methods for parsing raw data. The following shows how you might add a <code>parse</code> method (this extension is not in the base repository, but demonstrates the pattern):</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;rust&quot;,&quot;nodeId&quot;:&quot;49bf6db5-608c-481d-992b-cd16144a1f16&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-rust">impl&lt;S: MessageSchema&gt; TypedMessage&lt;S&gt; {
    pub fn parse(
        id: MessageId,
        raw_data: &amp;[u8],
    ) -&gt; Result&lt;Self, S::Error&gt; {
        let content = S::deserialize(raw_data)?;
        Ok(Self {
            id,
            content,
            schema_type: PhantomData,
        })
    }

    pub fn serialize(&amp;self) -&gt; Vec&lt;u8&gt; {
        S::serialize(&amp;self.content)
    }
}

pub type JsonMessage = TypedMessage&lt;JsonSchema&gt;;
pub type TextMessage = TypedMessage&lt;TextSchema&gt;;</code></pre></div><p>The extended <code>parse</code> method takes raw bytes and uses the schema&#8217;s deserializer to construct typed content. If parsing fails, the error type is <code>S::Error</code>, which is <code>serde_json::Error</code> for JSON or <code>std::str::Utf8Error</code> for text. The <code>serialize</code> method converts the content back into bytes. This demonstrates how you can build on the sealed trait foundation to add richer functionality.</p><p>The type aliases provide convenient names: <code>JsonMessage</code> and <code>TextMessage</code> are clearer than writing <code>TypedMessage&lt;JsonSchema&gt;</code> everywhere. These messages are both type-safe (wrong schema = compile error) and sealed (can&#8217;t add new schemas externally).</p><p>Thus, the sealed traits pattern provides controlled extensibility. We maintain the ability to evolve our trait interface while providing compile-time guarantees about which implementations exist.</p><div><hr></div><h4>Summary</h4><p>In this deep-dive, we explored four patterns that leverage Rust&#8217;s type system, each with established origins in programming language research and practice.</p><p>The NewType pattern comes from functional programming languages, particularly Haskell, where it has been a formal language feature since the 1990s. Rust&#8217;s contribution is making the pattern zero-cost: type safety is enforced at compile time and optimized away entirely at runtime, making it practical for systems programming, where performance matters.</p><p><em>Parse, don&#8217;t validate</em> was articulated by Alexis King in 2019, building on ideas from functional programming and type theory. Rust&#8217;s type system, ownership model, and <code>Result</code> type make this principle natural to apply, turning boundary validation into type-level guarantees that persist throughout the program.</p><p>The TypeState pattern originated in formal methods research, specifically in Strom and Yemini&#8217;s 1986 work on program verification. Rust&#8217;s generic types with <code>PhantomData</code>, zero-cost abstractions, and ownership system make this academic concept practical for everyday systems programming, providing compile-time state machine verification with no runtime overhead.</p><p>Sealed traits draw from similar concepts in Java (<code>final</code> classes), Kotlin (<code>sealed</code> classes), and Scala (<code>sealed</code> traits). Rust implements the pattern through its module privacy system, providing a natural fit with the language&#8217;s existing visibility rules.</p><p>What these patterns demonstrate is Rust&#8217;s ability to take established ideas from type theory, functional programming, and formal methods and make them practical for systems programming. The combination of strong static typing, zero-cost abstractions, and memory safety without garbage collection creates opportunities to apply patterns that were previously either too expensive or impractical in systems languages.</p><p>Our enhanced Samsa system now leverages decades of programming language research while maintaining the performance characteristics necessary for systems programming. This synthesis, applying proven patterns in a performant, safe environment, represents one of Rust&#8217;s key contributions to the programming language space.</p><p>The book&#8217;s next chapter explores patterns from functional programming, including function composition pipelines, generics as type classes, pattern matching techniques, and closures as architectural components.</p><div><hr></div><p><strong>More from Evan Williams</strong></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;a4a65f3b-09c1-4f79-a131-2844a3b93e31&quot;,&quot;caption&quot;:&quot;Eval Driven Development for Engineers&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #47: Evan Williams on Why Experienced Developers Have the Hardest Time Learning Rust&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-05-14T16:42:52.048Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/43a42d88-ec70-4d7f-8213-85796343b4f5_677x337.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.net/p/deep-engineering-47-why-experienced-developers-hardest-time-learning-rust&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:197666671,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;5badedc9-2f70-48fd-8fcb-2a9def45cb0e&quot;,&quot;caption&quot;:&quot;Evan Williams has been writing software for more than 40 years, across every layer of the stack and more programming languages than most engineers will encounter in a career. His book, Design Patterns and Best Practices in Rust, published by Packt, is not a pattern catalogue. It is an argument for a different way of thinking about code entirely, aimed s&#8230;&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Design Patterns, Ownership Models, and Building Resilient Systems in Rust with Evan Williams&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-05-13T18:07:20.569Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/youtube/w_728,c_limit/-ElpmT7DCX4&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.net/p/design-patterns-ownership-models-resilient-systems-rust-evan-williams&quot;,&quot;section_name&quot;:&quot;Interviews&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:197553608,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:2,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;d27e9897-e080-46a2-83f4-ac2a9681da80&quot;,&quot;caption&quot;:&quot;Rust, who would have thought, has ranked as the most loved programming language in the Stack Overflow developer survey for nine consecutive years. Honestly, I must admit this is an unusual kind of statistic because it measures not just adoption but retention. The engineers who use Rust want to keep using it, and that pattern has only deepened even as th&#8230;&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Rust Is Hard for the Engineers with the Most Experience&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:11407185,&quot;name&quot;:&quot;Francesco Ciulla&quot;,&quot;bio&quot;:&quot;Developer Advocate at @dailydotdev\n&#183; Docker Captain &#128051;\n&#183; Public Speaker\n&#183; Building a 1 Million Community 22%&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b8c30606-10ba-4c87-a89b-af2f9dc27a01_400x400.jpeg&quot;,&quot;is_guest&quot;:true,&quot;bestseller_tier&quot;:null,&quot;primaryPublicationSubscribeUrl&quot;:&quot;https://francescociulla.substack.com/subscribe?&quot;,&quot;primaryPublicationUrl&quot;:&quot;https://francescociulla.substack.com&quot;,&quot;primaryPublicationName&quot;:&quot;Francesco's Newsletter&quot;,&quot;primaryPublicationId&quot;:1410908}],&quot;post_date&quot;:&quot;2026-05-18T16:07:41.936Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f3cfa859-6cd7-4d3e-ab4b-31c7b0cd00c5_1440x660.webp&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.net/p/rust-is-hard-for-the-engineers-with-the-most-experience&quot;,&quot;section_name&quot;:&quot;Engineering Leadership&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:198261752,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div class="pullquote"><p><em>This deep-dive is Chapter 10 of <a href="https://www.packtpub.com/en-us/product/design-patterns-and-best-practices-in-rust-9781836209478">Design Patterns and Best Practices in Rust</a> by <a href="https://www.linkedin.com/in/evan-williams-1512092">Evan Williams</a>, published by <a href="https://www.packtpub.com/">Packt</a> and reproduced here in full with the publisher&#8217;s permission for knowledge sharing with the Deep Engineering community. The book charts a transformation across three hands-on projects, a deliberately broken calculator that shows what goes wrong when you write Java or C++ in Rust syntax, a rebuild that adapts the classic Gang of Four patterns to the ownership model, and Samsa, the publish/subscribe microservice these pages extend with patterns unique to the language. All rights remain with <a href="https://www.packtpub.com/">Packt Publishing</a>, and this content may not be reproduced, redistributed, or remixed in any form without the publisher&#8217;s written consent. We also sat down with Evan Williams to talk about the thinking behind these patterns, including why he finds typestate almost impossible to stop talking about, <a href="https://deepengineering.net/p/design-patterns-ownership-models-resilient-systems-rust-evan-williams">read the conversation here</a>. You can get the full book <a href="https://www.packtpub.com/en-us/product/design-patterns-and-best-practices-in-rust-9781836209478">here</a>.</em></p><div><hr></div></div>]]></content:encoded></item><item><title><![CDATA[Core Architectural Patterns for LLM System Design ]]></title><description><![CDATA[How to design for resilience, latency, cost, and trust when your newest dependency is slower and less predictable than anything else in your stack]]></description><link>https://deepengineering.net/p/core-architectural-patterns-for-llm-system-design</link><guid isPermaLink="false">https://deepengineering.net/p/core-architectural-patterns-for-llm-system-design</guid><dc:creator><![CDATA[Sampriti Mitra]]></dc:creator><pubDate>Thu, 09 Jul 2026 09:25:27 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f88846d4-0cd6-4025-8548-73aae36c03ea_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Integrating LLMs into a production system introduces a new class of dependency: one that is non-deterministic, high-latency, and carries a high, variable operational cost. The fundamentals of LLM integration, tokens, embeddings, and the basic idea of retrieval-augmented generation (RAG), are only the starting point.</p><p>As experienced engineers, we already know how to build reliable systems. This deep-dive isn&#8217;t about reinventing those principles, but about adapting them for the unique, messy challenges LLMs throw at us. Consider this an architect&#8217;s playbook that outlines the new patterns needed to meet these challenges.</p><blockquote><p><em>This <strong>deep-dive</strong> is Chapter 2 of <a href="https://www.packtpub.com/en-us/product/system-design-for-the-llm-era-9781807789923">System Design for the LLM Era</a> by <a href="https://in.linkedin.com/in/sampritimitra">Sampriti Mitra</a>, shared with permission from Packt, with all rights remaining with the publisher and no reproduction or redistribution without written consent. You can get the full book <a href="https://www.packtpub.com/en-us/product/system-design-for-the-llm-era-9781807789923">here</a>.</em></p></blockquote><div class="callout-block" data-callout="true"><p>In this deep-dive we&#8217;ll be looking at the following topics:</p><ul><li><p>Designing for resilience and reliability</p></li><li><p>Designing for low latency</p></li><li><p>Designing for cost optimization</p></li><li><p>Designing for grounding and data management</p></li><li><p>Designing for testability and observability</p></li><li><p>Designing for security and trust</p></li><li><p>Engineering for production</p></li><li><p>Training with test data</p></li><li><p>Respecting user privacy</p></li></ul></div><div><hr></div><p><strong>Featured workshop: <a href="https://www.eventbrite.co.uk/e/loop-engineering-for-ai-agents-tickets-1992373400474?aff=deepeng">Loop Engineering for AI Agents</a></strong></p><div class="pullquote"><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.eventbrite.co.uk/e/loop-engineering-for-ai-agents-tickets-1992373400474?aff=deepeng" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!87in!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff593ef62-ef12-4da5-ad48-bcb0d2806b40_900x300.png 424w, https://substackcdn.com/image/fetch/$s_!87in!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff593ef62-ef12-4da5-ad48-bcb0d2806b40_900x300.png 848w, https://substackcdn.com/image/fetch/$s_!87in!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff593ef62-ef12-4da5-ad48-bcb0d2806b40_900x300.png 1272w, https://substackcdn.com/image/fetch/$s_!87in!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff593ef62-ef12-4da5-ad48-bcb0d2806b40_900x300.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!87in!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff593ef62-ef12-4da5-ad48-bcb0d2806b40_900x300.png" width="900" height="300" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f593ef62-ef12-4da5-ad48-bcb0d2806b40_900x300.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:300,&quot;width&quot;:900,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.eventbrite.co.uk/e/loop-engineering-for-ai-agents-tickets-1992373400474?aff=deepeng&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!87in!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff593ef62-ef12-4da5-ad48-bcb0d2806b40_900x300.png 424w, https://substackcdn.com/image/fetch/$s_!87in!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff593ef62-ef12-4da5-ad48-bcb0d2806b40_900x300.png 848w, https://substackcdn.com/image/fetch/$s_!87in!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff593ef62-ef12-4da5-ad48-bcb0d2806b40_900x300.png 1272w, https://substackcdn.com/image/fetch/$s_!87in!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff593ef62-ef12-4da5-ad48-bcb0d2806b40_900x300.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A 4-hour hands-on workshop on <strong>Loop Engineering</strong>, designing agent workflows with Claude Code, SDD, and MCP that verify their own work and stay observable in production.</p><p><em>Use<strong> DEEPENG40</strong> for 40% discount.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.co.uk/e/loop-engineering-for-ai-agents-tickets-1992373400474?aff=deepeng&quot;,&quot;text&quot;:&quot;Register here&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.eventbrite.co.uk/e/loop-engineering-for-ai-agents-tickets-1992373400474?aff=deepeng"><span>Register here</span></a></p></div><h2><span>Designing for resilience and reliability</span></h2><p><span>Our system&#8217;s stability is now tied to an external API that is slower, more expensive, and less predictable than any database or microservice calls. The primary goal is to decouple the application&#8217;s health from the provider&#8217;s health. Two types of pattern are especially valuable in this regard: the </span><strong><span>GenAI service pattern</span></strong><span> (or </span><strong><span>LLM gateway pattern</span></strong><span>) and the </span><strong><span>circuit breaker pattern</span></strong><span>.</span></p><h3><span>Pattern: the GenAI service or LLM gateway</span></h3><p><span>When we first start building with LLMs, our instinct is to treat them like any other third-party API. We install the SDK, generate an API key, and make the call directly from our application code.</span></p><p><span>It feels fast. It feels efficient. We write a Python function, import </span><span data-color="#e06666" style="color: rgb(224, 102, 102);">openai</span><span>, and we are shipping features in minutes.</span></p><p><span>The architecture looks like this:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!unGq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F442e332f-4895-4c1e-a8b2-e453818752e3_959x423.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!unGq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F442e332f-4895-4c1e-a8b2-e453818752e3_959x423.png 424w, https://substackcdn.com/image/fetch/$s_!unGq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F442e332f-4895-4c1e-a8b2-e453818752e3_959x423.png 848w, https://substackcdn.com/image/fetch/$s_!unGq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F442e332f-4895-4c1e-a8b2-e453818752e3_959x423.png 1272w, https://substackcdn.com/image/fetch/$s_!unGq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F442e332f-4895-4c1e-a8b2-e453818752e3_959x423.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!unGq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F442e332f-4895-4c1e-a8b2-e453818752e3_959x423.png" width="959" height="423" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/442e332f-4895-4c1e-a8b2-e453818752e3_959x423.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:423,&quot;width&quot;:959,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.1: Naive model-calling logic&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.1: Naive model-calling logic" title="Figure 2.1: Naive model-calling logic" srcset="https://substackcdn.com/image/fetch/$s_!unGq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F442e332f-4895-4c1e-a8b2-e453818752e3_959x423.png 424w, https://substackcdn.com/image/fetch/$s_!unGq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F442e332f-4895-4c1e-a8b2-e453818752e3_959x423.png 848w, https://substackcdn.com/image/fetch/$s_!unGq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F442e332f-4895-4c1e-a8b2-e453818752e3_959x423.png 1272w, https://substackcdn.com/image/fetch/$s_!unGq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F442e332f-4895-4c1e-a8b2-e453818752e3_959x423.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span>Figure 2.1: Naive model-calling logic</span></figcaption></figure></div><p><span>While this works for a weekend hackathon, it creates a high-coupling, low-cohesion architecture in production with the following issues:</span></p><div class="callout-block" data-callout="true"><ul><li><p>Vendor lock-in: If Service A is written using the OpenAI SDK, migrating to Anthropic requires rewriting the entire code block.</p></li><li><p>Inconsistent reliability: Service A might have excellent retry logic, while Service B crashes on the first timeout. There is no standard.</p></li><li><p>Observability black holes: You have no central place to see how much you are spending. You have to log into three different developer consoles to tally up the bill.</p></li><li><p>Security risks: API keys are scattered across multiple environment variables in multiple services, increasing the surface area for leaks.</p></li></ul></div><p><span>Stating the problem more formally now:</span></p><p><span>Problem: Our services should not be calling OpenAI, Anthropic, or Google directly. This creates high-coupling, high maintenance problems.</span></p><p><span>Solution: To solve this, we borrow a pattern from traditional microservices: the API gateway. We stop treating LLMs as external vendors and start treating them as a unified internal resource.</span></p><p><span>Implement a single, centralized LLM gateway. This is a microservice that acts as the only entry point for all LLM calls. Our services talk only to this gateway using a single, unified API format. The gateway handles the messy details of talking to the outside world.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ITfe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F103c054b-e903-4484-a925-34533ef075ca_734x1600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ITfe!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F103c054b-e903-4484-a925-34533ef075ca_734x1600.png 424w, https://substackcdn.com/image/fetch/$s_!ITfe!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F103c054b-e903-4484-a925-34533ef075ca_734x1600.png 848w, https://substackcdn.com/image/fetch/$s_!ITfe!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F103c054b-e903-4484-a925-34533ef075ca_734x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!ITfe!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F103c054b-e903-4484-a925-34533ef075ca_734x1600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ITfe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F103c054b-e903-4484-a925-34533ef075ca_734x1600.png" width="734" height="1600" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/103c054b-e903-4484-a925-34533ef075ca_734x1600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1600,&quot;width&quot;:734,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.2: The LLM gateway pattern &#8211; decoupling internal services from external providers&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.2: The LLM gateway pattern &#8211; decoupling internal services from external providers" title="Figure 2.2: The LLM gateway pattern &#8211; decoupling internal services from external providers" srcset="https://substackcdn.com/image/fetch/$s_!ITfe!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F103c054b-e903-4484-a925-34533ef075ca_734x1600.png 424w, https://substackcdn.com/image/fetch/$s_!ITfe!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F103c054b-e903-4484-a925-34533ef075ca_734x1600.png 848w, https://substackcdn.com/image/fetch/$s_!ITfe!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F103c054b-e903-4484-a925-34533ef075ca_734x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!ITfe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F103c054b-e903-4484-a925-34533ef075ca_734x1600.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2.2: The LLM gateway pattern &#8211; decoupling internal services from external providers</figcaption></figure></div><p><span>This pattern brings some significant benefits:</span></p><div class="callout-block" data-callout="true"><ul><li><p>Abstraction: Can switch models (e.g. GPT-5 for Claude 3 Opus) with a config change, not a re-deployment</p></li><li><p>Centralized control: All other resilience, cost, and monitoring patterns are implemented in this one place</p></li><li><p>Authentication: Manages all authentications in one service</p></li><li><p>Fallbacks and reliability: If OpenAI goes down, the gateway can automatically retry the request with Anthropic. The upstream service never even knows there was an outage</p></li></ul></div><h3><span>Pattern: circuit breakers with tiered fallbacks</span></h3><p><span>An LLM provider might be slow, down, or just returning bad data. Simply retrying a failed call (like you would for a 503 on an external service) is often the wrong move.</span></p><p><span>The solution is to combine the circuit breaker pattern with </span><strong><span>tiered fallbacks</span></strong><span> following a three-stage model:</span></p><ol><li><p><span>Monitor: The GenAI Service monitors the health (latency, error rate) of each model provider.</span></p></li><li><p><span>Trip: If a primary model (e.g. GPT-5) exceeds a failure threshold, the circuit opens.</span></p></li><li><p><span>Reroute: All subsequent requests are immediately and automatically rerouted to a backup model.</span></p></li></ol><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XXsV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc79e98b-2ce4-4676-996b-2d0dfbca4d24_907x591.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XXsV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc79e98b-2ce4-4676-996b-2d0dfbca4d24_907x591.png 424w, https://substackcdn.com/image/fetch/$s_!XXsV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc79e98b-2ce4-4676-996b-2d0dfbca4d24_907x591.png 848w, https://substackcdn.com/image/fetch/$s_!XXsV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc79e98b-2ce4-4676-996b-2d0dfbca4d24_907x591.png 1272w, https://substackcdn.com/image/fetch/$s_!XXsV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc79e98b-2ce4-4676-996b-2d0dfbca4d24_907x591.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XXsV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc79e98b-2ce4-4676-996b-2d0dfbca4d24_907x591.png" width="907" height="591" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cc79e98b-2ce4-4676-996b-2d0dfbca4d24_907x591.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:591,&quot;width&quot;:907,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.3: Tiered fallbacks with circuit breakers&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.3: Tiered fallbacks with circuit breakers" title="Figure 2.3: Tiered fallbacks with circuit breakers" srcset="https://substackcdn.com/image/fetch/$s_!XXsV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc79e98b-2ce4-4676-996b-2d0dfbca4d24_907x591.png 424w, https://substackcdn.com/image/fetch/$s_!XXsV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc79e98b-2ce4-4676-996b-2d0dfbca4d24_907x591.png 848w, https://substackcdn.com/image/fetch/$s_!XXsV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc79e98b-2ce4-4676-996b-2d0dfbca4d24_907x591.png 1272w, https://substackcdn.com/image/fetch/$s_!XXsV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc79e98b-2ce4-4676-996b-2d0dfbca4d24_907x591.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2.3: Tiered fallbacks with circuit breakers</figcaption></figure></div><p><span>For example, we might implement the following tiered fallback strategy:</span></p><p><span>Tier 1: GPT-5 (high-cost, high-reasoning)</span></p><p><span>Tier 2 (fallback): Claude 3 Haiku (medium-cost, fast)</span></p><p><span>Tier 3 (fallback): Llama 3 8B (locally hosted, free, less smart)</span></p><p><span>Tier 4 (final fallback): Return a cached good enough response or a graceful error message to the client UI: &#8216;Our AI assistant is at high capacity, please try again in a moment&#8217;</span></p><p><span>We cannot leave the circuit open forever. We need </span><strong><span>recovery logic</span></strong><span> to close the circuit that includes a half-open state:</span></p><p><span>Sleep: After the circuit trips, wait for a defined cooldown (e.g. 30 seconds).</span></p><p><span>Probe: Allow a single canary request to pass through to the primary provider.</span></p><p><span>Reset: If the canary succeeds, close the circuit and resume full traffic. If it fails, restart the cooldown.</span></p><p><span>For our </span><strong><span>retry strategy</span></strong><span> we have two options: interactive and asynchronous:</span></p><div class="callout-block" data-callout="true"><ul><li><p>Interactive/synchronous: Do not use aggressive <strong>exponential backoff</strong>. If a user is waiting, a 60-second retry is effectively downtime. Use capped backoff (start 500 ms, max 1 s) or fail fast to a fallback model.</p></li><li><p>Asynchronous: Use exponential backoff (wait 1 min, 2 min, 4 min). Since no user is waiting, we can afford to wait out a 5-minute provider outage.</p></li></ul></div><div><hr></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h2><span>Designing for low latency</span></h2><p><span>LLM inference is fundamentally slow. A user requesting a search result expects a response in &lt; 500 ms, but an LLM might take 5&#8211;10 seconds to generate a full answer. We cannot change the speed of inference, but we can architect around it. Hybrid processing (synchronous vs. asynchronous), response streaming and caching patterns can help us.</span></p><h3><span>Pattern: hybrid processing (synchronous vs. asynchronous)</span></h3><p><span>Given that we cannot block a user&#8217;s web request for 10 seconds while waiting for an LLM, one option is to separate the workloads. This is the most important latency-saving pattern.</span></p><p><span>Synchronous path (&lt; 2 s): For immediate, low-latency needs. These are tasks that must be fast, like code complete or a real-time e-commerce search. These paths should use fast, cheap models or rely heavily on caching.</span></p><p><span>Asynchronous path (&gt; 10 s): For high-latency, long-running tasks use a </span><strong><span>message queue</span></strong><span> (like Kafka or SQS) to decouple the initial request from the actual LLM work. The client receives an immediate &#8216;202 Accepted&#8217; response. For example, use when generating an AI-powered report, in agentic workflows, or for offline content generation. The client can poll for the result or receive it via a WebSocket/callback.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CF42!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1af1d49e-c0ce-406e-ae41-52d86a8caa41_950x954.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CF42!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1af1d49e-c0ce-406e-ae41-52d86a8caa41_950x954.png 424w, https://substackcdn.com/image/fetch/$s_!CF42!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1af1d49e-c0ce-406e-ae41-52d86a8caa41_950x954.png 848w, https://substackcdn.com/image/fetch/$s_!CF42!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1af1d49e-c0ce-406e-ae41-52d86a8caa41_950x954.png 1272w, https://substackcdn.com/image/fetch/$s_!CF42!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1af1d49e-c0ce-406e-ae41-52d86a8caa41_950x954.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CF42!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1af1d49e-c0ce-406e-ae41-52d86a8caa41_950x954.png" width="950" height="954" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1af1d49e-c0ce-406e-ae41-52d86a8caa41_950x954.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:954,&quot;width&quot;:950,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.4: Sync and async flows&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.4: Sync and async flows" title="Figure 2.4: Sync and async flows" srcset="https://substackcdn.com/image/fetch/$s_!CF42!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1af1d49e-c0ce-406e-ae41-52d86a8caa41_950x954.png 424w, https://substackcdn.com/image/fetch/$s_!CF42!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1af1d49e-c0ce-406e-ae41-52d86a8caa41_950x954.png 848w, https://substackcdn.com/image/fetch/$s_!CF42!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1af1d49e-c0ce-406e-ae41-52d86a8caa41_950x954.png 1272w, https://substackcdn.com/image/fetch/$s_!CF42!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1af1d49e-c0ce-406e-ae41-52d86a8caa41_950x954.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2.4: Sync and async flows</figcaption></figure></div><h3><span>Pattern: response streaming</span></h3><p><span>Even a fast 3-second response feels slow if the user is staring at a loading spinner, so if possible stream the response token-by-token. Once the model generates the first word, send it to the client. This dramatically improves perceived latency by lowering the time-to-first-token (TTFT). This is a non-negotiable pattern for any conversational or chat application.</span></p><h4><span>How it works</span></h4><p><span>We use </span><strong><span>SSE</span></strong><span> (</span><strong><span>server-sent events</span></strong><span>) over standard HTTP. SSE is unidirectional (server -&gt; client) and runs over standard HTTP/2, making it firewall-friendly and easy to implement.</span></p><p><span>Client: Opens a persistent connection.</span></p><p><span>Server: Instead of returning a JSON object, it returns a generator (an iterator)</span></p><p><span>Header: The server must set content-type: text/event-stream</span></p><p><span>Format: Data is sent in chunks prefixed &#8216;data:&#8217;:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pLDx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67b2b3a5-eeec-4a72-8dc0-cb85b875ffa1_1194x1262.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pLDx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67b2b3a5-eeec-4a72-8dc0-cb85b875ffa1_1194x1262.png 424w, https://substackcdn.com/image/fetch/$s_!pLDx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67b2b3a5-eeec-4a72-8dc0-cb85b875ffa1_1194x1262.png 848w, https://substackcdn.com/image/fetch/$s_!pLDx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67b2b3a5-eeec-4a72-8dc0-cb85b875ffa1_1194x1262.png 1272w, https://substackcdn.com/image/fetch/$s_!pLDx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67b2b3a5-eeec-4a72-8dc0-cb85b875ffa1_1194x1262.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pLDx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67b2b3a5-eeec-4a72-8dc0-cb85b875ffa1_1194x1262.png" width="1194" height="1262" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/67b2b3a5-eeec-4a72-8dc0-cb85b875ffa1_1194x1262.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1262,&quot;width&quot;:1194,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.5: Streaming tokens&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.5: Streaming tokens" title="Figure 2.5: Streaming tokens" srcset="https://substackcdn.com/image/fetch/$s_!pLDx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67b2b3a5-eeec-4a72-8dc0-cb85b875ffa1_1194x1262.png 424w, https://substackcdn.com/image/fetch/$s_!pLDx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67b2b3a5-eeec-4a72-8dc0-cb85b875ffa1_1194x1262.png 848w, https://substackcdn.com/image/fetch/$s_!pLDx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67b2b3a5-eeec-4a72-8dc0-cb85b875ffa1_1194x1262.png 1272w, https://substackcdn.com/image/fetch/$s_!pLDx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67b2b3a5-eeec-4a72-8dc0-cb85b875ffa1_1194x1262.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2.5: Streaming tokens</figcaption></figure></div><h3><span>Pattern: caching strategies</span></h3><p><span>LLM calls are slow and expensive, while cache hits are the fastest, cheapest LLM calls you can make. Therefore implement a multi-level caching strategy:</span></p><h4><span>Level 1: exact match</span></h4><div class="callout-block" data-callout="true"><ul><li><p>Scenario: A viral product launch. 10,000 users ask &#8216;When is shipping?&#8217;.</p></li><li><p>Mechanism: Hash the prompt string <mark data-color="rgb(255, 255, 0)" style="background-color: rgb(255, 255, 0); color: rgb(0, 0, 0);">(sha256(&#8221;When is shipping?&#8221;))</mark>. Check Redis.</p></li><li><p>Result: 9,999 users get a 5 ms response.</p></li><li><p>Why: It catches the stampede of identical queries.</p></li></ul></div><h4><span>Level 2: semantic match</span></h4><div class="callout-block" data-callout="true"><ul><li><p>Scenario: User A asks &#8216;How do I reset password?&#8217; User B asks &#8216;Forgot password, help&#8217;.</p></li><li><p>Mechanism: L1 misses (strings don&#8217;t match). We convert &#8216;Forgot password, help&#8217; to a vector. We query the vector DB for similar past questions.</p></li><li><p>Result: The DB finds that &#8216;How do I reset password?&#8217; has a 0.98 similarity. It returns the cached answer for User A to User B.</p></li><li><p>Why: It catches different phrasings of the same intent, saving expensive reasoning costs.</p></li></ul></div><h4><span>Level 3: proactive caching</span></h4><div class="callout-block" data-callout="true"><ul><li><p>Scenario: A personalized &#8216;Daily Report&#8217; for 100,000 users.</p></li><li><p>Mechanism: Don&#8217;t wait for them to open the app at 9:00 AM. Run a batch job at 6:00 AM. Generate the reports and store them into Redis.</p></li><li><p>Result: When users log in, the AI generation feels instant because it happened 3 hours ago.</p></li><li><p>Why: It moves latency from online (user waiting) to offline. This is a core pattern in adaptive learning platforms and e-commerce search.</p></li></ul></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nDhi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6db30fb9-3000-4f5e-be1e-2e898e001948_875x1147.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nDhi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6db30fb9-3000-4f5e-be1e-2e898e001948_875x1147.png 424w, https://substackcdn.com/image/fetch/$s_!nDhi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6db30fb9-3000-4f5e-be1e-2e898e001948_875x1147.png 848w, https://substackcdn.com/image/fetch/$s_!nDhi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6db30fb9-3000-4f5e-be1e-2e898e001948_875x1147.png 1272w, https://substackcdn.com/image/fetch/$s_!nDhi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6db30fb9-3000-4f5e-be1e-2e898e001948_875x1147.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nDhi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6db30fb9-3000-4f5e-be1e-2e898e001948_875x1147.png" width="875" height="1147" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6db30fb9-3000-4f5e-be1e-2e898e001948_875x1147.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1147,&quot;width&quot;:875,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.6: Caching strategy&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.6: Caching strategy" title="Figure 2.6: Caching strategy" srcset="https://substackcdn.com/image/fetch/$s_!nDhi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6db30fb9-3000-4f5e-be1e-2e898e001948_875x1147.png 424w, https://substackcdn.com/image/fetch/$s_!nDhi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6db30fb9-3000-4f5e-be1e-2e898e001948_875x1147.png 848w, https://substackcdn.com/image/fetch/$s_!nDhi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6db30fb9-3000-4f5e-be1e-2e898e001948_875x1147.png 1272w, https://substackcdn.com/image/fetch/$s_!nDhi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6db30fb9-3000-4f5e-be1e-2e898e001948_875x1147.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2.6: Caching strategy</figcaption></figure></div><h3><span>Pattern: coalesce caching</span></h3><p><span>During high-traffic events, thousands of users might ask the exact same question simultaneously. A standard cache misses the first time for everyone, causing a stampede of identical requests to the LLM.</span></p><p><span>The solution is middleware that identifies identical in-flight requests. It pauses subsequent requests, waits for the first request to complete, and then serves that single LLM response to all waiting users. This reduces load on the provider by orders of magnitude.</span></p><div><hr></div><h2><span>Designing for cost optimization</span></h2><p><span>LLMs introduce a new, variable, and unbounded operational cost (</span><strong><span>COGS</span></strong><span> &#8211; cost of goods sold). A single complex query can cost dollars. Architectural decisions are now financial decisions. Important in this context are the </span><strong><span>model router</span></strong><span>, </span><strong><span>dynamic traffic control</span></strong><span> and </span><strong><span>prompt engineering</span></strong><span> and </span><strong><span>compression</span></strong><span> patterns.</span></p><h3><span>Pattern: the model router</span></h3><p><span>Not all tasks require the smartest (and most expensive) model. Using GPT-5 for a simple grammar check is like booking a helicopter to avoid traffic. By implementing a rule engine within your LLM Gateway to act as a model router you can dynamically route requests to the cheapest model that is good enough for the task.</span></p><p><span>Example rules:</span></p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;29627da5-2b1c-48b1-9bc2-b2f428e7031c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">IF task_type == 'simple_grammar_check' THEN route_to 'local_llama_8b'.
IF task_type == 'complex reasoning' AND user_tier == 'premium' THEN route_to 'GPT-5'.</code></pre></div><h3><span>Pattern: dynamic traffic control (utilization-based routing)</span></h3><p><span>A static rule for traffic routing, e.g. IF task == complex THEN GPT-5, can cause bottlenecks during traffic spikes. In this case we can use a production-grade router that acts as a traffic controller, implementing </span><strong><span>load shedding</span></strong><span> if the primary model&#8217;s latency breaches the P99 SLA (e.g. &gt; 2s) due to provider congestion. In those circumstances the router proactively shifts traffic to the faster or cheaper model, even for complex tasks. A </span><strong><span>cost ceiling</span></strong><span> can be established by enforcing a hard cap on tokens. If a prompt exceeds a threshold (e.g. 8k tokens), force-route to a lower-cost model to prevent a single query causing cost regression.</span></p><h3><span>Pattern: prompt engineering and compression</span></h3><p><span>Cost is based on the number of input and output tokens; large prompts are expensive. Therefore treat your prompt context as a cost to be optimized. Before sending a large document (e.g. 50 pages of chat history) to an LLM, use a cheaper, faster model to summarize or compress it first.</span></p><div><hr></div><h2><span>Designing for grounding and data management</span></h2><p><span>LLMs hallucinate (invent facts) and have knowledge cutoffs (their training data is stale). How can we build a reliable enterprise application on this foundation? Retrieval-augmented generation (RAG) approaches are a fundamental pattern for enabling modern LLM applications to address these fundamental issues.</span></p><h3><span>Pattern: retrieval-augmented generation (RAG)</span></h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-p-g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc2eebf-d8de-4fa6-8370-a93a1dde6308_403x833.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-p-g!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc2eebf-d8de-4fa6-8370-a93a1dde6308_403x833.png 424w, https://substackcdn.com/image/fetch/$s_!-p-g!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc2eebf-d8de-4fa6-8370-a93a1dde6308_403x833.png 848w, https://substackcdn.com/image/fetch/$s_!-p-g!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc2eebf-d8de-4fa6-8370-a93a1dde6308_403x833.png 1272w, https://substackcdn.com/image/fetch/$s_!-p-g!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc2eebf-d8de-4fa6-8370-a93a1dde6308_403x833.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-p-g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc2eebf-d8de-4fa6-8370-a93a1dde6308_403x833.png" width="403" height="833" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ccc2eebf-d8de-4fa6-8370-a93a1dde6308_403x833.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:833,&quot;width&quot;:403,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.7: Retrieval-augmented generation&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.7: Retrieval-augmented generation" title="Figure 2.7: Retrieval-augmented generation" srcset="https://substackcdn.com/image/fetch/$s_!-p-g!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc2eebf-d8de-4fa6-8370-a93a1dde6308_403x833.png 424w, https://substackcdn.com/image/fetch/$s_!-p-g!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc2eebf-d8de-4fa6-8370-a93a1dde6308_403x833.png 848w, https://substackcdn.com/image/fetch/$s_!-p-g!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc2eebf-d8de-4fa6-8370-a93a1dde6308_403x833.png 1272w, https://substackcdn.com/image/fetch/$s_!-p-g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc2eebf-d8de-4fa6-8370-a93a1dde6308_403x833.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2.7: Retrieval-augmented generation</figcaption></figure></div><p><span>Not uncommonly we need the LLM to answer questions about private, real-time, or domain-specific data. Instead of asking the LLM a question, the solution is to tell it the answer. As the RAG name suggests, there are three principal aspects to consider:</span></p><p><span>Retrieve: When a user asks, &#8216;What&#8217;s the status of order #123?&#8217;, you first query the database to get the order details.</span></p><p><span>Augment: You augment the prompt with this retrieved data.</span></p><p><span>Generate: You instruct the LLM &#8211; &#8216;Based only on the following context, generate a friendly response&#8217;.</span></p><p><span>Example context: </span><span data-color="#e06666" style="color: rgb(224, 102, 102);">{order_details_json}</span></p><h3><span>Pattern: the ingestion pipeline</span></h3><p><span>To build the retrieval component of a RAG application, we must prepare the knowledge base (documents, tickets, etc.) for retrieval. That means creating an ingestion pipeline, typically as an asynchronous, scalable job (e.g. using Spark). This pipeline:</span></p><ol><li><p><span>Chunks large documents into small, semantically meaningful pieces</span></p></li><li><p><span>Embeds each chunk by calling an embedding model API</span></p></li><li><p><span>Stores the chunk and its corresponding vector in a </span><strong><span>vector database</span></strong><span> (e.g. OpenSearch, Pinecone, or pg_vector [with postgres])</span></p></li></ol><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ccbq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb05a87bd-534b-4422-9cf8-e26a3ef8e419_307x645.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ccbq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb05a87bd-534b-4422-9cf8-e26a3ef8e419_307x645.png 424w, https://substackcdn.com/image/fetch/$s_!ccbq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb05a87bd-534b-4422-9cf8-e26a3ef8e419_307x645.png 848w, https://substackcdn.com/image/fetch/$s_!ccbq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb05a87bd-534b-4422-9cf8-e26a3ef8e419_307x645.png 1272w, https://substackcdn.com/image/fetch/$s_!ccbq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb05a87bd-534b-4422-9cf8-e26a3ef8e419_307x645.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ccbq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb05a87bd-534b-4422-9cf8-e26a3ef8e419_307x645.png" width="307" height="645" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b05a87bd-534b-4422-9cf8-e26a3ef8e419_307x645.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:645,&quot;width&quot;:307,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.8: Ingestion pipeline&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.8: Ingestion pipeline" title="Figure 2.8: Ingestion pipeline" srcset="https://substackcdn.com/image/fetch/$s_!ccbq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb05a87bd-534b-4422-9cf8-e26a3ef8e419_307x645.png 424w, https://substackcdn.com/image/fetch/$s_!ccbq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb05a87bd-534b-4422-9cf8-e26a3ef8e419_307x645.png 848w, https://substackcdn.com/image/fetch/$s_!ccbq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb05a87bd-534b-4422-9cf8-e26a3ef8e419_307x645.png 1272w, https://substackcdn.com/image/fetch/$s_!ccbq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb05a87bd-534b-4422-9cf8-e26a3ef8e419_307x645.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2.8: Ingestion pipeline</figcaption></figure></div><h4><span>Pattern: hybrid RAG</span></h4><p><span>RAG using just vector search (or even simple term search) is not enough for complex, interconnected data. It&#8217;s good at finding similar content, but bad at traversing relationships. Hybrid RAG approaches such as GraphRAG represent an advanced RAG pattern where the ingestion pipeline populates both a vector database (for semantic similarity) and a knowledge graph (e.g. Neptune or Neo4j) for structured relationships.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bkLW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0422109-57ac-465c-b9a1-743fa4eb97d5_1060x655.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bkLW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0422109-57ac-465c-b9a1-743fa4eb97d5_1060x655.png 424w, https://substackcdn.com/image/fetch/$s_!bkLW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0422109-57ac-465c-b9a1-743fa4eb97d5_1060x655.png 848w, https://substackcdn.com/image/fetch/$s_!bkLW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0422109-57ac-465c-b9a1-743fa4eb97d5_1060x655.png 1272w, https://substackcdn.com/image/fetch/$s_!bkLW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0422109-57ac-465c-b9a1-743fa4eb97d5_1060x655.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bkLW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0422109-57ac-465c-b9a1-743fa4eb97d5_1060x655.png" width="1060" height="655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d0422109-57ac-465c-b9a1-743fa4eb97d5_1060x655.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:655,&quot;width&quot;:1060,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.9: Hybrid RAG architecture &#8211; combining vector search with knowledge graphs&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.9: Hybrid RAG architecture &#8211; combining vector search with knowledge graphs" title="Figure 2.9: Hybrid RAG architecture &#8211; combining vector search with knowledge graphs" srcset="https://substackcdn.com/image/fetch/$s_!bkLW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0422109-57ac-465c-b9a1-743fa4eb97d5_1060x655.png 424w, https://substackcdn.com/image/fetch/$s_!bkLW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0422109-57ac-465c-b9a1-743fa4eb97d5_1060x655.png 848w, https://substackcdn.com/image/fetch/$s_!bkLW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0422109-57ac-465c-b9a1-743fa4eb97d5_1060x655.png 1272w, https://substackcdn.com/image/fetch/$s_!bkLW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0422109-57ac-465c-b9a1-743fa4eb97d5_1060x655.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2.9: Hybrid RAG architecture &#8211; combining vector search with knowledge graphs</figcaption></figure></div><p><span>The RAG orchestrator then queries both systems to build a much richer, more accurate context, as a customer support agent system demonstrates well.</span></p><h3><span>Pattern: function calling (tool usage)</span></h3><p><span>To overcome the problem that LLMs are generally bad at (for example) math and cannot interact with the outside world, we ask it to decide which tool to use. We provide a schema of tools (e.g. </span><span data-color="#e06666" style="color: rgb(224, 102, 102);">get_weather(city)</span><span>), and the LLM outputs a structured JSON object requesting that function. The application layer executes the code and feeds the result back to the LLM.</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Designing for testability and observability</span></h2><p><span>Testability for AI-powered systems is crucial in order to maintain the quality of response expected from the system, but for non-deterministic systems it is not trivial to write unit tests that would assert on a given known output.</span></p><p><span>We use a pattern of using an LLM as a judge by creating a golden dataset, which has input with their ideal responses, and using that dataset, the LLM judge is trained. This LLM is later used to evaluate the quality of response returned by the system against the ideal response.</span></p><p><span>Problem: How do you write a unit test for a system that is non-deterministic?</span></p><p><span>Something like assert(response) == &#8220;expected_string&#8221; will fail constantly.</span></p><h3><span>Pattern: golden datasets</span></h3><p><span>We need a reliable way to catch regressions in AI quality, and one way is to create a &#8216;golden dataset&#8217; of 50&#8211;100 representative inputs and their ideal outputs. In your CI/CD pipeline, run your system against this set to ensure that prompt changes or model upgrades haven&#8217;t broken core functionality.</span></p><h3><span>Pattern: LLM-as-a-Judge</span></h3><p><span>How do you assert the quality of a golden set test at scale? You can&#8217;t manually review 100 responses on every build. What you can do, however, is use a powerful LLM like GPT-5 as an evaluator or judge. You feed the judge the original prompt, the golden answer, and your system&#8217;s actual response, and then ask the Judge to score the actual response from 1 to 5 on metrics like accuracy, groundedness, and tone. This gives you a quantifiable quality metric you can track over time.</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mox2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0036a75-f6a1-46db-a9b2-0970a5def762_1020x221.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mox2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0036a75-f6a1-46db-a9b2-0970a5def762_1020x221.png 424w, https://substackcdn.com/image/fetch/$s_!mox2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0036a75-f6a1-46db-a9b2-0970a5def762_1020x221.png 848w, https://substackcdn.com/image/fetch/$s_!mox2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0036a75-f6a1-46db-a9b2-0970a5def762_1020x221.png 1272w, https://substackcdn.com/image/fetch/$s_!mox2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0036a75-f6a1-46db-a9b2-0970a5def762_1020x221.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mox2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0036a75-f6a1-46db-a9b2-0970a5def762_1020x221.png" width="1020" height="221" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d0036a75-f6a1-46db-a9b2-0970a5def762_1020x221.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:221,&quot;width&quot;:1020,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.10: The LLM-as-a-Judge evaluation pipeline for CI/CD&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.10: The LLM-as-a-Judge evaluation pipeline for CI/CD" title="Figure 2.10: The LLM-as-a-Judge evaluation pipeline for CI/CD" srcset="https://substackcdn.com/image/fetch/$s_!mox2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0036a75-f6a1-46db-a9b2-0970a5def762_1020x221.png 424w, https://substackcdn.com/image/fetch/$s_!mox2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0036a75-f6a1-46db-a9b2-0970a5def762_1020x221.png 848w, https://substackcdn.com/image/fetch/$s_!mox2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0036a75-f6a1-46db-a9b2-0970a5def762_1020x221.png 1272w, https://substackcdn.com/image/fetch/$s_!mox2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0036a75-f6a1-46db-a9b2-0970a5def762_1020x221.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Figure 2.10: The LLM-as-a-Judge evaluation pipeline for CI/CD</figcaption></figure></div><h3><span>Pattern: new observability metrics</span></h3><p><span>Existing performance dashboards (CPU, RAM, 5xx errors, etc.) are insufficient in the LLM era. We need to add a new layer of monitoring focused on the LLM itself, for example:</span></p><p><span>Cost: Track Cost_Per_Query and Total_Cost_Per_User. Set alerts for cost spikes.</span></p><p><span>Performance: Monitor P99 time to first token (TTFT) and tokens per second (TPS).</span></p><p><span>Provider health: Monitor 429_Rate_Limit_Errors and 5xx_Server_Errors per provider to feed your circuit breakers.</span></p><p><span>Quality: Track your LLM-as-a-Judge scores. Monitor the </span><strong><span>escalation rate</span></strong><span> &#8211; this measures the percentage of conversations where the AI fails to resolve the issue, forcing the system to transfer the work to a human (for example in case of customer support agents). A spike in this metric indicates a drop in model quality.</span></p><div><hr></div><h2><span>Designing for security and trust</span></h2><p><span>Treating an LLM as a simple API call is naive and dangerous. It&#8217;s a non-deterministic component that you are inviting inside your trusted system. We must architect our systems to defend against a new class of vulnerabilities. A variety of patterns can help us in this, including the following. In this section we summarize the main security patterns, the threats they address and the mitigations they provide.</span></p><h3><span>Pattern: mitigating prompt injection</span></h3><p><span>Threat: Attackers trick an LLM into ignoring its original instructions by embedding malicious commands in prompts, causing it to perform unintended actions.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tcKA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7558f49f-335a-4b63-87e7-5900b5950958_802x870.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tcKA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7558f49f-335a-4b63-87e7-5900b5950958_802x870.png 424w, https://substackcdn.com/image/fetch/$s_!tcKA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7558f49f-335a-4b63-87e7-5900b5950958_802x870.png 848w, https://substackcdn.com/image/fetch/$s_!tcKA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7558f49f-335a-4b63-87e7-5900b5950958_802x870.png 1272w, https://substackcdn.com/image/fetch/$s_!tcKA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7558f49f-335a-4b63-87e7-5900b5950958_802x870.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tcKA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7558f49f-335a-4b63-87e7-5900b5950958_802x870.png" width="802" height="870" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7558f49f-335a-4b63-87e7-5900b5950958_802x870.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:870,&quot;width&quot;:802,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.11: The firewall pattern &#8211; using a lightweight model to filter malicious intent before it reaches the core model&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.11: The firewall pattern &#8211; using a lightweight model to filter malicious intent before it reaches the core model" title="Figure 2.11: The firewall pattern &#8211; using a lightweight model to filter malicious intent before it reaches the core model" srcset="https://substackcdn.com/image/fetch/$s_!tcKA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7558f49f-335a-4b63-87e7-5900b5950958_802x870.png 424w, https://substackcdn.com/image/fetch/$s_!tcKA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7558f49f-335a-4b63-87e7-5900b5950958_802x870.png 848w, https://substackcdn.com/image/fetch/$s_!tcKA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7558f49f-335a-4b63-87e7-5900b5950958_802x870.png 1272w, https://substackcdn.com/image/fetch/$s_!tcKA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7558f49f-335a-4b63-87e7-5900b5950958_802x870.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2.11: The firewall pattern &#8211; using a lightweight model to filter malicious intent before it reaches the core model</figcaption></figure></div><h4><span>Mitigation</span></h4><p><span>Instruction/data separation: Clearly separate trusted instructions from untrusted data from the user. Use role-based API structures (system, user) and wrap all user input in clear delimiters like &lt;&gt;.</span></p><p><span>Input filtering: Use a second, simpler, faster LLM to classify the intent of a user&#8217;s prompt. If it detects a likely attack, reject it before it reaches your primary model.</span></p><p><span>Output filtering: Always validate the LLM&#8217;s response. If it contains any system prompt text or suspicious keywords, block it.</span></p><h3><span>Pattern: secure output handling</span></h3><p><span>Threat: Insecure output handling occurs when an LLM&#8217;s outputs are not properly sanitized before being used, potentially leading to attacks like </span><strong><span>cross-site scripting</span></strong><span> (</span><strong><span>XSS</span></strong><span>).</span></p><h4><span>Mitigation</span></h4><p><span>Treat as untrusted: Never eval() code output.</span></p><p><span>Sanitize and encode: If the output is HTML, sanitize it. If it&#8217;s text for a web page, encode it to prevent XSS.</span></p><p><span>Validate: If you expect JSON, parse it in a try/catch block and validate its schema.</span></p><p><span>Parameterize: If the LLM helps build a SQL query, have it generate the parameters for a pre-defined, parameterized query you control. Never execute a raw SQL string from an LLM.</span></p><h3><span>Pattern: mitigating excessive agency</span></h3><p><span>Threat: This happens when an LLM is given too much control and makes unauthorized decisions or actions without human oversight.</span></p><h4><span>Mitigation</span></h4><p><span>Dynamic permissions: Do not give the LLM agent a static set of all possible tools.</span></p><p><span>Plan-approve-execute: Implement a multi-step loop. The LLM proposes a plan. Your code approves this plan. Only then do you execute it with a scoped-down client.</span></p><p><span>Human-in-the-loop (HITL): For high-impact actions (e.g. &#8216;delete database&#8217;, &#8216;refund customer&#8217;), always require explicit human approval.</span></p><h3><span>Pattern: mitigating sensitive information disclosure</span></h3><p><span>Threat: The model may inadvertently reveal confidential data or </span><strong><span>personally identifiable information</span></strong><span> (</span><strong><span>PII</span></strong><span>) that was present in its training data.</span></p><h4><span>Mitigation (data hygiene)</span></h4><p><span>PII/data scrubbing: Aggressively sanitize all data before it is used for training or RAG ingestion.</span></p><p><span>Zero-retention policies: For ultra-sensitive data (like user code), process it in-memory and discard it immediately. Do not store it.</span></p><p><span>Tenant-level RAG filtering: This is a critical architectural pattern. All RAG queries must include a filter for tenant_id or user_id. Never perform a vector search on the entire database and hope the LLM picks the right data.</span></p><h3><span>Pattern: preventing model denial of service (MDoS)</span></h3><p><span>Threat: Attackers overload the LLM with resource-intensive requests, slowing it down or making it unavailable.</span></p><h4><span>Mitigation</span></h4><p><span>API rate limiting: Enforce strict per-user and per-IP rate limits at the API gateway</span></p><p><span>Input validation: Reject queries that are obviously abusive (e.g. a 50,000-token prompt)</span></p><p><span>Cost-based throttling: Implement logic to monitor the cost of a user&#8217;s queries in real-time. If a single user is incurring high costs, temporarily throttle their access. We would want to return a 429 in case of user spam/abuse or redirect to a lower model in case user hit their budget.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5be2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55606989-d179-4cf1-8df9-cb4889d961ab_620x917.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5be2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55606989-d179-4cf1-8df9-cb4889d961ab_620x917.png 424w, https://substackcdn.com/image/fetch/$s_!5be2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55606989-d179-4cf1-8df9-cb4889d961ab_620x917.png 848w, https://substackcdn.com/image/fetch/$s_!5be2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55606989-d179-4cf1-8df9-cb4889d961ab_620x917.png 1272w, https://substackcdn.com/image/fetch/$s_!5be2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55606989-d179-4cf1-8df9-cb4889d961ab_620x917.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5be2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55606989-d179-4cf1-8df9-cb4889d961ab_620x917.png" width="620" height="917" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/55606989-d179-4cf1-8df9-cb4889d961ab_620x917.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:917,&quot;width&quot;:620,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.12: Throttling to prevent DoS attacks on LLM&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.12: Throttling to prevent DoS attacks on LLM" title="Figure 2.12: Throttling to prevent DoS attacks on LLM" srcset="https://substackcdn.com/image/fetch/$s_!5be2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55606989-d179-4cf1-8df9-cb4889d961ab_620x917.png 424w, https://substackcdn.com/image/fetch/$s_!5be2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55606989-d179-4cf1-8df9-cb4889d961ab_620x917.png 848w, https://substackcdn.com/image/fetch/$s_!5be2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55606989-d179-4cf1-8df9-cb4889d961ab_620x917.png 1272w, https://substackcdn.com/image/fetch/$s_!5be2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55606989-d179-4cf1-8df9-cb4889d961ab_620x917.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2.12: Throttling to prevent DoS attacks on LLM</figcaption></figure></div><p><span>Here&#8217;s an example implementation:</span></p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;482c59db-8b95-45a5-8663-05a39f0fa72f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from fastapi import HTTPException

def route_request(user, prompt):
    # 1. HARD LIMIT CHECK (Redis Counter)
    current_rate = redis.get(f"rate:{user.id}")
    if current_rate &gt; 50:
        # Scenario A: Abuse -&gt; Hard Stop
        raise HTTPException(status_code=429, detail="Rate limit exceeded.")

    # 2. SOFT BUDGET CHECK (DB Query)
    daily_spend = db.get_spend(user.id)
    budget_limit = 10.00  # $10 limit
    if daily_spend &gt; budget_limit:
        # Scenario B: Over Budget -&gt; Downgrade (Soft Throttle)
        # We don't fail; we just swap the model
        print(f"User {user.id} over budget. Downgrading to Tier 3.")
        return call_llm(model="llama-3-8b", prompt=prompt)

    # 3. NORMAL PATH
    return call_llm(model="gpt-4", prompt=prompt)</code></pre></div><h3><span>Pattern: preventing data poisoning</span></h3><p><span>Threat: Malicious actors tamper with an LLM&#8217;s training data to corrupt its behavior, leading to biased or incorrect outputs.</span></p><p><span>Mitigation (ingestion control):</span></p><p><span>Trusted sources: Only use and ingest data from known, trusted sources.</span></p><p><span>Data lineage: Track the origin of all data used for training or RAG.</span></p><p><span>HITL review: For fine-tuning data, use human experts to review and validate the datasets before training.</span></p><h3><span>Pattern: mitigating supply chain and insecure plugin vulnerabilities</span></h3><p><span>Threat: The security of an LLM can be compromised through third-party services, plugins, or datasets that are themselves vulnerable. This includes insecure plugin design, which can be exploited for attacks like SQL injection.</span></p><h4><span>Mitigation</span></h4><p><span>Minimize functionality: Any plugin or tool given to the LLM should have the absolute minimal functionality needed (e.g. read_email, not delete_email).</span></p><p><span>Validate plugin inputs: Treat all data passed to a plugin as untrusted. Sanitize it to prevent injection attacks within the plugin itself.</span></p><p><span>Vulnerability scanning: Regularly scan all third-party libraries, containers, and models for known vulnerabilities.</span></p><p><span>Use the genAI service / LLM gateway: Your gateway allows you to quickly replace a provider or model that is found to be compromised.</span></p><div><hr></div><h2><span>Engineering for production</span></h2><p><span>We cannot manage what we do not track. Beyond standard APM (application performance monitoring), we must track:</span></p><p><span>TTFT (time to first token): The perceived latency. How long until the user sees the first character? High TTFT kills engagement.</span></p><p><strong><span>TPOT</span></strong><span> (</span><strong><span>time per output token</span></strong><span>): The generation speed. If this is high, the model is too heavy or the provider is overloaded.</span></p><p><span>Context utilization: Is the context window filled? Use the model optimized for token usage for the use case for cost effectiveness.</span></p><h2><span>Training with test data</span></h2><p><span>In earlier sections we have discussed golden datasets &#8211; inputs with their ideal outputs used for training the LLM as a judge. Let&#8217;s now look at how we can procure this ideal data for different use cases.</span></p><h3><span>Sourcing datasets</span></h3><div class="callout-block" data-callout="true"><ul><li><p>Curate diverse set of open source projects with varied languages, sizes and domains</p></li><li><p>Refresh the dataset regularly to incorporate up to date coding style and practices</p></li><li><p>Create custom open-source repos with intentional gaps and edge cases to stress test behaviors</p></li></ul></div><h3><span>Testing and training for code complete</span></h3><p><span>Randomly remove code blocks, functions, classes from the code. Start typing signatures of missing classes and functions. Compare the IDE code complete suggestions with the actual classes and functions to measure accuracy and semantic correctness. Iteratively tune the LLM model prompts based on error cases and coverage gaps identified.</span></p><h3><span>Testing for chat responses</span></h3><p><span>Clone the repo removing all the comments and docs. Ask the system to explain code blocks, functions, classes and compare against the original repo&#8217;s docs, README files or comments.</span></p><h3><span>Testing for agentic workflow and tasks</span></h3><p><span>Assign real world tasks like adding documentation, writing test cases or refactoring to the agent. Combine multiple tasks in scenarios.</span></p><p><span>Automatically check the code and tests and documentation for correctness and style.</span></p><p><span>Integrating these tests into CI/CDs helps keep models and prompts updated with evolving repos.</span></p><h4><span>Pattern: quantitative evaluation testing</span></h4><p><span>The goal of this pattern is to move beyond vibes and qualitative observations to a measurable correctness percentage.</span></p><p><span>The problem is that agentic flows are multi-step and non-deterministic. Traditional pass/fail unit tests often fail to capture the nuance of a complex task that is mostly correct but slightly off in tone or style.</span></p><p><span>To tackle this distinctive character we assign a numerical score to every agentic run by breaking the output into weighted criteria. This allows us to track an evaluation success rate, the percentage of tasks successfully completed to the golden standard.</span></p><p><span>To calculate the success rate we define weights to the different output facets that reflect their relative importance (e.g. logic: 50%, syntax: 30%, documentation: 20%).</span></p><p><span>Then we execute the agent across a golden dataset of at least 50 representative tasks, and score according to the following metric:</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3WNA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5751aa4e-7a98-46e0-ba13-26b404e4a499_830x73.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3WNA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5751aa4e-7a98-46e0-ba13-26b404e4a499_830x73.png 424w, https://substackcdn.com/image/fetch/$s_!3WNA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5751aa4e-7a98-46e0-ba13-26b404e4a499_830x73.png 848w, https://substackcdn.com/image/fetch/$s_!3WNA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5751aa4e-7a98-46e0-ba13-26b404e4a499_830x73.png 1272w, https://substackcdn.com/image/fetch/$s_!3WNA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5751aa4e-7a98-46e0-ba13-26b404e4a499_830x73.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3WNA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5751aa4e-7a98-46e0-ba13-26b404e4a499_830x73.png" width="830" height="73" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5751aa4e-7a98-46e0-ba13-26b404e4a499_830x73.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:73,&quot;width&quot;:830,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Eq&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Eq" title="Eq" srcset="https://substackcdn.com/image/fetch/$s_!3WNA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5751aa4e-7a98-46e0-ba13-26b404e4a499_830x73.png 424w, https://substackcdn.com/image/fetch/$s_!3WNA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5751aa4e-7a98-46e0-ba13-26b404e4a499_830x73.png 848w, https://substackcdn.com/image/fetch/$s_!3WNA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5751aa4e-7a98-46e0-ba13-26b404e4a499_830x73.png 1272w, https://substackcdn.com/image/fetch/$s_!3WNA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5751aa4e-7a98-46e0-ba13-26b404e4a499_830x73.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><span>Figure 2.13 shows the overall evaluation process:</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!htsq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7d0950b-bd89-429c-bee5-1b2b565c81f9_1338x312.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!htsq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7d0950b-bd89-429c-bee5-1b2b565c81f9_1338x312.png 424w, https://substackcdn.com/image/fetch/$s_!htsq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7d0950b-bd89-429c-bee5-1b2b565c81f9_1338x312.png 848w, https://substackcdn.com/image/fetch/$s_!htsq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7d0950b-bd89-429c-bee5-1b2b565c81f9_1338x312.png 1272w, https://substackcdn.com/image/fetch/$s_!htsq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7d0950b-bd89-429c-bee5-1b2b565c81f9_1338x312.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!htsq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7d0950b-bd89-429c-bee5-1b2b565c81f9_1338x312.png" width="1338" height="312" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d7d0950b-bd89-429c-bee5-1b2b565c81f9_1338x312.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:312,&quot;width&quot;:1338,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2.13: Weighted agentic evaluation&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2.13: Weighted agentic evaluation" title="Figure 2.13: Weighted agentic evaluation" srcset="https://substackcdn.com/image/fetch/$s_!htsq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7d0950b-bd89-429c-bee5-1b2b565c81f9_1338x312.png 424w, https://substackcdn.com/image/fetch/$s_!htsq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7d0950b-bd89-429c-bee5-1b2b565c81f9_1338x312.png 848w, https://substackcdn.com/image/fetch/$s_!htsq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7d0950b-bd89-429c-bee5-1b2b565c81f9_1338x312.png 1272w, https://substackcdn.com/image/fetch/$s_!htsq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7d0950b-bd89-429c-bee5-1b2b565c81f9_1338x312.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Figure 2.13: Weighted agentic evaluation</figcaption></figure></div><p><span>When to use it: Use this in your CI/CD pipeline. If a prompt change or model upgrade causes the correctness percentage to drop below a defined threshold (e.g. 90%) then the deployment should be automatically blocked to prevent quality regression.</span></p><h3><span>Mutation testing</span></h3><p><span>Add logic and syntax errors to your code. Evaluate whether the system can identify and fix these errors.</span></p><h3><span>Negative prompt testing</span></h3><p><span>Give confusing or risky actions like &#8216;Delete all databases&#8217;. Evaluate whether the system is able to flag it as a risk and ignore/refuse or ask the user for further clarification.</span></p><h2><span>Respecting user privacy</span></h2><p><span>We want to be able to use an LLM-powered system for code completions, reviews, generation etc. without leaking our codebase to the model. Techniques for ensuring that user privacy is maintained while the data is still served from the model include:</span></p><div class="callout-block" data-callout="true"><ul><li><p>Ephemeral data handling: Code snippets and related info, like filenames, should never be saved to disk or databases. Encrypted code should only be decrypted in the computer&#8217;s live memory for processing and deleted the moment it&#8217;s no longer needed.</p></li><li><p>Embedding-only search: We don&#8217;t store the actual code. Instead, code snippets are converted into a mathematical format (vectors) for searching. This process is irreversible, so the original code cannot be reconstructed from the database. (We also scramble all metadata. Real file and function names are replaced with anonymous, hashed IDs.)</p></li><li><p>Strict access control</p><ul><li><p>Minimized code transfer: Only code context required for a query or code complete is sent to the server, not the entire codebase.</p></li><li><p>Encryption policies: All requests sent to/from the client are encrypted, and all data stored on server side is encrypted.</p></li></ul></li></ul></div><h2><span>Summary</span></h2><p><span>The patterns we&#8217;ve just looked at, from the GenAI service and async queues to RAG and LLM-as-a-Judge, form an architect&#8217;s playbook for building production-grade, AI-powered systems. In the rest of the book we examine four case studies that utilize this playbook.</span></p><div><hr></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="pullquote"><p>This practical deep-dive is an excerpt, Chapter 2, from <strong><a href="https://www.packtpub.com/en-us/product/system-design-for-the-llm-era-9781807789923">System Design for the LLM Era: Patterns and principles for production-grade AI architecture</a></strong> by <a href="https://in.linkedin.com/in/sampritimitra">Sampriti Mitra</a>, published by Packt. It is shared here with the publisher&#8217;s permission for knowledge sharing with the Deep Engineering community. All rights remain with Packt Publishing. This content may not be reproduced, redistributed, or remixed in any form without the publisher&#8217;s written consent. You can get the full book <a href="https://www.packtpub.com/en-us/product/system-design-for-the-llm-era-9781807789923">here</a>.</p></div>]]></content:encoded></item><item><title><![CDATA[Deep Engineering #54: Sibasis Padhi on Governing Agentic Operations Before They Amplify Failures]]></title><description><![CDATA[Sibasis Padhi on Autonomic Reliability Governance, and how actuation budgets, blast radius limits, and reversibility keep self-healing from amplifying failure.]]></description><link>https://deepengineering.net/p/issue-54-governing-agentic-ai-operations</link><guid isPermaLink="false">https://deepengineering.net/p/issue-54-governing-agentic-ai-operations</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 02 Jul 2026 17:17:02 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/906c47f7-9323-4cb8-a7ab-0bb6c57756fb_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3><strong><a href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng">ARC 2026: Software Architecture in the Age of AI</a></strong></h3><p><em>Packt&#8217;s flagship <strong>two-day</strong> virtual summit </em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!a2PN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 424w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 848w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!a2PN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;ARC 2026: Software Architecture in the Age of AI&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="ARC 2026: Software Architecture in the Age of AI" title="ARC 2026: Software Architecture in the Age of AI" srcset="https://substackcdn.com/image/fetch/$s_!a2PN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 424w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 848w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!a2PN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef96fc39-8a27-4d7a-8ebc-f5d67d2154d9_1880x940.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>AI is reshaping software architecture, putting new demands on scalability, governance, reliability, and observability. <a href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng">ARC 2026</a> brings together architects, CTOs, and AI practitioners for keynotes, panels, and workshops on agentic system design, modernizing enterprise apps for AI, and building governable, observable AI systems.</p><blockquote><p>The lineup includes <a href="https://www.linkedin.com/in/ali-arsanjani">Dr. Ali Arsanjani</a> (Director, Applied AI, Google), <a href="https://www.linkedin.com/in/davidping">David Ping</a> (Head of Solutions Architecture, AWS), and <a href="https://www.linkedin.com/in/chi-wang-autogen">Chi Wang</a> (Google DeepMind), alongside other notable leaders distilling practical insights.</p></blockquote><p style="text-align: center;"><span>&#128467;&#65039;  </span><em><strong>25</strong> to <strong>26</strong> July, <strong>10:30</strong> am ET</em></p><p style="text-align: center;"><span>Use code </span><strong>DEEPENG50</strong><span> for 50% off the early bird price.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng&quot;,&quot;text&quot;:&quot;Reserve your spot&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.eventbrite.co.uk/e/arc-2026-software-architecture-in-the-age-of-ai-tickets-1991591030384?aff=deepeng"><span>Reserve your spot</span></a></p><div><hr></div><p><span>&#9997;&#65039; </span><strong><span>From the editor&#8217;s desk,</span></strong></p><p><span>Welcome to the </span><strong><span>54th</span></strong><span> issue of </span><strong><span>Deep Engineering</span></strong><span>!</span></p><p>Site reliability engineering is already agentic, and agents are moving from the edges of the incident into its core. New Relic unveiled Autopilot at its <a href="https://finance.yahoo.com/technology/ai/articles/relic-autopilot-relic-ground-truth-130000437.html">New Relic NOW</a> event on June 23, an automated SRE agent that triages incidents, identifies root causes, and scopes remediations the moment an alert fires. &#8220;Operations are going headless,&#8221; said New Relic&#8217;s Head of AI, <a href="http://linkedin.com/in/camden-swita-54a41aa/">Camden Swita</a>, describing agents that pull what they need through APIs and act.</p><p>That is exactly when governance starts to matter more than speed. Once an agent can decide and push a remediation, the reliability question is no longer whether automation moves fast but whether its actions stay bounded. Faster, wider change across a tightly coupled system amplifies trouble as easily as it absorbs it, and the retry, scaling, and routing reflexes that steady a healthy platform can drive a cascade once it is already stressed. Tellingly, even these launches arrive wrapped in the language of guardrails and human review.</p><p>The engineers who feel that trade-off most work where a wrong automated action shows up immediately on the balance sheet. <a href="https://www.linkedin.com/in/sibasis-padhi">Sibasis Padhi</a>, a Staff Software Engineer at <a href="https://www.linkedin.com/company/walmartglobaltech/">Walmart Global Tech</a>, builds large-scale financial platforms of exactly that kind. In his Deep Engineering article &#8220;<a href="https://deepengineering.net/p/autonomic-governance-for-agentic-systems">Autonomic Governance for Agentic Systems</a>,&#8221; he argues that reliability engineering must evolve from managing distributed systems to governing the automation that manages them.</p><p>Padhi makes the case that the fix is not less automation but automation you can bound, audit, and reverse.</p><p>Let&#8217;s get started.</p><div><hr></div><p><strong><span>Featured Newsletter: </span><a href="https://thehustlingengineer.substack.com/">The Hustling Engineer</a></strong></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://thehustlingengineer.substack.com/" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!INnN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe3928da-2936-4f40-810a-caa16995c9e1_500x500.png 424w, https://substackcdn.com/image/fetch/$s_!INnN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe3928da-2936-4f40-810a-caa16995c9e1_500x500.png 848w, https://substackcdn.com/image/fetch/$s_!INnN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe3928da-2936-4f40-810a-caa16995c9e1_500x500.png 1272w, https://substackcdn.com/image/fetch/$s_!INnN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe3928da-2936-4f40-810a-caa16995c9e1_500x500.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!INnN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe3928da-2936-4f40-810a-caa16995c9e1_500x500.png" width="238" height="238" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fe3928da-2936-4f40-810a-caa16995c9e1_500x500.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:500,&quot;width&quot;:500,&quot;resizeWidth&quot;:238,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://thehustlingengineer.substack.com/&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!INnN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe3928da-2936-4f40-810a-caa16995c9e1_500x500.png 424w, https://substackcdn.com/image/fetch/$s_!INnN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe3928da-2936-4f40-810a-caa16995c9e1_500x500.png 848w, https://substackcdn.com/image/fetch/$s_!INnN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe3928da-2936-4f40-810a-caa16995c9e1_500x500.png 1272w, https://substackcdn.com/image/fetch/$s_!INnN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe3928da-2936-4f40-810a-caa16995c9e1_500x500.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Read by <strong>25,000+</strong> software engineers, <a href="https://thehustlingengineer.substack.com/">The Hustling Engineer</a> covers career growth, AI, interviews, and productivity. Every week you get actionable insights, engineering deep dives, and lessons from top tech companies.</p><p>&#128233; Get one actionable lesson every week to accelerate your tech career.</p><p><strong><span>&#8594;</span><a href="https://thehustlingengineer.substack.com/subscribe"><span> </span>Subscribe to The Hustling Engineer</a></strong></p><div><hr></div><p>&#129504; Expert Insight</p><h2>Autonomic Governance for Agentic Systems</h2><p><em>by <a href="https://www.linkedin.com/in/sibasis-padhi">Sibasis Padhi</a></em></p><p>The incident often starts small. Usually one dependency slows, tail latency grows, and a few requests time out. And then the platform does what we trained it to do. All the while, customers automatically retry their failed requests, and very quickly there are far more retry attempts than real new customers trying to use the service. But the autoscaling system sees the growing queue and reads it as exploding traffic, so it adds more servers, often in the wrong place, which sends even more work to the spot that is already overloaded.</p><p>This eats into the bottom line and the return on engineering investment as the bill keeps mounting while the system handles less real work. Circuit breakers finally trip and try to move traffic elsewhere, but those other paths were never built to absorb a sudden flood.</p><p>Everything might still look like it is working on paper, but the whole situation keeps getting worse. Now compound that by adding agentic AI to operations, systems that recommend or execute operational actions, which increases the speed and breadth of change under conditions that are already chaotic.</p><p>AI-driven operational automation can increase fragility when it increases the rate, scope, or coupling of production actions without bounded actuation, explicit constraints, and auditability. Under governance, the same automation reduces fragility by enforcing safety envelopes and reversible decision paths.</p><p>This is not an argument against automation or AI. It is an argument that reliability engineering must evolve from managing distributed systems to governing the automation that manages them, especially under simultaneous SLO, cost, and compliance constraints. I call this discipline Autonomic Reliability Governance (ARG). ARG is not more automation. But governed automation that is policy-constrained, auditable, and rollback-capable by design. It ensures that AI-assisted or agentic operational decisions cannot escalate into amplification engines under stress.</p><h3>Why reliability automation amplifies in microservice ecosystems</h3><p>In microservice systems, outages rarely come from a single broken component. They emerge from how services interact, through feedback loops, dependencies, and cascading effects. Mechanisms that improve reliability in normal conditions can unintentionally amplify problems when the system is already under stress, which is the feedback behavior Karl Johan &#197;str&#246;m and Richard Murray formalize in <em><a href="https://press.princeton.edu/books/hardcover/9780691193984/feedback-systems">Feedback Systems</a></em> (Princeton University Press, 2008).</p><h4>Retries multiply load under stress</h4><p>Retries improve success rates when failures are transient and capacity is available. Under degradation, retries become a multiplier, the dynamic Jeffrey Dean and Luis Andr&#233; Barroso describe in <a href="https://cacm.acm.org/research/the-tail-at-scale/">&#8220;The Tail at Scale&#8221;</a> (Communications of the ACM, February 2013).</p><div class="callout-block" data-callout="true"><p>A dependency slows &#8594; timeouts increase &#8594; retries surge &#8594; downstream work increases &#8594; queues grow &#8594; latency rises &#8594; more timeouts &#8594; more retries.</p></div><p>This is how a minor regression becomes a self-inflicted flood. It can resemble a resource exhaustion event even when there is no attacker. The mechanism is simple. Unbounded retries consume the remaining capacity of an already-constrained dependency. The root issue is not that retries are bad. The issue is that unbounded retries are a powerful actuator that must be governed by system-level constraints such as SLO, cost, and blast radius.</p><h4>Autoscaling is a blunt actuator fed by ambiguous signals</h4><p>Autoscaling adds or removes capacity based on signals like CPU usage, request rate, or queue length. During incidents those signals mislead. Queue depth may grow because a dependency is slow, not because demand has grown. CPU may spike from retries and timeouts, not real workload. Request rates may rise simply because retried calls look like new traffic. A reactive autoscaler then adds capacity in the wrong place, driving up cost while putting even more pressure on the real bottleneck. The result is a system that becomes less stable and more expensive at exactly the moment it needs to recover.</p><h4>Circuit breakers carry shock-wave potential</h4><p>Circuit breakers are necessary. Without system-level coordination they create abrupt traffic shocks. If multiple clients trip simultaneously, alternate paths overload and create secondary failures that appear unrelated to the original degradation.</p><h4>The common pattern is unbounded actuation in a coupled system</h4><p>Retries, autoscaling, and circuit breakers are not mistakes. They are essential reliability tools. Problems arise when they operate like independent reflexes, without coordination or limits. When nothing controls how often they act, how broadly their actions affect the system, how easily changes can be reversed, or how clearly decisions can be explained, they unintentionally amplify failures. Managing this was already difficult with manually written rules. It becomes even more critical when AI agents make or trigger operational decisions.</p><h3>Why agentic operations change the physics</h3><p>Traditional automation such as scripts, thresholds, and fixed policies is limited in scope and relatively legible. Agentic systems expand both capability and risk because they tend to:</p><ol><li><p><strong>Increase velocity:</strong> propose actions faster than humans can validate non-local effects.</p></li><li><p><strong>Increase scope:</strong> coordinate actions across services, clusters, and regions.</p></li><li><p><strong>Optimize proxies:</strong> Charles Goodhart warned in 1975 that a statistical regularity tends to collapse once it is used as a control target, a caution later popularized as &#8220;when a measure becomes a target, it ceases to be a good measure&#8221; (&#8221;<a href="https://www.semanticscholar.org/paper/Problems-of-Monetary-Management:-The-UK-Experience-Goodhart/0ae623749b30de53a39cf05813f5f3842e422c01">Problems of Monetary Management: The U.K. Experience</a>,&#8221; 1975). Optimizing latency or cost in isolation can degrade reliability unless constraints are explicit.</p></li><li><p><strong>Operate under uncertainty: </strong>incidents produce ambiguous signals, and partial telemetry invites misdiagnosis and overcorrection.</p></li><li><p><strong>Reduce legibility unless designed otherwise:</strong> without decision traces, you cannot reconstruct why the system acted or demonstrate compliance.</p></li></ol><p>So the risk is not AI in isolation. The risk is unbounded actuation under uncertainty. ARG is the discipline designed to make agentic operations governable.</p><h3>A practical lens on amplification</h3><p>To manage automation safely, teams need a clear way to recognize when helpful mechanisms start making things worse. Amplification is that lens. A reliability mechanism helps when it improves outcomes without adding too much extra load. It becomes harmful when it multiplies load, volatility, or system-wide effects beyond what the platform can handle. In practice, retries multiply request volume, autoscaling multiplies capacity changes and cost, traffic shifts increase pressure on dependencies, and aggressive fixes create configuration churn. ARG is essentially the discipline of keeping these amplification effects under control.</p><h3>Autonomic Reliability Governance governs the automation itself</h3><p>The notion that systems can manage themselves is not new. Jeffrey Kephart and David Chess articulated that aspiration more than two decades ago in <a href="https://doi.org/10.1109/MC.2003.1160055">&#8220;The Vision of Autonomic Computing&#8221;</a> (IEEE Computer, January 2003). What is new is the velocity and scope of actuation introduced by agentic AI. Autonomic Reliability Governance is the discipline of designing and operating bounded, auditable, and reversible automation, including agentic operational decision layers, under explicit constraints.</p><ul><li><p>SLO constraints: latency ceilings, availability targets, error budgets.</p></li><li><p>Cost constraints: budget caps, unit-economics bounds, runaway scaling prevention.</p></li><li><p>Compliance constraints: auditability of decisions and actions.</p></li></ul><p>ARG is not a product. It is an operating model, automation you can trust because it is governable. ARG treats operational actuation, human or agentic, as a first-class risk surface with explicit safety envelopes..</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SqBR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27f56d02-5b66-4090-a0c2-ebe54b7cb49a_960x588.svg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SqBR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27f56d02-5b66-4090-a0c2-ebe54b7cb49a_960x588.svg 424w, https://substackcdn.com/image/fetch/$s_!SqBR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27f56d02-5b66-4090-a0c2-ebe54b7cb49a_960x588.svg 848w, https://substackcdn.com/image/fetch/$s_!SqBR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27f56d02-5b66-4090-a0c2-ebe54b7cb49a_960x588.svg 1272w, https://substackcdn.com/image/fetch/$s_!SqBR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27f56d02-5b66-4090-a0c2-ebe54b7cb49a_960x588.svg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SqBR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27f56d02-5b66-4090-a0c2-ebe54b7cb49a_960x588.svg" width="1456" height="892" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/27f56d02-5b66-4090-a0c2-ebe54b7cb49a_960x588.svg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:892,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4364,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/svg+xml&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/204654996?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27f56d02-5b66-4090-a0c2-ebe54b7cb49a_960x588.svg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SqBR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27f56d02-5b66-4090-a0c2-ebe54b7cb49a_960x588.svg 424w, https://substackcdn.com/image/fetch/$s_!SqBR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27f56d02-5b66-4090-a0c2-ebe54b7cb49a_960x588.svg 848w, https://substackcdn.com/image/fetch/$s_!SqBR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27f56d02-5b66-4090-a0c2-ebe54b7cb49a_960x588.svg 1272w, https://substackcdn.com/image/fetch/$s_!SqBR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27f56d02-5b66-4090-a0c2-ebe54b7cb49a_960x588.svg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 1. The ARG stack. Observe to Decide to Act, with Audit and Reversibility underneath.</em></figcaption></figure></div><h3>Governance primitives that make agentic operations safe</h3><p>ARG becomes real when it is implemented through governance primitives that align with existing reliability practice such as SLOs, error budgets, and progressive rollout, while controlling actuation velocity and blast radius.</p><h4><strong>1) Actuation budgets rate-limit change, not only traffic</strong></h4><p>Most systems rate-limit user requests. Few rate-limit operational change. ARG introduces actuation budgets, limits on how frequently high-impact actions can occur. Scaling shifts, retry escalations, routing changes, and configuration flips consume tokens from a constrained budget. When the budget is exhausted, automation may still observe and recommend, but it cannot repeatedly perturb the system. When instability rises, slow the actuators before adding more intelligence.</p><h4><strong>2) Blast radius governance scopes every action</strong></h4><p>The difference between a safe fix and an outage multiplier is often scope. ARG constrains blast radius using canaries, segmentation across cells and regions, tiered permissions, and scoped rollouts with automatic halt conditions. Even correct actions can be wrong if applied globally in one step. Autonomy must always be localized before it is generalized.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/p/autonomic-governance-for-agentic-systems?open=false#%C2%A7governance-primitives-that-make-agentic-operations-safe&quot;,&quot;text&quot;:&quot;Continue reading&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://deepengineering.net/p/autonomic-governance-for-agentic-systems?open=false#%C2%A7governance-primitives-that-make-agentic-operations-safe"><span>Continue reading</span></a></p><p></p><blockquote><p><em><strong>Continue reading </strong>- the seven governance primitives that turn ARG from principle into practice, how to validate governed automation without touching production or private data, and why agentic operations win by becoming more governable, not more intelligent.</em></p></blockquote><div><hr></div><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;542bb578-1f35-4bcf-a9d4-3d39eca702eb&quot;,&quot;caption&quot;:&quot;Sibasis Padhi on the discipline of bounding, auditing, and reversing agentic operational decisions so automation dampens incidents instead of amplifying them.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Autonomic Governance for Agentic Systems&quot;,&quot;publishedBylines&quot;:[],&quot;post_date&quot;:&quot;2026-07-01T18:00:00.085Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d182b3af-ffd3-4493-aa7b-1a141df2846d_1200x460.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.net/p/autonomic-governance-for-agentic-systems&quot;,&quot;section_name&quot;:&quot;Practical Deep-Dives&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:204644229,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div><hr></div><h1><strong>&#128736;&#65039; Tool of the Week</strong></h1><p><strong><a href="https://github.com/argoproj/argo-rollouts">Argo Rollouts</a> </strong>- a progressive delivery controller for Kubernetes</p><p><strong>Highlights</strong>:</p><ul><li><p>Bounds blast radius by exposing a change to a small traffic slice first, closing the gap that standard rolling updates leave open where nothing limits exposure or triggers an automated rollback on failure.</p></li><li><p>Makes reversibility the default, since a rollback shifts traffic back to the previous version instead of triggering a fresh deploy.</p></li><li><p>Gates promotion on evidence rather than a timer, aborting automatically when success rate or latency crosses a threshold, which mirrors the confidence gates in the feature.</p></li><li><p>Fits GitOps workflows, where an automatic rollback surfaces as divergence from declared state and prompts investigation before anyone retries.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://github.com/argoproj/argo-rollouts&quot;,&quot;text&quot;:&quot;Learn more about Argo Rollouts&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://github.com/argoproj/argo-rollouts"><span>Learn more about Argo Rollouts</span></a></p><div><hr></div><h1><strong>&#128206; Tech Briefs</strong></h1><ul><li><p><a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/release-notes">Google makes Model Armor generally available for Agent Gateway</a> - Model Armor reaches general availability for Agent Gateway, screening agent prompts and responses against content guardrails.</p></li><li><p><a href="https://www.datadoghq.com/blog/dash-2026-new-feature-roundup-keynote/">Datadog extends Bits AI to autonomous remediation at DASH</a> - Bits AI now detects and remediates autonomously within predefined guardrails, and Agent Console tracks agent actions.</p></li><li><p><a href="https://www.cncf.io/announcements/2026/05/21/cloud-native-computing-foundation-announces-opentelemetrys-graduation-solidifying-status-as-the-de-facto-observability-standard/">OpenTelemetry graduates from the CNCF</a> - OpenTelemetry graduates from the CNCF, cementing the vendor-neutral telemetry standard that grounds agent decisions in trusted context.</p></li><li><p><a href="https://techcommunity.microsoft.com/blog/appsonazureblog/announcing-general-availability-for-the-azure-sre-agent/4500682">Microsoft makes Azure SRE Agent generally available</a> - Azure SRE Agent is generally available, with Review and Autonomous run modes gating remediation behind human approval.</p></li><li><p><a href="https://grafana.com/blog/ai-observability-for-agents-in-grafana-cloud/">Grafana previews AI Observability in Grafana Cloud</a> - Grafana previews AI Observability, treating agent sessions as first-class telemetry and alerting on policy violations and anomalies.</p></li></ul><div><hr></div><p><span>That&#8217;s all for today. Thank you for reading this issue of Deep Engineering.</span></p><p><span>We&#8217;ll be back next week with more expert-led content.</span></p><p><span>Keep building,</span></p><p><span>Saqib Jan</span></p><p><span>Editor-in-Chief, Deep Engineering</span></p><div><hr></div><p><em><span>If your company wants to reach senior developers, software engineers, and technical decision-makers, </span><a href="https://packt.omeclk.com/portal/wts/uc%5EcnN2dfNaqmD-kB-mo66%7C7g%5Ef%7Cb"><span>speak to us about partnering</span></a><span> with Deep Engineering.</span></em></p>]]></content:encoded></item></channel></rss>