<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Packt Deep Engineering: Engineering Leadership]]></title><description><![CDATA[Deep Engineering talks to a lot of practitioners. Engineering Leadership is where we bring those conversations together to find the patterns that no single interview can surface. These are not interviews. They are the conclusions that the interviews make possible.]]></description><link>https://deepengineering.net/s/engineering-leadership</link><image><url>https://substackcdn.com/image/fetch/$s_!H5BJ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png</url><title>Packt Deep Engineering: Engineering Leadership</title><link>https://deepengineering.net/s/engineering-leadership</link></image><generator>Substack</generator><lastBuildDate>Sat, 26 Sep 2026 21:45:16 GMT</lastBuildDate><atom:link href="https://deepengineering.net/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Packt]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[deepengineering@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[deepengineering@substack.com]]></itunes:email><itunes:name><![CDATA[Packt]]></itunes:name></itunes:owner><itunes:author><![CDATA[Packt]]></itunes:author><googleplay:owner><![CDATA[deepengineering@substack.com]]></googleplay:owner><googleplay:email><![CDATA[deepengineering@substack.com]]></googleplay:email><googleplay:author><![CDATA[Packt]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[EY Bounds Agent Autonomy by Risk, Not by Capability]]></title><description><![CDATA[Dipanjan Sengupta, EY Distinguished Technologist, on the boundary that survives model churn and the operating model that gets AI past the pilot.]]></description><link>https://deepengineering.net/p/ey-agent-autonomy-governance-dipanjan-sengupta</link><guid isPermaLink="false">https://deepengineering.net/p/ey-agent-autonomy-governance-dipanjan-sengupta</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Tue, 22 Sep 2026 20:06:31 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/25db609b-c9e8-4106-a304-45687bd1c87d_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qDZd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89e3410a-5abd-4782-85da-ea2036b449d1_2400x1200.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qDZd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89e3410a-5abd-4782-85da-ea2036b449d1_2400x1200.png 424w, https://substackcdn.com/image/fetch/$s_!qDZd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89e3410a-5abd-4782-85da-ea2036b449d1_2400x1200.png 848w, https://substackcdn.com/image/fetch/$s_!qDZd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89e3410a-5abd-4782-85da-ea2036b449d1_2400x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!qDZd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89e3410a-5abd-4782-85da-ea2036b449d1_2400x1200.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qDZd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89e3410a-5abd-4782-85da-ea2036b449d1_2400x1200.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/89e3410a-5abd-4782-85da-ea2036b449d1_2400x1200.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:928948,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/216963352?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89e3410a-5abd-4782-85da-ea2036b449d1_2400x1200.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qDZd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89e3410a-5abd-4782-85da-ea2036b449d1_2400x1200.png 424w, https://substackcdn.com/image/fetch/$s_!qDZd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89e3410a-5abd-4782-85da-ea2036b449d1_2400x1200.png 848w, https://substackcdn.com/image/fetch/$s_!qDZd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89e3410a-5abd-4782-85da-ea2036b449d1_2400x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!qDZd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89e3410a-5abd-4782-85da-ea2036b449d1_2400x1200.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Most enterprise teams set the limits of agent autonomy by testing what the agent can do. They benchmark it, track the error rate across a few release cycles, and widen the scope as the numbers improve. The capability curve quietly becomes the permission curve, and nobody asks what an agent should be allowed to decide separately from what it can decide correctly.</p><p><a href="https://in.linkedin.com/in/dipanjan-sengupta-52aa716">Dipanjan Sengupta</a>, EY Distinguished Technologist and AI Engineering Leader for <a href="https://www.ey.com/">EY Consulting Global Delivery Services</a>, argues that this reverses the order of the decision. &#8220;In our experience, the boundary between agent-driven and human-driven decision-making is determined not by technical capability, but by risk, accountability, and business impact,&#8221; he reasons.</p><p>That position asks more of a leader than it first appears, because a better model earns an agent no new territory under it. The line gets drawn once, in terms that outlast whichever model you deploy this quarter.</p><h2>Reversibility marks the line, not capability</h2><p>Sengupta puts a specific class of work on the agent side of the boundary. &#8220;Agents are well suited for routine, low-risk, and reversible decisions such as data classification, document summarization, code suggestions, test generation, or workflow orchestration,&#8221; he explains. The other side of the line carries a different property entirely. &#8220;Humans remain responsible for high-impact decisions involving regulatory compliance, security, financial commitments, production releases, architectural trade-offs, and customer-facing outcomes.&#8221;</p><p>The test underneath both lists is reversibility. &#8220;Human oversight is particularly critical when decisions are irreversible or carry significant operational consequences,&#8221; he underscores. Reversibility belongs to the decision rather than to the system making it, which is exactly what gives the boundary its durability. A model upgrade does not move a production release from the irreversible column into the reversible one, so the classification survives the next six months of model churn without anyone renegotiating it.</p><p>The supervision pattern EY adopts follows from that classification rather than from an approval queue sitting in front of every action. &#8220;As a result, we increasingly adopt a &#8216;human-on-the-loop&#8217; model,&#8221; Sengupta shares. &#8220;Agents execute, recommend, and learn, while humans govern, approve, and intervene when required.&#8221; Agents still carry substantial work across requirements analysis, code generation, testing, documentation, knowledge retrieval, incident investigation, and operational monitoring. &#8220;Rather than replacing engineers, agents act as force multipliers, accelerating repetitive, information-intensive tasks while enabling teams to focus on architecture, innovation, and business outcomes,&#8221; he adds.</p><p>For a leader turning that into practice, the useful artifact is a written register of the decision types agents touch in your estate, with each entry marked reversible or not and carrying the name of the person accountable when an agent acts inside it. Build the register before the next expansion of scope rather than after an incident forces it, and review it when the work changes rather than when the model changes. An undocumented boundary widens on its own, one reasonable-looking exception at a time, and the register gives a manager something to point at when a team asks for more autonomy than the decision class warrants.</p><h2>Modernized platforms still stall before production</h2><p>Most organizations never reach the boundary argument because their pilots stop moving well before it becomes urgent, and Sengupta locates the cause somewhere other than where most postmortems put it. &#8220;The breakdown rarely occurs because of model limitations,&#8221; he points out. &#8220;It typically occurs because the surrounding ecosystem is not designed to operationalize AI at scale.&#8221;</p><p>Data readiness accounts for a large share of it. &#8220;Many enterprises continue to operate with fragmented, inconsistent, or poorly governed data landscapes,&#8221; he observes, and the gap only becomes visible once real traffic arrives. &#8220;AI models may perform well during experimentation, but production deployments expose issues related to data quality, lineage, ownership, freshness, and access control.&#8221; A pilot forgives all five of those because it operates on a curated corpus that somebody cleaned by hand, and that cleaning never appears in the cost of the next ten use cases.</p><p>The second gap compounds with every additional deployment. &#8220;Reliable AI deployment requires versioning, automated testing, CI/CD pipelines, monitoring, drift detection, retraining workflows, and end-to-end lifecycle management,&#8221; Sengupta explains. Read that list against the boundary argument and the two connect directly, because lineage, monitoring, and drift detection are what let a leader demonstrate that agents stayed inside the reversible column. Without them the boundary exists as policy rather than as something the organization can evidence after the fact.</p><p>The practical move here costs a week and saves a quarter. Take the five properties Sengupta lists and run them against the actual data sources your next pilot will depend on, then treat a source with no named owner as a blocker rather than a caveat in the risk register. Do the same with the operational list before the pilot ships instead of after, because versioning and drift detection retrofitted into a running system usually means rebuilding the deployment path rather than extending it.</p><h2>Governance retrofitted late becomes rework</h2><p>Data and operations gaps at least announce themselves through poor results. Governance arrives on a quieter schedule and costs more when it does. &#8220;Security, compliance, responsible AI controls, auditability, and model risk management are often introduced late in the adoption journey,&#8221; Sengupta warns. &#8220;By then, teams must retrofit controls into architectures that were never designed for enterprise-scale governance.&#8221;</p><p>Retrofitting carries a bill that rarely appears in the business case for a pilot, because the architecture that made the pilot fast is usually the same architecture that makes the controls hard. Direct database access, a single service account, and no request-level audit trail all accelerate a proof of concept and all have to be unwound before anything touches regulated data. The organizational version of the problem arrives alongside the technical one. &#8220;Data scientists, platform engineers, compliance teams, and business stakeholders frequently operate in silos, creating friction between experimentation and productionizing,&#8221; he notes. Each group optimizes for its own gate, and the handoff between experimentation and production becomes the place where the work quietly stops.</p><p>Put a security and compliance reviewer inside the pilot team from the first sprint, with one specific job, which is to write down the controls production will demand while the architecture remains cheap to change. That reviewer costs a few hours a week early and saves a rebuild later. Where a control cannot be implemented yet, record it as a known debt with a date attached rather than discovering it during a pre-production review, because a documented gap gets funded and an undocumented one gets argued about.</p><h2>Interoperability decides whether AI crosses boundaries</h2><p>Controls inside one organization handle only part of the problem, because enterprise AI rarely stops at the company boundary. Sengupta has contributed to industry standards work in integration and interoperability, and he draws a consistent lesson from it. &#8220;The most valuable AI systems are rarely standalone systems,&#8221; he puts it. Enterprise value arrives when AI operates across platforms, business functions, partners, suppliers, and regulatory environments, and crossing those lines introduces problems harder than the models themselves.</p><p>His principle is that &#8220;openness and standardization drive scalability,&#8221; and the architectural consequence gets specific quickly. &#8220;Systems that expose clear contracts, support discoverability, and maintain robust auditability are significantly easier to integrate and govern across organizations,&#8221; he explains. The same three properties that make a system integrable make it governable, which is why the interoperability question and the autonomy question resolve together rather than separately. &#8220;Interoperability is not simply a technical challenge,&#8221; he contends. &#8220;It is also a trust challenge. Organizations must be confident about data lineage, access controls, explainability, accountability, and compliance before AI systems can collaborate effectively.&#8221;</p><p>For teams designing against a moving target, he offers a rule worth carrying into architecture review. &#8220;Resilience comes from abstraction rather than dependency,&#8221; he says. &#8220;AI architectures designed around open standards, modular components, and loosely coupled services are better equipped to adapt as models, platforms, and regulations evolve.&#8221;</p><p>Apply that as a test on every integration decision. Ask whether you could replace the component in twelve months without renegotiating with the partner on the other side of the interface, and where the answer is no, put a documented contract between your system and theirs before the coupling hardens. Publish the interface descriptions and the audit surface as first-class artifacts rather than as documentation debt, since those are the things a partner or a regulator will ask for, and producing them on demand takes far longer than maintaining them as you go.</p><h2>Platform thinking replaces project thinking</h2><p>Every fix above becomes expensive when a team does it once per project, which is what makes the operating model the real unit of change. &#8220;In my experience, successful AI modernization requires a shift from project thinking to platform thinking,&#8221; Sengupta maintains. &#8220;Enterprises need shared foundations for data, governance, observability, evaluation, and reusable AI services.&#8221;</p><p>Under project thinking, the tenth use case costs roughly what the first one did, because each team rebuilds the evaluation harness and rediscovers which controls production demands. Under platform thinking, the tenth use case inherits the controls, the harness, and the audit trail from everything built before it, and the reversibility boundary becomes a property of the platform rather than a policy each team reinterprets for itself. &#8220;The organizations that scale AI effectively view cloud infrastructure as the starting point, not the destination,&#8221; Sengupta says. &#8220;Sustainable success comes from building a disciplined operating model where data, engineering, governance, and business objectives evolve together rather than independently.&#8221;</p><p>Two moves carry the most weight for a leader deciding where to put effort this quarter. Classify every decision your agents currently touch by reversibility and blast radius rather than by how well the model performs on it, and give that register an owner so it stays current as scope grows. Then move audit, lineage, and evaluation out of individual projects and into the platform layer, so the next team inherits controls instead of rebuilding them under deadline. Both are organizational decisions rather than technical ones, which is why they need a leader to make them and why they rarely emerge from a delivery team working alone.</p><p>The framing Sengupta closes on is worth keeping in front of any team drawing these lines. &#8220;The future is not about choosing between human and artificial intelligence,&#8221; he says. &#8220;It is about designing trusted human-AI systems where agents handle scale and speed, and humans provide judgment, context, ethics, and accountability.&#8221;</p>]]></content:encoded></item><item><title><![CDATA[Engineering Teams Are Paying Back Their AI Speed Gains With Interest on Review]]></title><description><![CDATA[Checksum CEO Gal Vered on AI code debt, review cycles growing 25% or more, and why verification has to run before a human ever opens the pull request.]]></description><link>https://deepengineering.net/p/ai-speed-gains-review-burden</link><guid isPermaLink="false">https://deepengineering.net/p/ai-speed-gains-review-burden</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Mon, 31 Aug 2026 17:32:59 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b17ed8ff-a9f1-4175-a1ca-a49c45899d04_2400x1200.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Two years ago the difficult part of AI-assisted development was getting a model to produce code that looked correct, and for many teams that is no longer the main constraint, with agents now opening pull requests faster than their review processes can absorb. What replaced it is a harder and less visible problem that begins the moment the code exists and someone has to decide whether to trust it.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Vjsb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde4c1bdf-a228-4c26-a859-beef9f8081fa_2400x1000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Vjsb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde4c1bdf-a228-4c26-a859-beef9f8081fa_2400x1000.png 424w, https://substackcdn.com/image/fetch/$s_!Vjsb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde4c1bdf-a228-4c26-a859-beef9f8081fa_2400x1000.png 848w, https://substackcdn.com/image/fetch/$s_!Vjsb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde4c1bdf-a228-4c26-a859-beef9f8081fa_2400x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!Vjsb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde4c1bdf-a228-4c26-a859-beef9f8081fa_2400x1000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Vjsb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde4c1bdf-a228-4c26-a859-beef9f8081fa_2400x1000.png" width="1456" height="607" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/de4c1bdf-a228-4c26-a859-beef9f8081fa_2400x1000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:607,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:553654,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/213550519?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde4c1bdf-a228-4c26-a859-beef9f8081fa_2400x1000.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Vjsb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde4c1bdf-a228-4c26-a859-beef9f8081fa_2400x1000.png 424w, https://substackcdn.com/image/fetch/$s_!Vjsb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde4c1bdf-a228-4c26-a859-beef9f8081fa_2400x1000.png 848w, https://substackcdn.com/image/fetch/$s_!Vjsb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde4c1bdf-a228-4c26-a859-beef9f8081fa_2400x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!Vjsb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde4c1bdf-a228-4c26-a859-beef9f8081fa_2400x1000.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>&#8220;The part nobody solved is what happens after the code is written,&#8221; says</span><a href="https://www.linkedin.com/in/gal-vered"><span> Gal Vered</span></a><span>, CEO and co-founder of</span><a href="https://checksum.ai/"><span> Checksum</span></a><span>, whose work with engineering teams centers on the testing infrastructure that determines whether generated code holds up against a real database, a rate-limited API, and a permission system nobody documented.</span></p><p><span>Checksum&#8217;s research shows how wide that gap has become. Over the previous 90 days, 61% of surveyed engineering leaders reported a production incident caused by AI-generated code, and 74.3% had rolled back an AI change that their own unit tests never flagged. These findings do not describe teams that simply skipped every safeguard, because respondents reported writing tests, conducting code review, and using AI tools to check AI output, which suggests the problem lies in the design of the verification loop rather than in the diligence of the people operating it.</span></p><h2><span>AI code debt looks correct and passes every check you have</span></h2><div class="pullquote"><p><em><span>&#8220;That spread is the signature of a blind spot, not a bug type.&#8221;<br></span></em><span>Gal Vered</span></p></div><p><span>Vered draws a distinction that changes what teams think they are accumulating when they scale up generation.</span></p><p><span>&#8220;It&#8217;s not sloppy code,&#8221; he says of what he calls AI code debt. &#8220;It&#8217;s code that looks completely correct and passes every check you have, but was never tested against the conditions it will actually run in.&#8221; The model wrote it without ever seeing the production database, the third-party rate limits, or the business logic that lives only in a senior engineer&#8217;s head, so it can pass the available checks and still fail under the conditions that determine whether the system works.</span></p><p><span>The root-cause data from the same research supports that framing because the incidents did not cluster around a single failure category. When Checksum asked leaders what caused their most recent AI-related incident, no single cause dominated, and performance issues at scale, logic errors, and integration failures the model could not have anticipated all clustered together in the high teens to low twenties.</span></p><p><span>For Vered, that flat distribution is the finding rather than a gap in the data, because it points to a blind spot rather than a single bug type. &#8220;You don&#8217;t fix a blind spot by writing more unit tests,&#8221; he says. &#8220;You fix it by giving the agent visibility into the environment before the code ships, not after.&#8221; Vered uses the CrowdStrike and Cloudflare incidents as examples of the same pattern at scale, where individually tested components can still fail through their interaction at runtime.</span></p><p><span>For leaders who want to act on that reading rather than file it away, one practical move is to change how incidents get classified after the fact. Instead of sorting them only by bug type, sort them by whether the failing condition was observable anywhere in the pre-merge environment, because that distinction shows whether the team needs another test or a more realistic verification environment.</span></p><h2><span>Review time absorbs the volume gains from AI code generation</span></h2><p><span>The workflow most teams followed a few years ago moved from writing code to reviewing it to shipping it, and the current one looks closer to prompting, generating, reviewing, re-prompting, and reviewing again. Review never went away under that shift, because it quietly absorbed the hours that used to go into writing along with a share of the hours that were never budgeted anywhere.</span></p><p><span>Checksum&#8217;s data puts the cost in plain terms, with 64.8% of leaders reporting that AI-generated code takes more review time than human-written code rather than less, and half reporting that their review cycles have grown by 25% or more since adopting AI coding tools. Faros AI&#8217;s telemetry, which Vered cites, reports that teams with high AI adoption merge 98% more pull requests while review time on those pull requests rises 91%.</span></p><p><span>&#8220;The volume gains on the writing side are getting paid back with interest on the review side,&#8221; Vered says.</span></p><p><span>Adding reviewers is the obvious response, and the same research suggests many leaders do not think it will hold, because only 28.6% believe they could hire their way out of the review burden and a comparable share describe it as a structural problem that headcount cannot solve. Those responses reflect a review task that differs from reading a colleague&#8217;s pull request, because engineers must search for subtle mistakes inside code that often appears locally plausible.</span></p><p><span>What works instead is a change in sequence rather than staffing, and Vered describes the target state in terms most teams will recognize from their own product usage. He expects the same prompt, generate, and verify loop used in vibe-coded applications to shape enterprise software within the next 12 months, with verification becoming a simulation layer that tests the application under realistic conditions before it ships.</span></p><p><span>Vered argues that teams getting this right move verification ahead of the human, so that by the time an engineer opens a diff the automated checks have already covered the mechanical questions and review can focus on design and intent. Practically, that means setting a rule about ordering, where no pull request reaches a human reviewer until automated verification has completed and reported what it found.</span></p><h2><span>Coding agents cannot see what actually decides runtime behavior</span></h2><p><span>Underneath both the debt and the review tax is a visibility gap that Vered calls the &#8220;Context Void,&#8221; and naming it matters because it changes what kind of problem leaders think they are solving.</span></p><p><span>A coding agent sees code, but it does not see database state, API behavior under load, the permission system, feature flags, or the actual shape of production traffic. Everything that determines runtime behavior remains outside what the agent can observe, and Vered is explicit that this is structural rather than a temporary limitation that better models will close, because a more capable model still cannot reason about a system it was never shown.</span></p><p><span>That is where his comparison to autonomous vehicles becomes more than a convenient analogy. &#8220;Nobody would put a self-driving car on the road without a world model,&#8221; he says. &#8220;You don&#8217;t make the driving model safe by making it smarter in isolation, you give it millions of simulated miles to practice on first.&#8221;</span></p><p><span>In Vered&#8217;s comparison, autonomy in the physical world depends on simulation rather than trust in the model alone, with varying weather, traffic, pedestrian behavior, and sensor noise rehearsed before anything touches a public road.</span></p><p><span>Vered argues software carries the harder version of that problem, since a production system&#8217;s state space, meaning every combination of configuration, data, and timing, is arguably larger than what a car encounters on a city block and cannot be exhaustively enumerated. Because teams cannot exhaustively enumerate that state space, they have to simulate representative conditions, which is why Vered expects the next few years of software development to resemble autonomous vehicle development more than the review-heavy process most teams use today.</span></p><h2><span>Verification belongs in infrastructure rather than at the end of the pipeline</span></h2><p><span>Vered&#8217;s central recommendation to engineering leaders is to reclassify verification as infrastructure, because that changes both where it operates and what gets funded. That means making it as permanent and automatic as a CI pipeline, because most teams currently have AI writing code and a patchwork of humans and point tools trying to catch what it missed, and that patchwork does not scale as generated-code volume grows.</span></p><p><span>The concrete starting point is an audit of which stage each existing check occupies and what environment it can see. Unit tests, code review, and security scanning all earn their place, and Checksum&#8217;s research found that AI unit-test generation, which Vered describes as the most direct counterweight to AI-written code, remains the least adopted of the major verification categories at 48.6%. Adoption is only half the question, because none of those layers can see what production sees, and a team can raise coverage across all of them while leaving the actual blind spot untouched.</span></p><p><span>The standard Vered suggests leaders hold themselves to is the one they already apply to their toolchain without thinking about it, which is trusting code the way they trust a compiler, not by reading every line but by trusting the verification underneath it. Reaching that bar is what makes the volume sustainable, and teams that set it now will get there before the rate of generated code outruns anyone&#8217;s ability to review it by hand.</span></p><h2><span>Three stages take a team from manual QA to simulation-driven validation</span></h2><p><span>For teams relying on manual QA today and looking at simulation-driven validation as the destination, Vered describes a sequence rather than a leap, and each stage produces value on its own.</span></p><p><span>The first stage brings verification inside the loop the AI already uses for writing code, with tests generated and executed automatically on every pull request and targeted to what actually changed, so that nothing merges on the strength of code review alone. Teams that complete only this stage can still reduce failures caused by approving a diff that nobody executed.</span></p><p><span>The second stage moves from testing code in isolation to testing it against production-like conditions, meaning real data shapes, real API behavior, and real load rather than mocks. Checksum&#8217;s research shows that 69.5% of teams already verify against production-like conditions in some form, although those checks often remain bolted onto a process that cannot fully see the interactions between systems where the expensive failures emerge.</span></p><p><span>The third stage closes the loop so that the agent receives more than a pass or a fail. It gets told what broke and why, in terms it can act on, so it can fix the problem and re-verify without a human in the middle of every cycle. The payoff at that point is not only a lower incident count, because the larger return appears in how the team uses its most expensive engineering hours.</span></p><p><span>&#8220;They&#8217;ll get their senior engineers back,&#8221; Vered says, &#8220;because those are the people currently absorbing the gap by hand.&#8221;</span></p><div><hr></div><p><em><a href="https://www.linkedin.com/in/gal-vered">Gal Vered</a> is CEO and co-founder of <a href="https://checksum.ai/">Checksum</a>, which builds AI-generated end-to-end Cypress and Playwright tests. This is not a sponsored article. Research cited in this article was conducted by Checksum unless otherwise noted.</em></p>]]></content:encoded></item><item><title><![CDATA[Forward Deployed Engineer Jobs Are Out There. Hiring Is Harder Than The Postings Suggest]]></title><description><![CDATA[Forward deployed engineer postings rose 729 percent in April 2026. Three leaders explain when to hire FDEs, how to structure the team, and what holds good ones.]]></description><link>https://deepengineering.net/p/forward-deployed-engineer-jobs-hiring</link><guid isPermaLink="false">https://deepengineering.net/p/forward-deployed-engineer-jobs-hiring</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Tue, 18 Aug 2026 20:22:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/78dec1ee-5bce-4c13-83a7-c4e7f53dac88_2760x1096.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>As companies integrate advanced AI systems into core workflows, they need engineers who can turn persuasive pilots into reliable production systems. And so, a disciplined hiring approach should define the job, evaluate candidates, structure the team, and keep field learning connected to the product.</p><p><span>Forward deployed engineer postings were roughly </span><strong><span>729 percent</span></strong><span> higher in</span><strong><span> April 2026</span></strong><span> than a year earlier, according to an Indeed index reported by </span><a href="https://www.businessinsider.com/forward-deployed-engineer-jobs-in-demand-2026-5"><span>Business Insider</span></a><span>. While the figure covers one job board rather than the whole market, the direction is hard to ignore.</span></p><p><span>Anthropic, OpenAI, Palantir, Stripe, and Google Cloud have all recruited for this role as AI vendors push beyond model access and into enterprise deployment. </span><a href="https://openai.com/careers/forward-deployed-engineer-%28fde%29-sf-san-francisco/"><span>OpenAI&#8217;s current San Francisco role</span></a><span> covers discovery through production rollout and measures success through adoption, workflow impact, and feedback that reaches product and model roadmaps.</span></p><p><span>That broad scope explains the demand, while it also exposes the risk created by weak job descriptions and vague accountability. Hiring teams often select polished customer engineers, strong coders who avoid ambiguity, or heroic generalists who burn out under responsibilities that should belong to a team.</span></p><p><span>But the strongest operating models define a production outcome, test candidates inside realistic uncertainty, and protect a formal path from customer exceptions back into the product. And they treat forward deployment as a repeatable engineering system rather than a collection of individual rescues performed by unusually resilient employees.</span></p><blockquote><p><strong>Deep Engineering</strong> spoke with <strong>three women</strong> leading AI delivery, data science, and forward deployed engineering to understand what the role owns and where its operating model fails. Their responses point toward a practical hiring model that values production adoption, technical judgment, customer context, and reusable product learning in equal measure.</p></blockquote><h2><span>Production adoption is what the role actually owns</span></h2><p><span>Useful FDE definitions converge on the same outcome because the engineer remains accountable until a system works inside the customer&#8217;s environment and changes a measurable workflow. Discovery meetings, architecture reviews, implementation support, and custom code all matter, but they describe activities rather than the business result that justifies the role.</span></p><p><a href="https://www.linkedin.com/in/ritikasingh"><span>Ritika Singh</span></a><span>, COO at </span><a href="https://www.datagol.ai/"><span>DataGOL</span></a><span>, focuses that accountability on the difficult distance between an impressive demonstration and an operating system that survives real data, permissions, compliance gates, undocumented processes, and stakeholders with conflicting incentives. &#8220;An FDE is not about a demo, not a signed pilot, not an architecture diagram everyone nods at in a conference room,&#8221; she says. Her framing gives hiring leaders a better first line for the job description than the familiar list of customer-facing responsibilities.</span></p><div class="pullquote"><p></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!x_Fj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe116e9ca-5143-4a5c-995b-901003ab73cf_200x200.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!x_Fj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe116e9ca-5143-4a5c-995b-901003ab73cf_200x200.jpeg 424w, https://substackcdn.com/image/fetch/$s_!x_Fj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe116e9ca-5143-4a5c-995b-901003ab73cf_200x200.jpeg 848w, https://substackcdn.com/image/fetch/$s_!x_Fj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe116e9ca-5143-4a5c-995b-901003ab73cf_200x200.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!x_Fj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe116e9ca-5143-4a5c-995b-901003ab73cf_200x200.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!x_Fj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe116e9ca-5143-4a5c-995b-901003ab73cf_200x200.jpeg" width="256" height="256" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e116e9ca-5143-4a5c-995b-901003ab73cf_200x200.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:200,&quot;width&quot;:200,&quot;resizeWidth&quot;:256,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Ritika Singh - DataGOL | LinkedIn&quot;,&quot;title&quot;:&quot;Ritika Singh - DataGOL | LinkedIn&quot;,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Ritika Singh - DataGOL | LinkedIn" title="Ritika Singh - DataGOL | LinkedIn" srcset="https://substackcdn.com/image/fetch/$s_!x_Fj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe116e9ca-5143-4a5c-995b-901003ab73cf_200x200.jpeg 424w, https://substackcdn.com/image/fetch/$s_!x_Fj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe116e9ca-5143-4a5c-995b-901003ab73cf_200x200.jpeg 848w, https://substackcdn.com/image/fetch/$s_!x_Fj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe116e9ca-5143-4a5c-995b-901003ab73cf_200x200.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!x_Fj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe116e9ca-5143-4a5c-995b-901003ab73cf_200x200.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://www.linkedin.com/in/ritikasingh"><span>Ritika Singh</span></a></strong><span><br>COO at </span><a href="https://www.datagol.ai/"><span>DataGOL</span></a></p><p><span>&#8220;The problem with most FDE job descriptions is that they list activities, not accountability.&#8221;</span></p></div><p><span>&#8220;Shipping without extracting the pattern makes you a very expensive contractor. Extracting patterns without shipping makes you an analyst,&#8221; Singh explains. Because that balance matters, leaders should translate production adoption into one customer metric and one engineering metric before opening the role for recruitment. The customer metric might track cycle time, exception rate, revenue recovery, or user adoption, while the engineering metric should show whether reliability, evaluation coverage, deployment speed, or reuse improves across engagements.</span></p><p><span>Write the success line before writing the qualifications, and make it concrete enough that a candidate can explain how they would establish a baseline, Singh recommends. &#8220;A strong version says the engineer owns production adoption and measurable workflow impact through rollout, while a weak version merely promises exposure to strategic customers and frontier technology.&#8221;</span></p><h2><span>A new title covers an old operating gap</span></h2><p><span>Palantir popularized the forward deployed model by embedding technical teams with customers, but the current AI hiring wave has broadened the title across product companies, frontier laboratories, consultancies, and infrastructure vendors. That matters because many employers now use one label for several jobs with different incentives, customer loads, and definitions of finished work.</span></p><p><span>In many organizations, solutions engineering helps a buyer understand the product, validates a possible architecture, and reduces technical risk before a purchase. A builder-style FDE stays through production, contributes working code, owns adoption, and converts repeated exceptions into product capabilities that make later deployments faster and safer.</span></p><p><span>The practical distinction appears in the weekly calendar and the performance scorecard rather than the title printed on an offer. A role built around many accounts, demonstrations, technical qualification, and revenue-linked compensation behaves like pre-sales, while a role built around a few deep deployments, production ownership, and documented reuse behaves like forward deployed engineering.</span></p><p><span>Companies can operate either model successfully, but candidates and managers need an accurate description of the work before they commit. Renaming a solutions role without changing the customer load, decision rights, engineering expectations, or product feedback loop creates confusion for employees and weak results for customers.</span></p><div class="callout-block" data-callout="true"><p><strong>&#9889; <a href="https://www.eventbrite.co.uk/e/forward-deployed-engineering-fde-workshop-from-ai-demo-to-production-tickets-1996779962620?aff=deepeng">Forward Deployed Engineering Workshop</a></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mGbL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c7d369d-0fbc-480c-9f35-5f9963cb8fc5_1880x879.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mGbL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c7d369d-0fbc-480c-9f35-5f9963cb8fc5_1880x879.jpeg 424w, https://substackcdn.com/image/fetch/$s_!mGbL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c7d369d-0fbc-480c-9f35-5f9963cb8fc5_1880x879.jpeg 848w, https://substackcdn.com/image/fetch/$s_!mGbL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c7d369d-0fbc-480c-9f35-5f9963cb8fc5_1880x879.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!mGbL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c7d369d-0fbc-480c-9f35-5f9963cb8fc5_1880x879.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mGbL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c7d369d-0fbc-480c-9f35-5f9963cb8fc5_1880x879.jpeg" width="1880" height="879" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3c7d369d-0fbc-480c-9f35-5f9963cb8fc5_1880x879.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:879,&quot;width&quot;:1880,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:258363,&quot;alt&quot;:&quot;Forward Deployed Engineering (FDE) Workshop: From AI Demo to Production&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Forward Deployed Engineering (FDE) Workshop: From AI Demo to Production" title="Forward Deployed Engineering (FDE) Workshop: From AI Demo to Production" srcset="https://substackcdn.com/image/fetch/$s_!mGbL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c7d369d-0fbc-480c-9f35-5f9963cb8fc5_1880x879.jpeg 424w, https://substackcdn.com/image/fetch/$s_!mGbL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c7d369d-0fbc-480c-9f35-5f9963cb8fc5_1880x879.jpeg 848w, https://substackcdn.com/image/fetch/$s_!mGbL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c7d369d-0fbc-480c-9f35-5f9963cb8fc5_1880x879.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!mGbL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c7d369d-0fbc-480c-9f35-5f9963cb8fc5_1880x879.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://www.linkedin.com/in/keithbourne/">Keith Bourne</a></strong>, Forward Deployed AI Engineer at <strong>Tribe AI</strong>, and <strong><a href="https://www.linkedin.com/in/tanya-dixit-computer-vision/">Tanya Dixit</a></strong>, Forward Deployed Engineer at <strong>Google</strong>, will lead two live sessions. You will map the FDE skill stack, scope a 90-day agent deployment for a regulated customer, and defend it in a CISO hot seat.</p><p><strong>September</strong> <strong>19</strong> and <strong>20</strong>. <strong><a href="https://www.eventbrite.co.uk/e/forward-deployed-engineering-fde-workshop-from-ai-demo-to-production-tickets-1996779962620?aff=deepeng">Reserve your seat</a></strong></p></div><h2><span>Hire the function only when the economics support it</span></h2><p><span>Dedicated FDE headcount earns its place when strategic customers have genuinely different environments, repeated deployment friction keeps revealing useful product patterns, and the contract economics can support deep engineering attention. Without those conditions, a company usually needs clearer templates, stronger onboarding, better product defaults, or a conventional post-sales motion before it needs a new function.</span></p><p><span>Singh argues that teams should reserve FDE capacity for complex, high-value accounts instead of spreading expensive engineers across every customer. &#8220;Spread across every deal, the role burns out fast,&#8221; Singh reasons. &#8220;Applied surgically to the deals that actually move the company, it is one of the highest-leverage hires you can make.&#8221; Her threshold protects the role from becoming an unlimited customization queue, while preserving enough concentrated field exposure to uncover patterns that the core product team could not see from roadmap discussions alone.</span></p><p><span>A useful planning test compares the expected customer value with the fully loaded deployment cost, including travel, security reviews, support load, and productization time. The model improves only when later deployments reuse earlier work, so leaders should expect cycle time and custom code to decline as the function matures.</span></p><p><span>Travel and customer load belong in that economic model because they shape performance, retention, and the number of deployments one engineer can own. </span><a href="https://openai.com/careers/forward-deployed-engineer-%28fde%29-sf-san-francisco/"><span>OpenAI&#8217;s current posting</span></a><span> allows travel up to 50 percent, which shows why leaders must disclose the real operating rhythm instead of treating travel as a minor line near the end.</span></p><h2><span>Write the job around decisions and boundaries</span></h2><p><span>A strong job description opens with the production outcome and then states the ownership arc from discovery through stable rollout. It explains whether the engineer writes customer-specific code, changes the core product, carries an on-call obligation, manages the deployment plan, or relies on separate delivery and program leadership.</span></p><p><span>Next, define decision rights across scope, architecture, security, and customer commitments, because ambiguity becomes expensive when a deployment is already under pressure. Candidates should know which tradeoffs they can make independently, which decisions need customer approval, and who resolves conflict between a near-term delivery request and the long-term product direction.</span></p><p><span>Then expose the workload through expected travel, number of concurrent customers, engagement length, escalation coverage, and protected productization time. These details help serious candidates assess the job, while discouraging applicants who want customer visibility without the sustained engineering and operational responsibility that follows the sale.</span></p><p><span>State the non-goals with equal clarity, because FDEs should not become permanent support engineers, unbounded professional services, or convenient owners for every cross-functional gap. The job should end a one-off dependency by creating a reusable component, a tested integration pattern, a documented limitation, or a clear product decision.</span></p><h2><span>Interview for how agentic systems fail</span></h2><p><span>Once the role is clear on paper, the hiring loop must test whether candidates can diagnose the failures they will encounter inside a customer environment. Agentic AI makes that evaluation harder because a system can complete a task while taking the wrong path, using weak context, calling unnecessary tools, violating an approval boundary, or consuming more budget than the business value it creates. Traditional pass or fail testing cannot reveal enough about that behavior when permissions, exceptions, costs, and customer conditions keep changing.</span></p><p><a href="https://www.linkedin.com/in/1swaatii/"><span>Swati Tyagi</span></a><span>, Senior Manager in Data Science and AI/ML at </span><a href="https://www.tredence.com/"><span>Tredence</span></a><span>, argues that trace-level behavior now matters alongside final output quality. Her view expands the technical bar for FDE hiring beyond model familiarity and toward the observability, evaluation, permissions, and business controls required for reliable production operation.</span></p><div class="pullquote"><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!v113!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85a02479-a420-4d42-a0e2-2181e3a86079_1280x1019.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!v113!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85a02479-a420-4d42-a0e2-2181e3a86079_1280x1019.png 424w, https://substackcdn.com/image/fetch/$s_!v113!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85a02479-a420-4d42-a0e2-2181e3a86079_1280x1019.png 848w, https://substackcdn.com/image/fetch/$s_!v113!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85a02479-a420-4d42-a0e2-2181e3a86079_1280x1019.png 1272w, https://substackcdn.com/image/fetch/$s_!v113!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85a02479-a420-4d42-a0e2-2181e3a86079_1280x1019.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!v113!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85a02479-a420-4d42-a0e2-2181e3a86079_1280x1019.png" width="314" height="249.9734375" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/85a02479-a420-4d42-a0e2-2181e3a86079_1280x1019.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1019,&quot;width&quot;:1280,&quot;resizeWidth&quot;:314,&quot;bytes&quot;:1293607,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://deepengineering.net/i/211748893?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85a02479-a420-4d42-a0e2-2181e3a86079_1280x1019.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!v113!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85a02479-a420-4d42-a0e2-2181e3a86079_1280x1019.png 424w, https://substackcdn.com/image/fetch/$s_!v113!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85a02479-a420-4d42-a0e2-2181e3a86079_1280x1019.png 848w, https://substackcdn.com/image/fetch/$s_!v113!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85a02479-a420-4d42-a0e2-2181e3a86079_1280x1019.png 1272w, https://substackcdn.com/image/fetch/$s_!v113!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85a02479-a420-4d42-a0e2-2181e3a86079_1280x1019.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://www.linkedin.com/in/1swaatii/"><span>Swati Tyagi</span></a></strong><span><br>Senior Manager in Data Science and AI/ML at </span><a href="https://www.tredence.com/"><span>Tredence</span></a></p><p><em><span>&#8220;With agents, the system may technically be working while still choosing the wrong tool, using incomplete context, taking an inefficient path or consuming far more tokens than expected.&#8221;</span></em></p></div><p><span>&#8220;An automated refund agent, for example, may execute perfectly but still create financial leakage by approving the wrong refunds,&#8221; Tyagi says. She recommends replacing a generic AI coding screen with &#8220;a broken agent trace that shows plausible output, inefficient tool selection, a permissions flaw, and a business loss hidden behind acceptable aggregate metrics.&#8221; Ask the candidate to isolate the failure, define the missing instrumentation, propose an evaluation set, and describe the rollback or human review boundary. This exercise shows how the candidate investigates ambiguous behavior and whether that person can protect the business while repairing the underlying system.</span></p><p><span>And the strongest response will connect technical behavior to business consequences without treating either side as somebody else&#8217;s responsibility, while distinguishing a fix for this customer from a reusable improvement to the harness, evaluation framework, permission model, or product interface that prevents the same failure elsewhere.</span></p><h2><span>Build a unit instead of hiring a superhero</span></h2><p><span>Tyagi also warns that companies often write one job description for several people, combining customer discovery, domain expertise, AI engineering, data architecture, organizational change, and production ownership into a single heroic profile. &#8220;The scalable unit is the team, not the individual,&#8221; she cautions. Exceptional generalists exist, but a hiring plan that assumes every seat will hold one creates a fragile operating model and a narrow talent funnel.</span></p><p><span>Anthropic&#8217;s current hiring model offers a useful example as a separate </span><a href="https://www.anthropic.com/careers/jobs/5017903008"><span>Technical Deployment Lead</span></a><span> owns scoping, stakeholder management, value measurement, and delivery complexity while working alongside FDEs who build the technical solution. That separation keeps technical execution close to customer reality without forcing every engineer to carry the entire commercial and organizational burden alone.</span></p><p><span>A scalable deployment unit pairs the FDE with a product counterpart, a customer domain owner, and shared platform, security, or governance support that can move quickly. The exact composition can change by engagement, but every responsibility needs an explicit owner before the team enters a production environment.</span></p><p><span>Because the FDE remains the technical integrator closest to the customer, the team should not dilute that person&#8217;s ownership through endless handoffs. The surrounding unit exists to remove specialist bottlenecks and clarify decisions, while the FDE keeps the system, workflow, and customer outcome connected from discovery through adoption.</span></p><h2><span>Put the function near engineering and protect the product loop</span></h2><p><span>Reporting lines shape the product loop because incentives decide whether field learning becomes durable software or disappears inside delivery work. A revenue-only system naturally rewards closing the current engagement, while an engineering system can also reward reducing future deployment effort and improving the underlying product.</span></p><p><span>Singh (of DataGOL) places the function inside engineering or a tight engineering and product hybrid, with constant collaboration across go-to-market teams but technical leadership over goals and development standards. &#8220;The mistake is having their goals, comp, and reporting line resolve to revenue rather than to product,&#8221; she says. That structure protects customer urgency without turning FDEs into consultants whose success ends when the statement of work closes.</span></p><p><span>Keep the technical reporting line, then add shared objectives with product and a regular review where field patterns receive an explicit disposition. Every recurring exception should become a reusable component, a roadmap decision, a documented product boundary, or a deliberate services commitment, rather than an unresolved note in a customer channel.</span></p><p><span>Singh proposes roughly 70 to 80 percent of capacity on customer delivery and about 20 percent on productization. The exact ratio will vary, but a protected allocation matters because the product loop disappears whenever leaders treat reuse as optional work that can wait until customer pressure falls.</span></p><p><span>Measure that loop through deployment cycle time, reuse rate, recurring defect classes, custom code retired, and field-originated product changes that reach general availability. These metrics show whether forward deployment compounds into a better platform or simply accumulates expensive customer-specific debt behind a fashionable title.</span></p><h2><span>Design the first ninety days around one reusable win</span></h2><p><span>The first ninety days should prove the operating model rather than test how much ambiguity a new hire can absorb without help. Leaders need to provide customer access, technical context, product sponsorship, and a deployment with enough importance to matter but enough support to become a learning environment.</span></p><p><span>Singh describes strong early performers as people who separate the blocker a customer states from the constraint that actually prevents adoption, then ship one visible improvement before the relationship loses momentum. &#8220;Those are rarely the same thing,&#8221; she says. Misfires remain in discovery, stay distant from the customer, or produce throwaway work that cannot help the next deployment.</span></p><p><span>During the first thirty days, the new hire should map the customer workflow, establish baseline measures, document system boundaries, and identify the decisions that could block production. The manager should evaluate the quality of diagnosis and access gained, rather than rewarding a premature volume of code.</span></p><p><span>Between days thirty-one and sixty, the engineer should deliver one safe, visible win with evaluation, observability, rollback, and adoption responsibilities included. The result should improve a customer metric while producing enough technical evidence to guide the larger deployment plan.</span></p><p><span>By day ninety, the engineer should convert one lesson into a reusable asset and present the pattern to product and engineering leadership. That artifact might be an integration component, evaluation suite, deployment playbook, permission pattern, or product proposal with evidence from the customer environment.</span></p><h2><span>Make the career path visible before the first offer</span></h2><p><span>Retention depends on whether forward deployment expands an engineer&#8217;s authority or traps that person inside a permanent queue of exceptions. Candidates will accept demanding customer work when it builds technical range, domain credibility, product influence, and a visible path toward broader leadership.</span></p><p><a href="https://www.linkedin.com/in/jayeeta-putatunda/"><span>Jayeeta Putatunda</span></a><span>, Forward Deployed AI Engineering Lead at </span><a href="https://www.turing.com/">Turing</a><span>, sees the role as a durable response to the execution gap between controlled AI capability and enterprise production. At Turing, she also sees the career ceiling arriving when companies consume deployment effort without returning agency or product influence.</span></p><div class="pullquote"><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!n4k3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d5ad6e0-e16a-47cc-bf5f-989aed1366d4_400x400.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!n4k3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d5ad6e0-e16a-47cc-bf5f-989aed1366d4_400x400.jpeg 424w, https://substackcdn.com/image/fetch/$s_!n4k3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d5ad6e0-e16a-47cc-bf5f-989aed1366d4_400x400.jpeg 848w, https://substackcdn.com/image/fetch/$s_!n4k3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d5ad6e0-e16a-47cc-bf5f-989aed1366d4_400x400.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!n4k3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d5ad6e0-e16a-47cc-bf5f-989aed1366d4_400x400.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!n4k3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d5ad6e0-e16a-47cc-bf5f-989aed1366d4_400x400.jpeg" width="248" height="248" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3d5ad6e0-e16a-47cc-bf5f-989aed1366d4_400x400.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:400,&quot;width&quot;:400,&quot;resizeWidth&quot;:248,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Jayeeta Putatunda - AI Loves Data&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Jayeeta Putatunda - AI Loves Data" title="Jayeeta Putatunda - AI Loves Data" srcset="https://substackcdn.com/image/fetch/$s_!n4k3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d5ad6e0-e16a-47cc-bf5f-989aed1366d4_400x400.jpeg 424w, https://substackcdn.com/image/fetch/$s_!n4k3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d5ad6e0-e16a-47cc-bf5f-989aed1366d4_400x400.jpeg 848w, https://substackcdn.com/image/fetch/$s_!n4k3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d5ad6e0-e16a-47cc-bf5f-989aed1366d4_400x400.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!n4k3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d5ad6e0-e16a-47cc-bf5f-989aed1366d4_400x400.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://www.linkedin.com/in/jayeeta-putatunda/"><span>Jayeeta Putatunda</span></a></strong><span><br>Forward Deployed AI Engineering Lead at </span><a href="https://www.turing.com/"><span>Turing</span></a></p><p><em><span>&#8220;After two years, strong FDEs could move into product, platform engineering, vertical leadership, solution architecture, or larger deployment leadership. The ceiling would appear only when companies treat them as permanent exception handlers, moving from one bespoke proof of concept to another without giving them influence over the product.&#8221;</span></em></p></div><p><span>&#8220;What will hold good FDEs is agency, deeper domain ownership, a voice in the roadmap, and recognition for turning lessons from one client into capabilities that can benefit many,&#8221; Putatunda says. Turn that career map into explicit levels before recruiting the first team, with progression based on deployment complexity, reusable leverage, technical influence, and the ability to develop other engineers. But promotion should not depend on accepting more customers at once, because volume alone rewards the behavior that eventually damages quality and retention.</span></p><p><span>Compensation should recognize travel, incident responsibility, customer pressure, and the market value of engineers who can operate across technical and organizational boundaries. But autonomy, roadmap influence, recovery time, and movement into product or platform leadership will often determine whether strong people stay after the initial learning curve.</span></p><h2><span>A better hiring model starts with a narrower promise</span></h2><p><span>The FDE title may change as AI roles continue to evolve, but the operating gap will remain wherever a capable model meets a complicated customer environment. Companies still need engineers who can diagnose the real workflow, ship production software, guide adoption, and carry field evidence back into the product.</span></p><p><span>Engineering leaders should narrow the promise before expanding the headcount, because the role works when its outcome, customer scope, decision rights, technical bar, productization time, and career path are visible. That clarity produces better interviews, more honest offers, stronger deployments, and a healthier relationship between customer urgency and product quality.</span></p><p><span>And because the market is moving quickly, disciplined role design now creates an advantage that compensation alone cannot sustain. The companies that learn from every deployment will build stronger products, while the companies that celebrate individual heroics will keep paying for the same exception in different customer environments.</span></p><div><hr></div><p><strong>&#128227; Contribute to Deep Engineering</strong></p><p>If you lead a team, we would like to interview you and build an <a href="https://deepengineering.net/s/engineering-leadership">engineering leadership</a> feature around your story. If you are a senior engineer with something you have learned in production, pitch a <a href="https://deepengineering.net/s/practical-deep-dives">practical deep dive</a> under your byline.</p>]]></content:encoded></item><item><title><![CDATA[Materials Science Gets Real Value From Quantum Before Anything Else]]></title><description><![CDATA[Why materials science and small molecule chemistry reach real quantum value first, what higher fidelity simulation changes for drug discovery, and why optimization and cryptography wait.]]></description><link>https://deepengineering.net/p/materials-science-first-real-quantum-value</link><guid isPermaLink="false">https://deepengineering.net/p/materials-science-first-real-quantum-value</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 06 Aug 2026 13:48:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b91e798a-acfd-477e-94d1-3eff12314f55_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><em><span>By </span><a href="https://www.linkedin.com/in/shassinger"><span>Sebastian Hassinger</span></a><span>, author of</span><a href="https://www.packtpub.com/en-it/product/the-new-quantum-era-9781807787370"><span> </span></a><a href="https://www.packtpub.com/en-us/product/the-new-quantum-era-9781807787370"><span>The New Quantum Era</span></a><span> and former quantum lead at </span><strong><span>AWS</span></strong><span> and </span><strong><span>IBM</span></strong><span> | This piece is adapted from his live Deep Engineering session, </span><a href="https://deepengineering.net/p/quantum-computing-beyond-the-hype-sebastian-hassinger"><span>Quantum Computing Beyond the Hype</span></a><span>. Edited by</span><a href="https://substack.com/@saqibjan"><span> Saqib Jan</span></a></em></p></blockquote><p><span>People list the same five domains whenever they ask where quantum computing will actually pay off, so chemistry, materials, optimization, cryptography, and machine learning. </span><strong><span>Of those, materials science is closest to real value, with small molecule chemistry a close second</span></strong><span>, and small molecule chemistry is sometimes treated as a type of materials science anyway. Optimization, cryptography, and machine learning all need thousands of logical qubits, so those are considerably further off.</span></p><p><span>The reason materials leads is that materials science is condensed matter physics. It is about how atoms pack together into a lattice, into a crystalline structure, and how the material behaves as a result of that packing. All of that comes down to physics calculations, which is precisely the kind of work a quantum computer should do well. There is already interesting research going on around battery technology, using quantum information and quantum computing to help with designing and inventing new battery chemistries.</span></p><p><span>What sits behind that work is more ambitious. The aspiration is that if we can simulate materials precisely enough at sufficient scale, we can start creating designer materials. Maybe something much lighter for the same strength, or much stronger for the same weight. Maybe photosynthesis built into the material itself, so you coat your car in something that generates electricity without any cells on the outside. The imagination gets stimulated by the idea of engineering the attributes of a new material at atomic scale, and that is one of the genuinely underestimated parts of this whole story.</span></p><p><span>Chemistry follows for the same underlying reason, because chemistry is quantum mechanical at its core. It is the way atoms interact with one another, and that interaction is a quantum mechanical phenomenon happening at scale, which is what makes it so difficult to simulate precisely. People talk about small molecule chemistry as the next frontier beyond materials, and there is a natural leap from small molecule chemistry to pharmaceuticals, which tend to be larger and more complex molecules. You need a bigger machine to simulate them.</span></p><p><span>Provided we can build one large enough, you can easily imagine drug discovery and research happening inside a very high fidelity simulation, with all the acceleration we are used to getting from computer simulation. The pipeline could compress quite a bit, because you might screen a whole set of drug candidates in simulation before you ever formulate anything, and by the time you are making the drug in the real world you already know how it will interact with the human system.</span></p><p><span>The word doing the work in all of this is </span><strong><span>fidelity</span></strong><span>. Classical simulation of chemistry is approximate by necessity. Methods like</span><a href="https://en.wikipedia.org/wiki/Density_matrix_renormalization_group"><span> DMRG</span></a><span> throw away a great deal of information about the reaction because there is simply too much to calculate, so we have devised ways to focus algorithmically on the small slice of the problem we hope matters most to the answer. The promise of quantum computing is not throwing away as much, and getting a simulation that captures more of the real dynamics of the system accurately.</span></p><p><span>So if you are working out where to point attention, watch the physical sciences rather than your own industry for the first real result. And treat any near term claim about quantum optimization, quantum machine learning, or breaking encryption with the logical qubit count in mind, because those applications are waiting on hardware that does not exist yet.</span></p><div><hr></div><p><strong><span>Read the full issue</span></strong></p><p><span>This piece comes from a longer conversation on how to read quantum progress honestly. The complete interview and the rest of this week&#8217;s </span><strong><span>Deep Engineering</span></strong><span> issue are</span><a href="https://deepengineering.net/s/newsletter-issues"><span> available here</span></a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Eighty Years of Boolean Thinking Is the Real Barrier to Quantum Adoption]]></title><description><![CDATA[Why building quantum intuition inside an engineering organization takes years, what the JPMorgan research model gets right, and where to put effort before the hardware matures.]]></description><link>https://deepengineering.net/p/boolean-thinking-barrier-quantum-adoption</link><guid isPermaLink="false">https://deepengineering.net/p/boolean-thinking-barrier-quantum-adoption</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 06 Aug 2026 13:42:31 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/db4c1548-4f53-4c80-bf42-35e9e3cb67c4_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><em><span>By </span><a href="https://www.linkedin.com/in/shassinger"><span>Sebastian Hassinger</span></a><span>, author of</span><a href="https://www.packtpub.com/en-it/product/the-new-quantum-era-9781807787370"><span> </span></a><a href="https://www.packtpub.com/en-us/product/the-new-quantum-era-9781807787370"><span>The New Quantum Era</span></a><span> and former quantum lead at </span><strong><span>AWS</span></strong><span> and </span><strong><span>IBM</span></strong><span> | This piece is adapted from his live Deep Engineering session, </span><a href="https://deepengineering.net/p/quantum-computing-beyond-the-hype-sebastian-hassinger"><span>Quantum Computing Beyond the Hype</span></a><span>. Edited by</span><a href="https://substack.com/@saqibjan"><span> Saqib Jan</span></a></em></p></blockquote><p><span>It is smart for enterprises to start investing in the skills you need for understanding quantum information, and in exploring potential algorithms now, even though the computers that could actually run anything useful do not exist yet. That sounds premature until you look at what the work actually involves, because it takes a great deal of effort to recast a business problem into quantum terms.</span></p><p><span>We have been thinking about these problems in classical terms, in Boolean logic terms, since the middle of the last century. That is seventy or eighty years now, and it will be a hundred before too long. So we are carrying a very deeply ingrained set of preconceptions about problem solving. We look at our world through the lens of how our laptop could fix the problem in front of us. We write code on those machines, and our minds run along the lines of decomposing the problem, working out what the dynamics are, and figuring out how to represent them efficiently. All of that rests on an invisible reliance on the assumption that at the ground level you are doing Boolean algebra to solve the problem.</span></p><p><span>I do not think many people are fully aware of how much of their problem solving thinking is rooted in the way classical computing works. And it is so different in quantum computing that </span><strong><span>the gap becomes the real obstacle</span></strong><span>. People talk about developing </span><strong><span>quantum intuition</span></strong><span>, which means building the habit of looking at the world through the lens of linear algebra and through the dynamics and capabilities of quantum information. That takes a lot of work, and it is not the kind of work you can compress once the hardware arrives.</span></p><p><span>The smartest enterprise approaches to quantum computing I have seen understand this. They hire a small number of strong people and run research alongside quantum hardware companies and academic researchers at the top of the field. JPMorgan does this well. You can imagine an organization that size has a very large set of challenges, algorithms, and tasks it has to carry out to operate as a financial entity, and the team there looks at theoretical problems that have some mapping to an aspect of those business processes, then does exploratory research against them. They</span><a href="https://www.jpmorganchase.com/about/technology/research/applied-research"><span> publish open science papers</span></a><span>, so everybody benefits from the work they put in.</span></p><p><span>What they are really doing is building the muscle. When quantum technologies mature to sufficient scale, JPMorgan will know how to use them, because the people there will already have years of thinking in the right terms behind them. </span><strong><span>That head start cannot be bought later.</span></strong><span> The hardware will become available to everyone at roughly the same moment, and the differentiator will be whether your organization has anyone who can look at a business problem and recognize the high dimensional structure inside it.</span></p><p><span>Two things follow from this if you are deciding where to put effort right now. Pick one or two people who are genuinely curious and give them real time to build quantum intuition, rather than sending the whole team to an introductory session that changes nothing. And point them at a problem you actually have, something with heavily interconnected variables, so the learning attaches to your business instead of staying abstract.</span></p><div><hr></div><p><strong><span>Read the full issue</span></strong></p><p><span>This piece comes from a longer conversation on how to read quantum progress honestly. The complete interview and the rest of this week&#8217;s </span><strong><span>Deep Engineering</span></strong><span> issue are</span><a href="https://deepengineering.net/s/newsletter-issues"><span> available here</span></a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Site Reliability Engineer (SRE) Interviews Measure a Depreciating Skill]]></title><description><![CDATA[Notable engineering leaders on motivation, ambiguity, tribal knowledge and the measurement problem behind every reliability hire that did not work out]]></description><link>https://deepengineering.net/p/sre-hiring-interviews-measure-depreciating-skills</link><guid isPermaLink="false">https://deepengineering.net/p/sre-hiring-interviews-measure-depreciating-skills</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Wed, 29 Jul 2026 16:08:58 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/eeacb1d8-a1ef-492f-b24c-3a1a426dbe4f_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Reliability teams are hiring against a job description that describes work their agents increasingly do. That mismatch surfaces as a strong candidate who cleared every round and then underperformed, as a loop where three finalists looked interchangeable and the decision came down to feel, or as a new hire who put their first year into building instrumentation nobody had told them was missing.</p><p>For most of the discipline&#8217;s history the binding constraint during an incident was human attention, because somebody had to correlate logs across services, hold a working model of the system, and form a hypothesis fast enough to matter. Interview loops were built to find that person and they were right to. Agents now take a growing share of that first pass, and the loop has not moved with them. A coding exercise, a debugging scenario, and a systems round all measure how fast a candidate converges on a cause. Those are the skills losing value fastest, while the ones that decide whether the hire works go untested.</p><p><a href="https://www.linkedin.com/in/reid-s/"><span>Reid Savage</span></a><span>, Senior Engineering Manager for the Site Reliability Engineering team at </span><a href="https://www.honeycomb.io/"><span>Honeycomb</span></a><span>, shares how he has repeatedly seen that gap widen across the hires he has made and overseen.</span></p><p><span>&#8220;Standard SRE hiring loops are far too skills-based,&#8221; he says. And his own team acted on that conclusion, removing coding exercises from the loop at Honeycomb&#8217;s current stage.</span></p><h2><span>Motivation decides the hire, not the skills matrix</span></h2><p><span>The failure pattern Savage describes does not look like a skills gap, which is what makes it so easy for a conventional loop to miss. &#8220;I&#8217;ve had excellent hires that have never used our cloud of choice, and some who didn&#8217;t work out that had nearly everything,&#8221; he says. The candidates who arrived with the full stack on their r&#233;sum&#233; were not reliably the ones who lasted, and the ones missing the specific platform experience frequently were. That inversion is not a story about credentials being worthless, and it is a story about credentials answering a question that turns out not to be decisive.</span></p><p><span>What Savage found underneath the pattern was a mismatch of a different kind. The problem in the hires that did not work out was &#8220;usually a mismatch between true motivation (which requires significant self-insight on their part) and what the team needs,&#8221; he explains. Motivation in this framing is not enthusiasm for the role or eagerness in the interview, both of which candidates supply readily and neither of which predicts much. It is the specific thing that makes the work satisfying to that person, which requires enough self-knowledge on their part to name it accurately. Savage anchors the distinction to a line he paraphrases from </span><em><span>First, Break All The Rules</span></em><span>, about hiring the accountant who cannot sleep until the books are balanced rather than the one who knows Excel.</span></p><p><span>The scale of Honeycomb&#8217;s problem space is part of why this holds. &#8220;At Honeycomb&#8217;s size, the skills required for each problem we have wouldn&#8217;t fit in 5 job descriptions,&#8221; Savage says, which means any specific skill the loop screens for covers a small fraction of what the hire will actually encounter. A loop optimized for skills coverage is therefore optimizing a variable that cannot be made to matter, because the surface area defeats it. The organizations most exposed to this are the ones whose reliability function touches many systems rather than one, which describes most companies past their first platform consolidation.</span></p><p><span>Honeycomb&#8217;s replacement for the coding exercise is the actionable part. The technical screen is now a pull request review, chosen because, in Savage&#8217;s words, &#8220;it covers the important things: reading between the lines, social skills, eagerness to help, mindset, and integrity.&#8221; A PR review works as a screen because it presents the candidate with someone else&#8217;s decisions rather than a blank editor, and reliability work consists largely of forming a fast, correct read on choices other people made under constraints that are no longer visible. Leaders who want to run this should use a real pull request from their own history, ideally one where the right call was contested, and score the review on what the candidate noticed and chose to raise rather than on whether they found a seeded bug.</span></p><p><span>Savage&#8217;s position is that the technical bar clears more reliably in the presence of those qualities than the reverse, because &#8220;if you have those, and they display the level of technical aptitude you need, they&#8217;ll be able to catch up.&#8221;</span></p><h2><span>Profile of the hire changed inside a year</span></h2><p><span>Fixing the screen addresses the process. The harder change is to the role itself, and it has moved faster than most loops have been revised. &#8220;The profile of a great SRE has fundamentally changed over the last year,&#8221; says </span><a href="https://www.linkedin.com/in/anish-agarwal-io"><span>Anish Agarwal</span></a><span>, co-founder and CEO of </span><a href="https://www.traversal.com/"><span>Traversal AI</span></a><span>, which builds AI systems for production incident investigation. His account of what companies previously optimized for matches what most loops still measure, which is the engineer who could reason through dashboards, logs, and metrics under pressure and arrive at a root cause. That was the correct thing to optimize while humans carried every investigation from alert to resolution.</span></p><p><span>Agarwal&#8217;s argument is that agentic systems now take on the most time-consuming parts of that work, and the consequence for hiring follows directly. &#8220;Raw troubleshooting ability is becoming less of the differentiator than it was even a year ago,&#8221; he says. A capability stops being a differentiator once it stops being scarce, and a loop weighted toward a non-scarce capability produces candidates who cluster indistinguishably at the top. Leaders tend to notice this as a symptom well before they diagnose it, in loops where several finalists all pass the debugging round cleanly and the final decision comes down to feel.</span></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Packt Deep Engineering! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>What replaces it, in Agarwal&#8217;s view, is the part of the role that always belonged there and rarely got attention. &#8220;What has become far more important is the ability to engineer resilient systems,&#8221; he says, being precise about why that capability went underexercised for so long. Most teams never found time for it because troubleshooting consumed their days, so resilience work stayed aspirational rather than scheduled. Two changes now arrive together, with AI removing much of the investigative burden while production systems grow more complex through new services, new dependencies, and new failure modes each quarter. His conclusion is that resilience engineering &#8220;is no longer the work SREs wish they had time for. It&#8217;s increasingly the work the role demands.&#8221;</span></p><p><span>That reframes what a strong candidate has to demonstrate in an interview. Agarwal describes the highest-leverage SREs as the ones &#8220;thinking about what the system will need a year from now, which SLAs actually matter, and which technologies best meet those SLAs given the tradeoffs between cost, reliability, and performance.&#8221; A loop that wants to surface this has to ask forward-looking questions rather than diagnostic ones, and the practical version is to hand the candidate a real architecture with real growth assumptions and ask what breaks first at ten times the load. The answer reveals whether the candidate reasons about failure modes that do not exist yet, which no incident replay exercise can establish. Agarwal adds one further capability he would weight more heavily, describing it as &#8220;AI fluency. Not just expertise, but genuine curiosity,&#8221; and locating its value in engineers who use agents to compress investigations so that the time freed goes into resilience work.</span></p><h2><span>Judgment is the scarce input, and it leaves when people do</span></h2><p>Hiring toward resilience assumes the judgment to do that work can be found and kept, and that assumption meets an organizational problem that no interview round surfaces.<span> </span><a href="https://www.ghantzaras.com/"><span>George Hantzaras</span></a><span>, Director of Engineering, Core Platforms at </span><a href="https://www.mongodb.com/"><span>MongoDB</span></a><span>, speaking at the AI-Powered Platform Engineering workshop we hosted some time back, argued that the problems facing platform and operations teams have stopped being problems of tooling or frameworks and have become problems of people, team dynamics, and how humans interact with the systems they built. His first category is the knowledge gap, where the answers engineers need lie buried across dozens of systems, dated documentation, and sprawling wikis. He described his own developers as having become archaeologists, running scavenger hunts through internal documentation to recover a single answer.</span></p><p><span>The part of that gap most relevant to hiring is the part nobody wrote down. Hantzaras underscored tribal knowledge as the real problem, meaning the unwritten rules and undocumented workarounds that exist only in the minds of the few people who have been at a company long enough to accumulate them. He described one team where new engineers took close to six months to reach full productivity, with the primary cause being the difficulty of navigating internal systems and discovering rules that appeared on no wiki. That knowledge carries a second liability beyond slow onboarding, because it leaves the organization when the people holding it leave, and nothing in a standard hiring process accounts for either effect.</span></p><p><span>Two further problems he named sharpen what the reliability hire is actually for. Golden paths handle roughly the first eighty percent of cases well and then become what he calls golden cages for the remaining twenty percent, which tend to be the most innovative work and the cases where a developer needs something the template will not express. He also described a platform lead whose team was putting more time into maintaining YAML templates for their portal than into building new platform capability, an overhead he calls the tooling tax. Both problems consume senior judgment on work that produces nothing durable, which is exactly the capacity a resilience-focused hire is supposed to create.</span></p><p><span>Hantzaras reaches the case for automating the first pass from the opposite direction, arriving through the operations burden rather than through the hiring loop. He pointed out that the industry has fixated on mean time to recovery while, in his reading, the largest component is mean time to identify, and that teams are drowning in logs, metrics, traces, and alerts without a good way to act on any of it. The resulting alert fatigue leads engineers to tune out noise until a critical alert has a high probability of being lost. His read on responsibility for that is unambiguous, because it is almost never a case of an engineer not working and almost always a case of a system not working as it should, and his conclusion is that filtering data at that volume &#8220;shouldn&#8217;t be a human job. This is a machine job.&#8221;</span></p><p><span>For leaders hiring into this, the implication is that the scarce input is accumulated judgment rather than throughput, and that judgment currently has no home outside individual heads. The practical response runs on two tracks. Weight the loop toward candidates who can reconstruct the reasoning behind decisions they did not make, since that is the skill tribal knowledge recovery actually requires, and a PR review from an unfamiliar codebase tests it directly. Then treat the documentation of undocumented judgment as an explicit deliverable in the first six months rather than as good citizenship, because the alternative is rehiring the same knowledge every time someone resigns.</span></p><h2><span>Foundations come before fluency</span></h2><p><span>Hiring for AI fluency assumes the environment can support it, and that assumption fails more often than the hiring conversation admits. </span><a href="https://www.linkedin.com/in/chankramath"><span>Ajay Chankramath</span></a><span>, CTO of </span><a href="https://platformengineering.org/"><span>Platform Engineering Community</span></a><span>, a member organization for DevOps and cloud native practitioners, led the platform engineering workshop sessions for Deep Engineering recently, where he described platform maturity as a progression through four states, moving from reactive to responsive, then to predictive, and finally to autonomous. His caution for teams reaching for the later states is that observability, runbooks, and service level objectives have to be in place before AI gets layered on top, because introducing AI into an environment without those foundations amplifies the existing chaos rather than resolving it. That ordering has a direct hiring consequence, because a candidate hired for AI fluency into a team without instrumented systems will put their first year into building the substrate instead.</span></p><p><span>Chankramath is equally clear that the tooling does not displace the role. His framing of AI in incident response is that it hands the engineer a capability rather than taking one away, reading the runbook on their behalf and proposing the next step, while the engineer remains the expert who decides whether that step gets taken. He grounds this in an observation most people on call will recognize, which is that nobody reads documentation during a live incident, because when the system is failing and customers are escalating, engineers act on what they already know rather than opening a markdown file. AI closes that gap by surfacing the relevant procedure at the moment of the decision, and the judgment about whether to follow it stays where it was.</span></p><p><span>He also names a failure mode that should change how leaders think about seniority in these hires. Platform capabilities frequently go unused by the most experienced engineers, who treat self-service tooling as something built for juniors who do not know how to do the work directly, and because those engineers are the ones whose opinions carry weight internally, their dismissal propagates. A reliability hire brought in specifically to raise AI adoption will therefore run into resistance from exactly the people whose endorsement determines whether the change holds. Leaders should test for this in the interview by asking how the candidate has previously won over a skeptical senior engineer, rather than assuming enthusiasm for the tooling transfers to the team.</span></p><p><span>The last piece of Chankramath&#8217;s argument concerns measurement, and it constrains how any of this gets evaluated. He treats the standard delivery metrics as lagging indicators, useful for confirming what already happened and poor at telling a team where things are heading, which makes running an organization solely on them a mistake. Leaders hiring for resilience need leading indicators to judge whether the hire is working, because a resilience improvement shows up as incidents that never occurred, and no lagging metric records an absence. </span>The workable substitutes are the observable precursors, meaning coverage of critical paths by service level objectives, the share of alerts that map to a documented response, and error budget burn rate read before the budget is exhausted rather than after.</p><h2><span>Ambiguity behavior predicts the outcome better than anything else</span></h2><p><span>Across everything Savage has observed, one behavior separates the hires that worked from the hires that did not, and it is not a technical one. &#8220;The largest determinant I see is what they do when faced with an ambiguous problem,&#8221; he says, and he lays out the range of possible responses as moving away from it, tagging in someone else, asking for help, owning the solution, helping someone else own it, or investigating why the problem was ambiguous in the first place. The question underneath all of those, in his phrasing, is whether the candidate accepts or rejects unclear problems. Reliability work arrives almost exclusively as unclear problems, which is what makes this behavior more predictive than any skill the loop can measure.</span></p><p><span>The cost of getting it wrong is not usually a visible failure, which is part of why it goes undetected. Savage describes one SRE who was passionate and knowledgeable, and who struggled to bring others into the work or carry projects across the finish line, producing a long trail of initiatives that started without much to show for them. Because SREs touch every part of the business, that pattern compounds in both directions, generating leverage when it works and opportunity cost plus real cloud spend when it does not. </span>His summary of the missing skill is that &#8220;SREs need to know when to open, peek inside, rearrange, and put back problems.&#8221;<span> The judgment is as much about which problems to close as which to open.</span></p><p><span>Testing for this requires giving candidates something genuinely underspecified rather than a puzzle with a hidden solution. The practical version is to present a real ambiguous situation from the team&#8217;s history, withhold the resolution, and pay attention to whether the candidate starts solving or starts questioning the framing. Both responses can be correct, and what matters is that the candidate demonstrates a deliberate choice between them rather than defaulting.</span></p><h2><span>So define the role before you post it</span></h2><p><span>None of the above helps if the role itself was never specified, which Savage treats as the prior question. &#8220;You need to know what you&#8217;re hiring the SRE for,&#8221; he says, and the diagnostic he runs through is a series of concrete organizational facts rather than an abstract job description. </span>Those facts are whether reliability genuinely matters to the business and to what degree, how many nines the company actually needs, whether anyone currently has responsibility for predicting what will fail a year out, whether developers are resisting on-call, and whether the team is drowning in toil.<span> Each of those describes a different job, and an SRE can address any of them, though not all of them at once.</span></p><p><span>The nines question is the one most organizations answer loosely, and it is the one that determines everything downstream. A target expressed as a service level objective with an error budget attached tells a candidate what the job actually involves, because it establishes how much unreliability the business has already agreed to tolerate and therefore where the engineering effort goes. A posting that claims high availability without naming a target is describing an aspiration rather than a role.</span></p><p><span>His standard for handling that honestly is where the recommendation lands. It is fine to need some of those things and not others, as long as the job posting states what the role is expected to be today and how that might change. Most mis-hires that look like candidate failures are role definition failures surfacing late, because a candidate optimized for toil reduction will underperform in a role that turns out to require reliability forecasting, and neither party did anything wrong in the interview. The discipline is to write the posting from the specific problem rather than from a template, and to name the parts of the role that are unsettled instead of leaving them out.</span></p><p><span>The same clarity should extend into the first ninety days, and Savage&#8217;s approach here inverts the usual instrument. The senior hires who succeeded did not treat his 30/60/90 as gospel, and read it instead as an aspirational list with goals attached, oriented toward making social connections, becoming familiar with the code, and starting the long climb through the tech stack. He asks new hires early on to disagree with him and to offer at least one piece of constructive feedback, which he uses to open room for them to work on problems he cannot see himself. That is a deliberate test of the same quality the loop was built to find, because an engineer who will surface an uncomfortable observation in week three is demonstrating the disposition toward unclear problems that predicts everything else.</span></p><p><span>What connects all four positions is that none of them argues for a lower technical bar. The bar holds, and the argument is that the loop currently puts most of its measurement on the part of the bar that agents are absorbing, while the parts that decide whether the hire succeeds go untested. The reliability hire that works out is the person who accepts unclear problems, reasons about failures that have not happened, carries judgment that was never documented, and improves a system nobody else wanted to own. A process built to find that person looks different from the one most companies are still running.</span></p>]]></content:encoded></item><item><title><![CDATA[System Design Hiring Is Really a Judgment Test]]></title><description><![CDATA[Some repeat architecture while others jump straight to scaling, but the ones who stand out reason through failure, cost, and change.]]></description><link>https://deepengineering.net/p/system-design-hiring-judgment-test</link><guid isPermaLink="false">https://deepengineering.net/p/system-design-hiring-judgment-test</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 09 Jul 2026 18:22:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f3576bfb-c703-4d4d-b8c2-6e0f48dea57f_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Tech hiring is tighter, and AI has raised the bar for what companies expect from senior engineers. When companies do open senior engineering roles, they are paying closer attention to quality of hire, adding more signal to the process, and trusting the old formats less.</p><p>And the system design round is where that scrutiny concentrates, because it is the one round an AI assistant will not get you through. The system design interview may look the same as it did a couple of years ago, but the grading rubric underneath it has changed a lot.</p><p><span>Consider </span><a href="https://karat.com/engineering-interview-trends-2026/"><span>Karat&#8217;s 2026 survey</span></a><span> of 400 engineering leaders. It found that AI has widened the gap between strong and weaker engineers rather than closing it, and 73 percent of leaders now say a strong engineer is worth at least three times their total compensation. CoderPad&#8217;s </span><a href="https://coderpad.io/blog/hiring-developers/new-research-the-2026-state-of-tech-hiring-what-ai-means-for-developers-and-hiring-teams/"><span>State of Tech Hiring Report 2026</span></a><span> shows technical assessments up 48 percent globally since mid 2023, with 60 percent of talent leaders naming quality of hire as their top priority for the year.</span></p><p><span>We asked engineering leaders who personally conduct system design loops what they actually look for in a session. These insights are not usually found in prep guides and they converge on a shift most candidates have not yet noticed.</span></p><h2><span>Buzzword-first designs are a huge turn off</span></h2><p><a href="https://in.linkedin.com/in/architagarwal984"><span>Archit Agarwal</span></a><span>, Principal Member of Technical Staff at </span><a href="https://www.oracle.com/"><span>Oracle</span></a><span>, has in his engineering purview interviewed hundreds of engineers, especially for ultra-low-latency authorization work. And the failure he catches most often has nothing to do with AI. &#8220;I&#8217;ve seen a lot of engineers come in to a system design interview and, as soon as I give a problem, they start with &#8216;let&#8217;s use microservices,&#8217; and start using distributed cache,&#8221; he shares, and when he asks how many users they are planning for, the answer rarely matches the architecture they just proposed. &#8220;That is a key difference between any interview-ready engineer and a genuinely good engineer,&#8221; he explains, because &#8220;a genuinely good engineer would not want to implement everything up front.&#8221;</span></p><p><span>Agarwal&#8217;s baseline is quite blunt, that &#8220;if the problem isn&#8217;t complex yet, don&#8217;t overengineer it.&#8221; And he, like many notable leaders, holds the position that &#8220;microservices aren&#8217;t the magical fix that fixes bad architecture. They just distribute that over the network.&#8221; What he seeks instead is one to two minutes of genuine alignment at the start, functional requirements first to establish what is being built and what the user needs, then the nonfunctional requirements that set scale, consistency, and latency. &#8220;Nonfunctional requirements are the ones that decide the architecture&#8212;not the other way around.&#8221; And not every system earns planet-scale treatment, since a tool used only by a company&#8217;s own engineers needs no multi-region deployment, and proposing one tells him the candidate is performing rather than designing.</span></p><p><span>He also weighs cost. It is something that always runs through his evaluation the same way, and he shared with us a line that reframed his own career, &#8220;a good engineer would design for performance, but a great engineer would design for performance per dollar,&#8221; and he pushes candidates to weigh a latency win against the infrastructure bill it creates. That cost lens is exactly where the interview has changed most, because the most expensive and least predictable component on the whiteboard is now the model.</span></p><blockquote><p>Earlier this year, we spoke with Agarwal about <a href="https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews">trade-offs in modern system design</a>, and later published another piece featuring his candidate-side advice on <a href="https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews">why senior engineers fail system design interviews</a>.</p></blockquote><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Treat every AI component as one that will fail</span></h2><p><span>AI is on every whiteboard now, and the reliability assumptions underneath it are exactly what interviewers have started to stress test. </span><a href="https://www.linkedin.com/in/rohitrpoduval"><span>Rohit Poduval</span></a><span>, Senior Software Engineer on </span><strong><span>Prime Video Trust</span></strong><span> and </span><strong><span>Safety</span></strong><span> at </span><strong><span>Amazon</span></strong><span>, opens with a test that did not exist two years ago. &#8220;Does the candidate treat an AI component as an unreliable dependency or a magic box?&#8221; &#8220;I ask them to design a system that uses an LLM. Content classification, recommendation filtering, whatever fits the role,&#8221; he explains, and then he watches what they do with it. &#8220;Weak candidates draw a box labeled LLM and move on, like it&#8217;s a database that always returns the right answer,&#8221; while strong candidates, he shares, &#8220;immediately ask: What happens when it&#8217;s wrong? What&#8217;s the fallback? How do we know it&#8217;s drifting?&#8221;</span></p><p><span>That one reaction most often separates people who have shipped AI systems from people who have read about them. LLMs return different outputs for the same input, and Poduval points out that you can enforce structure on the output, valid JSON or responses from an allowed set, but you cannot assert that the decision itself is correct with a traditional pass and fail test. Candidates who have built these systems talk about evaluation suites, guardrails that constrain what the model can do regardless of its output, and monitoring for drift in production. But candidates who have not will handwave past every one of those concerns. And that handwaving is quite visible within minutes.</span></p><p><span>&#8220;The next thing I probe: how do they test it?&#8221; Poduval shares, because traditional assertions do not apply to AI components. &#8220;Strong candidates talk about benchmark evaluations: a defined set of inputs with expected behavioral boundaries (not exact outputs) that the system must satisfy before shipping.&#8221; They also talk about continuous auditing against real traffic after launch, not just checks before deployment, because a model that passed benchmarks last month might be drifting today. His deeper observation is that with AI systems you often cannot enumerate every right answer before launch, so the old playbook of define requirements, build to spec, and ship no longer closes. The engineers he rates highly design for safe iteration instead, with shadow mode, holdback groups, evaluation frameworks from day one, and human escalation paths, because they expect to keep refining the definition of correct after launch. &#8220;That&#8217;s exactly where the memorized answers fall apart.&#8221;</span></p><p><span>Taking into consideration the resource-intensive workloads and the shrinking budgets, cost has to be something a candidate can defend out loud. Poduval changes a constraint mid session and tells the candidate the system now needs to handle ten times the traffic. &#8220;Candidates who&#8217;ve never dealt with inference costs at scale just say &#8216;add more instances&#8217;. Candidates who&#8217;ve lived it start talking about which requests actually need an expensive model vs which can be routed to cheaper classifiers,&#8221; he shares. They design tiered systems with routing logic that decides in real time which compute budget each request deserves.</span></p><p><span>The practical takeaway for anyone walking into a loop this year is direct, so state the fallback, the evaluation plan, and the per-request cost tier for any AI component before the interviewer has to ask. Everything on that checklist has an architectural counterpart, and we have captured the nuances of each in </span><a href="https://in.linkedin.com/in/sampritimitra"><span>Sampriti Mitra&#8217;s</span></a><span> practical deep-dive on</span><a href="https://deepengineering.net/p/core-architectural-patterns-for-llm-system-design"><span> core architectural patterns for LLM system design</span></a><span>, which walks through the gateway, tiered fallback, model routing, and evaluation patterns behind it.</span></p><h2><span>Constraint changes reveal who patches and who rethinks</span></h2><p><a href="https://www.linkedin.com/in/chandu-p-a5a896118/"><span>Chandu Putta</span></a><span>, Senior Software Engineer at </span><strong><span>MissionSquare Retirement</span></strong><span>, in our email interview shared what earns his confidence. &#8220;The signal that I trust the most is not the architecture a candidate produces, it&#8217;s how they think if I change a constraint in the middle of the interview.&#8221; In a recent loop he had candidates design a document processing pipeline, and midway through he added an LLM call to the extraction step, with variable latency and non-deterministic output. The majority of candidates bolted a retry mechanism onto their existing design and moved on. One candidate stopped and asked, &#8220;What&#8217;s the acceptable failure mode &#8212; an explicit error to the engineer or a silent degradation?&#8221; he recalls.</span></p><p><span>That question, to his mind, is the level marker, because it treats the new component as a design problem rather than a network hiccup. Putta is blunt about how he sizes this up, as &#8220;rehearsed candidates patch the existing design, but the real senior engineers discard it and start fresh with this new anchor.&#8221; The engineers he considers qualified for the job also ask about context boundaries and cost at scale without being prompted, he adds, and they raise a question that prep articles don&#8217;t teach, which is who decides whether the model output is good enough, the technology team or the business team. An engineer who asks that has sat in the meeting where the answer was contested.</span></p><p><a href="https://www.linkedin.com/in/ed-tian"><span>Edward Tian</span></a><span>, creator of </span><strong><span>GPTZero</span></strong><span>, sees the same moment through a different lens. He shares with us how his interviews confirm the pattern from another direction. &#8220;When we see candidates redesign their system following the introduction of a new constraint, we see that they are able to quickly identify the most significant aspect of their original design that has become the bottleneck,&#8221; he explains, and they design new components to replace it.</span></p><p><span>But weak engineers make incremental changes to accommodate the new requirement without revisiting the larger trade-offs the original design was built on. His sharpest distinction is about posture rather than knowledge, because &#8220;most engineers will assume that their original design is now incorrect after the constraint has been added and defend the changes made to it,&#8221; he shares. &#8220;Senior engineers will simply continue to increase the rate of evolution of their design until it meets their new constraint.&#8221;</span></p><p><span>Tian also shares how he has noticed a newer failure mode that improved AI tooling has made common. &#8220;Many candidates can produce beautiful pipeline diagrams for their LLM, but when we dig through more complex failure modes, for example context window, they really struggle,&#8221; and they default to talking about scaling the model. The diagram generated fluency, but probing exposes it.</span></p><p><span>Agarwal looks for something similar when he changes the requirements mid session. The candidate who absorbs the change, restates it in their own words to confirm alignment, and then highlights which parts of the system must change and which remain intact is showing him a structured redesign instinct rather than attachment to a diagram. He also likes to give away something candidates rarely hear from the other side of the table. &#8220;The curveballs that the interviewer gives you will never be in a way that you will have to scrap the complete diagram, the complete architecture, unless you were already off the track.&#8221; A constraint change is a test of surgical judgment, so the working advice is to name out loud what survives and what dies before you touch the whiteboard.</span></p><h2><span>Can you defend the choice you just made?</span></h2><p><a href="https://www.linkedin.com/in/prakharchaube/"><span>Prakhar Chaube</span></a><span>, Senior Software Engineer at </span><a href="https://www.insurgrid.com/"><span>InsurGrid</span></a><span> who previously held the same level at </span><strong><span>Whatfix</span></strong><span> and has conducted more than 25 technical rounds, argues the design round should never hunt for a correct answer. He assesses the thought process and then questions the candidate on their component choices, because that questioning mirrors what happens inside real companies. &#8220;Most companies have some form of an Architectural Review Board, meaning it&#8217;s rare that a new implementation will not be questioned even if it is the right choice.&#8221; Candidates should know why the standard patterns work as standards rather than assuming a plug and play model, and the ones who can stand by their choice under pressure usually survive a real review. In his loops he is categorical about telemetry, since &#8220;if a candidate is not mentioning that in an interview it is a big no,&#8221; because &#8220;without proper observability the system is doomed to fail at scale and cause chaos during debugging.&#8221;</span></p><p><span>Chaube has also built a ladder for assessing AI fluency that moves well past the 2024 questions. &#8220;I first assess the basics like parameters, prompt structure, prompt security, hallucination,&#8221; he shares, then he wants to dig deeper into implementation scenarios drawn from real products, autonomous tasks orchestrated by a central model such as a code review agent or a login and scraping system. The final rung, he explains, is how candidates handle agent-guided coding, because &#8220;I look at their queries and flow&#8221; rather than their claimed experience. This method, he says, reliably separates candidates who prepared just for the interview from candidates who are genuinely strong. The preparation gap shows up in the follow-ups, never in the opening answer.</span></p><p><a href="http://fr.linkedin.com/in/sandor-dargo"><span>S&#225;ndor Darg&#243;</span></a><span>, senior software engineer at </span><a href="https://engineering.atspotify.com/about"><span>Spotify</span></a><span>, has observed the same gap from years of interviewing in the C++ world, and his experience reinforces Chaube&#8217;s from a different stack. He describes a depth gap first, because a language whose standard runs about 2,000 pages guarantees that even genuine experts have blind areas, and the honest move is admitting them.</span></p><p><span>Darg&#243; in </span><a href="https://deepengineering.net/p/clean-c-code-and-the-hidden-cost"><span>our live interview session</span></a><span> shared how he once told an interviewer he did not know a topic and did not want to guess. The interviewer replied that their team did not use it either, and he got hired. His second observation lands harder, as he has seen senior engineers who handle architectural questions with real thoughtfulness struggle to write simple algorithms live under pressure. Rehearse defending your choices against a hostile follow-up, not presenting them to a friendly one, and practice the small problems you assume you have outgrown.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Safety-critical systems raise the same bar higher</span></h2><p><span>With more than 15 years in secure embedded systems, </span><a href="https://www.linkedin.com/in/hareesha-rameshappa/"><span>Hareesha KoratikereRameshappa</span></a><span>, Senior Technical Product Manager at </span><a href="https://phantom.ai/"><span>Phantom AI</span></a><span>, likes to conduct design interviews where the system under discussion can kill someone. He evaluates whether a candidate can translate high-level product requirements into a structured, developable architecture without compromising safety, cost, accuracy, reliability, and power consumption. He looks for system-level thinking across hardware, firmware, application software, and flashing methodologies like OTA. And what he expects, he shares, is that &#8220;the candidate should understand the key elements of system design, such as system and operating modes, error and fault detection, active and history fault mode managements, transition to the safe mode, and recovery from the fault mode.&#8221; They must also supply the inputs verification depends on, KPIs, requirements, corner cases, and they must plan to monitor the system after deployment rather than treating launch as the finish line.</span></p><p><span>His trade-off vocabulary sounds exactly like the web scale interviewers we heard from earlier, and that essentially is the point. Rameshappa presses candidates on functionality versus performance and accuracy versus resource consumption, and he expects them to justify design decisions with data and adjust the design when another team&#8217;s feedback is valuable. At Phantom AI he shared how he struggled to find candidates who combined production experience with sensor selection, calibration integration, and corner case validation, which tells you how rare the full reasoning package is even among experienced engineers.</span></p><p><span>This means that what the interviewers are testing is not a property of a tech stack. It is engineering judgment, and it transfers from a recommendation pipeline to a multi-camera ADAS system without translation.</span></p><h2><strong><span>Trade-off reasoning still decides the level</span></strong></h2><p><span>Everything above rests on a foundation that has not moved, and </span><a href="https://www.linkedin.com/in/dhirendra-sinha"><span>Dhirendra Sinha</span></a><span>, Software Engineering Manager at </span><a href="https://about.google/"><span>Google</span></a><span>, described it </span><a href="https://deepengineering.net/p/designing-for-scale-and-resilience"><span>last year in one of our conversations</span></a><span> with him and Tejas Chopra. &#8220;It&#8217;s easy to choose between good and bad solutions,&#8221; Sinha told us, &#8220;but senior engineers often have to choose between two good options. I want to hear their reasoning.&#8221; He also carries a line from a chief architect at Yahoo that reframes how scale changes judgment, because when an engineer dismissed a corner case as one in a million, the architect replied, &#8220;One in a million happens every hour here.&#8221; Scale invalidates assumptions, and mostly every leader in this piece is testing whether a candidate&#8217;s assumptions survive contact with it.</span></p><p><span>Agarwal&#8217;s version of the same evaluation focuses on where a candidate chooses to go deep. When a candidate takes one area, distributed storage or authentication or performance engineering, and drives into real depth, he takes that as an engineer who &#8220;understands the gravity of things&#8221; rather than one collecting surface vocabulary. And he calibrates the questions a candidate asks against their level, because clarification questions are always welcome, but a flood of questions too basic for someone with a strong resume tells him the candidate has never actually thought about the system in front of them.</span></p><p><a href="https://www.linkedin.com/in/chopratejas"><span>Tejas Chopra</span></a><span>, Senior Engineer at </span><strong><span>Netflix</span></strong><span>, closes the loop between the old test and the new one with a single follow-up he has asked for years. When a candidate picks SQL over NoSQL or strong over eventual consistency, he asks what changes if the user base grows tenfold, and the AI era version of that question is exactly what some interviewers in technical rounds now ask about model cost and failure modes. The component changed, but the muscle is the same. For readers who want the underlying vocabulary these interviews keep invoking, consistency, availability, partition tolerance, and the read and write trade-offs that connect them, Sinha and Chopra&#8217;s practical deep-dive on</span><a href="https://deepengineering.net/p/distributed-system-attributes"><span> distributed system attributes</span></a><span> from their book </span><a href="https://www.packtpub.com/en-in/product/system-design-guide-for-software-professionals-9781805124993"><span>System Design Guide for Software Professionals</span></a><span> walks through all of it with worked examples.</span></p><h2><strong><span>What to do in your next loop</span></strong></h2><p><span>Ask for the acceptable failure mode before you design anything, because that single question outranks any diagram you will draw. Treat every AI box in your architecture as a component that will be wrong, and price its errors with a fallback, an evaluation plan, and a cost tier per request class. When a constraint changes mid session, say out loud which parts of your design survive and which parts die before you redraw anything. Say what you would monitor without being asked, and say I don&#8217;t know about the right things, because every interviewer quoted here treats honesty about limits as seniority rather than weakness.</span></p><p><span>The rehearsed candidate optimizes for the first answer, and the bar has moved to the third follow-up. That is the whole shift.</span></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p><strong><span>Contributors</span></strong></p><ul><li><p><strong><span>Archit Agarwal</span></strong><span>, Principal Member of Technical Staff, Oracle.</span><a href="https://in.linkedin.com/in/architagarwal984"><span> LinkedIn</span></a><span>. Read his candidate-side advice in</span><a href="https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews"><span> Why Senior Engineers Fail System Design Interviews</span></a><span>.</span></p></li><li><p><strong><span>Rohit Poduval</span></strong><span>, Senior Software Engineer, Amazon (Prime Video Trust &amp; Safety).</span><a href="https://www.linkedin.com/in/rohitrpoduval/"><span> LinkedIn</span></a></p></li><li><p><strong><span>Naga Chand Putta</span></strong><span>, Senior Software Engineer, MissionSquare Retirement. </span><a href="https://www.linkedin.com/in/chandu-p-a5a896118/"><span>LinkedIn</span></a></p></li><li><p><strong><span>Edward Tian</span></strong><span>, Founder, GPTZero. </span><a href="https://www.linkedin.com/in/ed-tian"><span>LinkedIn</span></a></p></li><li><p><strong><span>Prakhar Chaube</span></strong><span>, Senior Software Engineer (SDE-3), InsurGrid. </span><a href="https://www.linkedin.com/in/prakharchaube"><span>LinkedIn</span></a></p></li><li><p><strong><span>Hareesha KoratikereRameshappa</span></strong><span>, Sr. Technical Product Manager, Phantom AI. </span><a href="https://www.linkedin.com/in/hareesha-rameshappa/"><span>LinkedIn</span></a></p></li><li><p><strong><span>S&#225;ndor Darg&#243;</span></strong><span>, Senior Software Engineer, Spotify.</span><a href="https://sandordargo.com/"><span> sandordargo.com</span></a></p></li><li><p><strong><span>Dhirendra Sinha</span></strong><span> (Google) and </span><strong><span>Tejas Chopra</span></strong><span> (Netflix), authors of </span><em><span>System Design Guide for Software Professionals</span></em><span> (Packt). Read their</span><a href="https://deepengineering.net/p/distributed-system-attributes"><span> free chapter</span></a><span>.</span></p></li></ul><p><strong><span>Sources</span></strong></p><ul><li><p><span>Karat, &#8220;Engineering Interview Trends in 2026,&#8221; March 2026. </span><a href="https://karat.com/engineering-interview-trends-2026/"><span>https://karat.com/engineering-interview-trends-2026/</span></a></p></li><li><p><span>CoderPad, &#8220;State of Tech Hiring Report 2026,&#8221; March 2026. </span><a href="https://coderpad.io/blog/hiring-developers/new-research-the-2026-state-of-tech-hiring-what-ai-means-for-developers-and-hiring-teams/"><span>https://coderpad.io/blog/hiring-developers/new-research-the-2026-state-of-tech-hiring-what-ai-means-for-developers-and-hiring-teams/</span></a></p></li><li><p><span>Deep Engineering, &#8220;Why Senior Engineers Fail System Design Interviews,&#8221; May 2026. </span><a href="https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews"><span>https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews</span></a></p></li><li><p><span>Deep Engineering #36, Archit Agarwal on System Design Trade-offs, February 2026. </span><a href="https://deepengineering.substack.com/p/deep-engineering-36-archit-agarwal"><span>https://deepengineering.substack.com/p/deep-engineering-36-archit-agarwal</span></a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[How SUSE Runs AI Without Losing Control]]></title><description><![CDATA[Open source enterprises treat data sovereignty, MCP governance, and cost predictability as one connected problem]]></description><link>https://deepengineering.net/p/how-suse-runs-ai-without-losing-control</link><guid isPermaLink="false">https://deepengineering.net/p/how-suse-runs-ai-without-losing-control</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Wed, 24 Jun 2026 20:20:29 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c14a30e1-69ef-4264-9338-6a7b82cc9e59_2760x1240.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most companies adopting AI are making a quiet trade in exchange for speed. They accept that their data passes through systems they do not control, priced on terms they did not set, governed by rules someone else can change. For a consumer app that trade is often fine. But for an enterprise running infrastructure under compliance regimes, it is not. <a href="https://www.linkedin.com/in/rickspencer3">Rick Spencer</a>, General Manager for Technology and Product at <a href="https://www.suse.com/">SUSE</a>, has for the last two years been working out how an open source enterprise adopts AI at scale without making that trade, and his approach is a useful template for any organization that takes data sovereignty, auditability, and cost predictability seriously.</p><p>SUSE&#8217;s position is unusual in a way that sharpens the problem. &#8220;All the software that we write is open source,&#8221; Spencer explains. &#8220;We&#8217;re not worried about, oh, we leaked the code, we publish the code.&#8221; The concern is the data that belongs to others. When an engineer debugs a customer environment, the logs are not SUSE&#8217;s to hand to a third-party model. &#8220;They trust us to not do those kinds of things,&#8221; he says, and that trust is the thing the entire approach is built to protect.</p><div><hr></div><p><em>You can watch our full interview or read the <a href="https://deepengineering.substack.com/p/sovereign-ai-agentic-infrastructure-rick-spencer-suse">Q&amp;A article here</a>.</em></p><div id="youtube2-8PdtwqLL6YI" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;8PdtwqLL6YI&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/8PdtwqLL6YI?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div><hr></div><h2><strong>Sovereignty means how, not whether</strong></h2><p>The first instinct when AI collides with compliance is to draw a boundary around where AI is allowed to go. Spencer rejects that framing, saying &#8220;I don&#8217;t think it&#8217;s can or cannot. It&#8217;s more so how.&#8221; The distinction matters because a can-or-cannot policy ends with engineers either blocked from useful tools or quietly routing around the rules. A question of how keeps the capability available while controlling the conditions under which it runs.</p><p>The clearest illustration is SUSE&#8217;s build process. A lot of what the company ships is built in an internal instance of Open Build Service, and the defining property of those builds is that they happen offline. &#8220;All the builds are offline. They literally are not connected to the internet,&#8221; he points out. &#8220;This is super important because you need to be able to prove that nothing happened during the build process.&#8221; Proving a negative is far easier when there was no connection through which anything could have happened. It means doing things the hard way, making sure every source is present ahead of time because nothing can be pulled live during the build, but the payoff is provable integrity.</p><p>Applying AI inside that environment is where the how becomes concrete. The example Spencer gives is backporting a patch to previous stable releases, the kind of repetitive, knowledge-intensive work AI is well suited to. The question is whether it can be done in a sovereign way, and his answer is yes, on conditions. &#8220;As long as we are running AI in a way that it&#8217;s able to run disconnected from the internet, and we can have complete visibility into everything it&#8217;s doing.&#8221; SUSE goes further than running existing models in isolation. &#8220;In some cases we even train our own models to accomplish these things,&#8221; he shares, &#8220;and that way we know the model doesn&#8217;t have some naughty time bombs built into it.&#8221; Because owning the model end to end removes the last category of thing the enterprise would otherwise have to take on trust.</p><p>The foundation under all of this is SUSE AI, the company&#8217;s own stack for running AI workloads on private infrastructure, which it uses heavily internally and runs Llama on. &#8220;It&#8217;s all within our private infrastructure, so we make sure there&#8217;s no chance that any data can escape,&#8221; Spencer says. &#8220;We only use models which can be vetted effectively.&#8221; The principle is consistent. Keep the data inside the boundary, and only run models you can actually inspect.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Packt Deep Engineering! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><strong>MCP as the control layer, not just the connector</strong></h2><p>Most of the conversation about Model Context Protocol treats it as plumbing that turns a chatbot into an agent by giving it tools to act with. Spencer agrees that is what it does, but his more interesting argument is that MCP is where an enterprise regains the control that autonomous agents would otherwise erode. &#8220;MCP servers do another really important thing,&#8221; he explains, &#8220;which is provide a place where you can, as an enterprise, bring some sanity and control to the usage.&#8221;</p><p>The mechanism is straightforward once you see the server as infrastructure rather than glue. &#8220;If you have MCP servers running, they&#8217;re just servers,&#8221; Spencer says. &#8220;That means you can provide access ACLs to them.&#8221; A server can be told that a given user&#8217;s agent may use these tools and not those. The usage can be logged. Gateways can sit in front, and the company runs its own alongside a partnership with StackLok. The architectural rule that holds it together is that the language model never touches tools directly. &#8220;You don&#8217;t give the LLMs access directly to tools, only the MCP servers,&#8221; he reasons, &#8220;and then you can have that oversight, meet your compliance needs.&#8221;</p><p>He takes the containment idea down to the operating system. &#8220;You can put the MCP server, I call it, in jail,&#8221; Spencer says, describing a systemd process scoped to present only the compute resources the server actually needs. The reasoning is a security posture rather than a convenience. &#8220;For every MCP server you&#8217;re running, there&#8217;s an LLM out there that&#8217;s trying to use it, and who knows what kind of prompt injections people are running.&#8221; The same boundary defends against the model&#8217;s own failures, not only malicious input. An agent cannot delete a production server with a tool it was never given. &#8220;They guard against things like the AI hallucinating something and deleting your production server, because you simply don&#8217;t provide that tool to it,&#8221; he warns. Control here is not a policy document. It is the set of tools the agent is and is not handed.</p><p>There is a quality dimension to MCP that reinforces the control argument, because a well-built server encodes expert knowledge rather than leaving the model to guess. SUSE ships MCP servers with its products, crafted by the people who know those products best. &#8220;It would be like, instead of you sitting down in front of a chatbot saying I need to figure out how to use Rancher, you&#8217;re sitting down with the whole Rancher development team telling you how to prompt the chatbot,&#8221; Spencer says. An agent working against an expert-built server is not interpreting raw APIs and making guesses a human cannot easily validate. The encoded knowledge makes the agent both more capable and more predictable, which is control of a subtler kind.</p><h2>Cost sovereignty belongs in the same conversation</h2><p>The control story is incomplete if it stops at data and governance, because an enterprise that cannot predict its AI spend has lost a different kind of control. Spencer folds cost into the sovereignty argument directly. &#8220;Sometimes they call it cost sovereignty,&#8221; he says, &#8220;because no one can come back later and say, oh, by the way, we&#8217;re changing our model.&#8221; He has watched a supplier move developers from seat-based to usage-based pricing, a shift his teams did not control and could not prevent. Hosting your own infrastructure changes the nature of the exposure. &#8220;There&#8217;s a maximum cost there,&#8221; he notes of self-hosted AI, and the question shifts from whether you will overrun a variable bill to whether you have the observability to use a fixed capacity fully.</p><p>For the cases where variable cost is unavoidable, SUSE uses circuit breakers that cut off runaway spend in real time when usage spikes past a threshold in a given minute. Spencer is honest that this frustrates engineers who get rate limited mid-task, but the alternative is an autonomous agent running up cost with no human in the loop. The same discipline runs through the company&#8217;s use of frontier models, reserved for high-value strategic work rather than routine completion, and through the practice of using a frontier model once to build an agent that then runs on a cheaper model or a private one. Each of these is the same instinct expressed at the level of money, keeping the expensive capability available for the work that justifies it and putting hard limits around everything else.</p><p>What makes the SUSE approach instructive beyond its own walls is that none of it depends on slowing engineers down. The sovereignty, the MCP governance, and the cost controls exist so that engineers can move fast inside a boundary the enterprise can actually stand behind. &#8220;You don&#8217;t want to stop them from getting that 100X improvement,&#8221; Spencer says. &#8220;You need to give them the right tools for the job.&#8221; Control, in his telling, is not the opposite of speed. It is the thing that makes speed safe to allow.</p><p></p>]]></content:encoded></item><item><title><![CDATA[Compute Obsession Is Slowing Down AI Systems]]></title><description><![CDATA[Why data movement costs more than computation and what most engineers building AI systems are getting wrong because of it]]></description><link>https://deepengineering.net/p/compute-obsession-slowing-down-ai-systems</link><guid isPermaLink="false">https://deepengineering.net/p/compute-obsession-slowing-down-ai-systems</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Tue, 26 May 2026 05:30:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1d53f5cf-8656-4090-8fec-27c78e53b139_690x330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Engineers building AI systems today tend to focus on compute first. It is typically about how many GPU cores, how many parameters, how much VRAM, and how to extract more from all of it. While the benchmarks are about throughput and inference speed, the infrastructure conversations are about scaling horizontally across more hardware.</p><p><a href="https://www.linkedin.com/in/jimledin">Jim Ledin</a>, a seasoned engineering leader, CEO of <a href="https://ledin.com/">Ledin Engineering</a> and author of <a href="https://www.packtpub.com/en-us/product/modern-computer-architecture-and-organization-9781806028023">Modern Computer Architecture and Organization</a> (third edition, Packt), thinks that framing misses the most important constraint in production AI systems. The bottleneck holding back real-world AI performance is not compute but data movement.</p><p>&#8220;Data movement can often be more expensive than the actual computation steps,&#8221; Ledin says. &#8220;The latency, especially moving large data structures across different levels of the memory hierarchy, can dominate and leave a lot of your compute bandwidth idle.&#8221; This is not a niche embedded systems concern. It is happening in the largest AI deployments in the world, and it is the reason hardware vendors like NVIDIA are designing systems the way they are today.</p><blockquote><p><strong>Continue reading</strong> or watch the full conversation with Jim Ledin below.</p></blockquote><div id="youtube2-Q21CJXfQLOk" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;Q21CJXfQLOk&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/Q21CJXfQLOk?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p></p><h2>Memory bandwidth is slowing your AI system more than your GPU is</h2><p>When a CPU or GPU requests data from memory and that data is not available in cache, the processor waits while the computation units sit idle. In a consumer application, that idle time seems like a minor inconvenience. But in an AI system processing large tensors continuously, it accumulates into a significant fraction of total runtime.</p><p>&#8220;AI workloads are becoming increasingly memory bandwidth limited,&#8221; Ledin shares, pointing to a dynamic that is reshaping how AI hardware gets built. &#8220;It is taking more time to bring data into the GPU or TPU memory than it is taking for the computation to take place on the data.&#8221; The raw ability to multiply matrices is no longer the binding constraint. But getting the data to the multipliers fast enough is.</p><p>This is exactly why high bandwidth memory exists. HBM modules are stacks of RAM chips built into a cube, physically close to the processing units, with far higher data transfer rates than conventional DRAM. &#8220;On a TPU card, you typically have several of these HBM modules,&#8221; Ledin explains, &#8220;and they have a far higher data rate for transferring data in and out of the GPU processing components than on a typical consumer grade GPU.&#8221; The engineering bet being made with systems like NVIDIA&#8217;s Blackwell architecture is that memory bandwidth is worth more than raw core count, because the cores are already faster than the data can reach them.</p><p>But there is a side effect that touches anyone buying consumer hardware. &#8220;A lot of the production capacity for memory is going into these high bandwidth memory modules, which cost a lot more for the purchaser and make a lot more money for the vendor,&#8221; Ledin observes. That is a direct reason DDR5 has been difficult to find and expensive when available. The memory fabs are prioritizing the more profitable HBM production, and consumer DRAM is downstream of that decision.</p><h2>The hardware cost your cloud bill is hiding</h2><p>Most software engineers, especially those working in cloud environments, treat the hardware as someone else&#8217;s concern. The abstraction is good enough, the managed services handle the infrastructure, and the code runs somewhere. Ledin&#8217;s argument is that this hands-off relationship with hardware has a real cost that shows up in performance and in cloud bills.</p><p>&#8220;If your code is accessing memory in inefficient patterns, if you are not using the cache memory within the processor in an effective manner, and if you are just moving data around more than is necessary, that can all have significant performance impacts,&#8221; he warns. The CPU requests data from memory, and if it is not in cache, it waits. &#8220;A lot of the time it is unavoidable, but the amount of latency can be minimized by different ways of optimizing algorithms.&#8221;</p><p>The mechanics are specific. When a modern CPU reads from DRAM, even a single byte triggers a 64-byte cache line transfer. The processor brings in a block of adjacent memory whether it needs all of it or not. If the algorithm then jumps to a different memory location, causes that block to be evicted from cache, and later needs it again, it has to re-read it from DRAM. That is wasted time. &#8220;For best efficiency, you would want your code to be working with data from that block before it moves on to something else,&#8221; Ledin explains, &#8220;rather than bouncing around to other memory locations.&#8221;</p><p>In a cloud environment, this inefficiency does not just slow things down. It costs money, and there is no incentive for cloud providers to surface it clearly. &#8220;You are paying for the usage of the system whether the CPU is actually crunching instructions or the CPU is idle waiting for a data item to come in from memory,&#8221; he points out. The cloud bill does not distinguish between productive cycles and stall cycles. Engineers who understand cache locality can write code that reduces stalls and therefore reduces cost, not just latency. Optimizing for cost comes down to understanding your memory access patterns and engineering around them, not just choosing the right managed tooling stack.</p><p>Drawing from his engineering work across embedded and production systems, Ledin shares a useful example. A Linux web server called Tux, which ran in kernel space to avoid user-to-kernel data transfers, developed a performance problem under high load because its per-request state data grew large enough to exceed the CPU&#8217;s level two cache. &#8220;Performance dropped off sharply,&#8221; he recalls. Engineers analyzed the cache behavior, restructured the data layout to keep per-request state smaller, and did the same for instruction caching by batching related processing together. &#8220;Fixes that they implemented increased the application performance by about 40%.&#8221; No new hardware, no architectural overhaul. Just understanding where the memory ceiling was and designing around it.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Deep Engineering publishes weekly expert-led insights on systems design, architecture, and real-world engineering. <strong>Subscribe for free!</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>GPUs are the right tool, but not always for the reason you think</h2><p>The assumption that GPUs are the correct architecture for AI workloads is not wrong, but it is incomplete in a way that matters for engineers making infrastructure decisions. Ledin draws a distinction that is often glossed over in the mainstream conversation about AI hardware.</p><p>&#8220;GPUs are probably the ideal architecture today for people and small companies that want to run language models locally,&#8221; he says, drawing from personal experience. He recently ran the Gemma 4 26-billion-parameter model on an NVIDIA RTX 4090, and for that use case the GPU is the right tool. But for larger-scale deployments running the much larger frontier models, the picture is different. &#8220;The trend there is for dedicated TPUs,&#8221; he notes.</p><p>The distinction matters because GPUs carry silicon dedicated to graphics work that has nothing to do with tensor operations. A consumer GPU has hardware for real-time video rendering, gaming pipelines, and display output. A TPU does not. &#8220;TPUs do not use up silicon for that purpose and focus everything on the tensor work,&#8221; Ledin explains. When you are running thousands of inference requests at scale, that difference in silicon allocation translates directly into efficiency at the workload that actually matters.</p><p>There is also the SIMT execution model to understand. Modern NVIDIA GPUs run 32 threads in lockstep, all executing the same instruction on different data streams simultaneously. This is efficient for linear, parallel workloads. When those threads hit a branch, a conditional where some threads take the if path and some take the else path, the hardware executes one side then goes back and executes the other. &#8220;You basically have effectively a pipeline stall where it has to go back and execute a different thread in that kind of situation,&#8221; Ledin highlights. The flexibility is there, but it comes at a cost. &#8220;Avoiding branching if possible can have a significant impact on performance.&#8221;</p><p>For engineers deciding where to run inference workloads, Ledin offers a practical heuristic. &#8220;The GPU only really becomes attractive when you have enough work for it to do that it can be parallelized and enough that it will amortize the costs associated with moving data onto the GPU, launching the kernels, and doing the management work to transfer data to and from the GPU.&#8221; If the workload is not large enough to keep the GPU busy, the CPU implementation may be faster because it avoids all that overhead entirely.</p><h2>Frameworks are hiding costs that engineers need to see</h2><p>Frameworks and libraries have made it possible to build sophisticated AI systems without ever thinking about what is happening in hardware. That is mostly a good thing. The abstraction accelerates development and reduces mistakes. But there is a point where abstraction stops being a benefit and starts hiding costs that need to be visible.</p><p>&#8220;Where it becomes dangerous to use too much abstraction is when it obscures what is happening with the data layout in memory and the execution patterns,&#8221; Ledin cautions. In performance-critical applications, the framework is making decisions about how data is structured and how the processor interacts with it. If the engineer does not know what those decisions are, they cannot tell when they are working against the hardware.</p><p>The practical approach Ledin recommends is a two-layer architecture. &#8220;Use the most expressive code at the edges of the system, and in the core, use more performance-aware code.&#8221; The boundary between those layers is not always obvious in advance, and finding it usually requires benchmarking rather than reasoning. But the principle is clear: abstractions are appropriate where they preserve meaning across the team, and they become a problem where they hide costs that affect the system&#8217;s ability to meet its requirements.</p><p>One specific pattern worth knowing is the array of structures versus structure of arrays tradeoff. A common data layout is an array of objects, where each object holds all the fields for one entity. For CPU cache efficiency, it can be significantly better to restructure this as a structure of arrays, where each field is stored as a separate array for all entities. &#8220;That might have a big impact on performance,&#8221; Ledin notes, because the CPU cache loads contiguous memory, and if the algorithm is operating on one field across many entities, the structure of arrays layout means each cache load is full of useful data rather than fields the algorithm is not touching.</p><h2>The skills that will matter when the hardware changes again</h2><p>The specific technologies that matter five years from now are difficult to predict, and Ledin is honest about that. &#8220;Four years ago when the previous version of my book came out, it was not at all clear to me, or I think a lot of people, what was going to be happening with AI in the coming years,&#8221; he says. Predicting which hardware architectures or AI frameworks will dominate is not the point. Building the mental model to understand them when they appear is.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.packtpub.com/en-us/product/modern-computer-architecture-and-organization-9781806028023" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qFZO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 424w, https://substackcdn.com/image/fetch/$s_!qFZO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 848w, https://substackcdn.com/image/fetch/$s_!qFZO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 1272w, https://substackcdn.com/image/fetch/$s_!qFZO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qFZO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775" width="336" height="414.46153846153845" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:336,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Modern Computer Architecture and Organization&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.packtpub.com/en-us/product/modern-computer-architecture-and-organization-9781806028023&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Modern Computer Architecture and Organization" title="Modern Computer Architecture and Organization" srcset="https://substackcdn.com/image/fetch/$s_!qFZO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 424w, https://substackcdn.com/image/fetch/$s_!qFZO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 848w, https://substackcdn.com/image/fetch/$s_!qFZO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 1272w, https://substackcdn.com/image/fetch/$s_!qFZO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>3rd Edition</strong></figcaption></figure></div><p>The foundational skill is the ability to reason across abstraction layers. &#8220;The way to really understand the system requires the ability to reason across all of the abstraction layers from the software framework that you are working on at the top level, all the way down to the hardware that runs the code,&#8221; Ledin underscores. That does not mean reading assembly code for every application. It means understanding how pipelines and caches work and orienting code to work within those environments rather than against them.</p><p>The other shift is heterogeneous computing. Writing code that runs on a CPU is no longer sufficient context for many engineering problems. &#8220;It is also becoming more critical to understand heterogeneous computing environments,&#8221; Ledin says. &#8220;It is not just writing code that runs on a CPU. You might also have code that interacts with the GPU if you are running a parallelized algorithm on that, whether it is a language model or something else.&#8221; Domain-specific accelerators, TPUs, RISC-V implementations, and specialized inference chips are all becoming part of the environments that production engineers have to reason about. The engineers who will be most effective in that landscape are the ones who understand why those architectures make the tradeoffs they do, not just how to call their APIs.</p><div><hr></div><p><em>This article is based on <strong>Deep Engineering #46</strong>. You can read the full issue, including additional insights from Jim Ledin on modern computer architecture and AI infrastructure,</em></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;b15102d9-d90c-4b9c-a26f-9fb34ceb12fa&quot;,&quot;caption&quot;:&quot;Memory bandwidth, GPU trade-offs, and the infrastructure decisions that determine whether AI systems are resilient up in production.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #46: Jim Ledin on Modern Computer Architecture and the AI Infrastructure Layer&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:284339563,&quot;name&quot;:&quot;Jim Ledin&quot;,&quot;bio&quot;:&quot;Embedded system developer and cybersecurity tester. Author of \&quot;Modern Computer Architecture and Organization\&quot; and \&quot;Architecting High-Performance Embedded Systems.\&quot;&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d7d3f2aa-9d0a-426d-b514-634881eff942_144x144.png&quot;,&quot;is_guest&quot;:true,&quot;bestseller_tier&quot;:null,&quot;primaryPublicationSubscribeUrl&quot;:&quot;https://jimledin.substack.com/subscribe?&quot;,&quot;primaryPublicationUrl&quot;:&quot;https://jimledin.substack.com&quot;,&quot;primaryPublicationName&quot;:&quot;Jim Ledin&quot;,&quot;primaryPublicationId&quot;:8979199}],&quot;post_date&quot;:&quot;2026-05-07T15:03:11.734Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8adf3738-2898-4207-9a54-e3709b1e9c3c_850x400.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.substack.com/p/issue-46-jim-ledin-computer-architecture-ai-infrastructure-layer&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:196768803,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[Why Senior Engineers Fail System Design Interviews]]></title><description><![CDATA[Scoping, adaptability, and being exceedingly clear matter more than just knowing your tech stack in system design interviews.]]></description><link>https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews</link><guid isPermaLink="false">https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Tue, 19 May 2026 20:19:10 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/bb2a37a3-e396-4fe6-8a55-f331e5572a15_1418x736.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most engineers presume that because they know their tech stack well enough, the system design interview will be easy. And why should they think any differently. They have shipped distributed systems at scale, debugged race conditions at 3am, and made the architectural calls that kept production stable under pressure. But then they walk into a system design interview with confidence and walk out having failed, often without understanding exactly why.</p><p><a href="https://in.linkedin.com/in/architagarwal984">Archit Agarwal</a>, Principal Member of Technical Staff at <a href="https://www.oracle.com/">Oracle </a>where he builds ultra-low-latency authorization services in Go, has interviewed hundreds of engineers. His observation about why experienced engineers fail is the most direct and honest assessment of this problem: they do not fail because they do not know what Kafka is or how DynamoDB handles consistency. They fail because of how they communicate. That single factor, how clearly and deliberately an engineer narrates their thinking, determines the outcome of most system design interviews more than any technical knowledge does.</p><h2>They jump to solutions before understanding the problem</h2><p>Agarwal described a pattern he sees play out repeatedly across interviews at every level of seniority. An interviewer gives a problem and within thirty seconds the candidate is already saying &#8220;I&#8217;ll use Redis, I&#8217;ll use Kafka, let&#8217;s go with microservices.&#8221; The interviewer has not said anything about scale. The candidate has not asked how many users the system needs to support, whether it is read-heavy or write-heavy, what the latency requirements are, or whether there are compliance constraints based on the geography of operation. They have skipped the part of the conversation that actually determines what should be built.</p><p>Those questions are not warm-up questions. They are the questions that drive the architecture. Nonfunctional requirements determine architecture, not the other way around. How many requests per second, what consistency model you need, whether you have a strict latency ceiling, these are the inputs. The architecture is the output. Engineers who skip to the output without gathering the inputs are designing in a vacuum, and the interviewer can see it the moment it happens. Agarwal&#8217;s recommendation is to spend the first one to two minutes of any system design interview doing nothing but alignment: gather functional requirements on what is being built and what the user actually needs, then gather nonfunctional requirements on scale, consistency, latency, and compliance. If you ask the right questions in those first two minutes, Agarwal says, you have already impressed the interviewer. They are listening properly now, engaged, and following where you are going rather than waiting for you to stumble.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Packt Deep Engineering! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2><strong>They design everything at Google scale</strong></h2><p>Senior engineers have worked on large systems and that experience is genuinely valuable, but it also creates a bias that hurts them in interviews: the instinct to design for the most demanding possible version of any problem, whether the problem actually requires it or not. Agarwal is direct about this. Not every system needs to scale to Google. If you are designing an internal tool that will only ever be used by your company&#8217;s engineers, you do not need multi-region deployment, and you do not even need cloud infrastructure. You could run it on a local area network and it would be perfectly adequate for the problem at hand. The engineer who reaches for global infrastructure for a problem that does not need it is demonstrating a failure of judgment, not a depth of knowledge.</p><p>Good system design is about matching the architecture to the requirements you gathered in those first two minutes, not about showcasing every pattern you have ever learned across a career. The interviewer is not evaluating whether you know how to design at Google scale. They are evaluating whether you understand when to use which level of complexity and why, and that distinction is entirely invisible if you default to maximum complexity regardless of the constraints in front of you.</p><div><hr></div><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;45f78df8-1ffd-4b85-a665-d571d1e41c5f&quot;,&quot;caption&quot;:&quot;From monolith-to-services signals to &#8220;performance per dollar&#8221; and practical resilience under real attacks&#8212;clear choices you can defend in production and interviews.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #36: Archit Agarwal on System Design Trade-offs &quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:140662997,&quot;name&quot;:&quot;Divya Anne Selvaraj&quot;,&quot;bio&quot;:&quot;Editor-in-Chief of Deep Engineering by Packt&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/309a6f07-27a6-40bf-ab99-d042556d816b_400x400.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:14853060,&quot;name&quot;:&quot;Archit Agarwal&quot;,&quot;bio&quot;:&quot;Principal Member of Technical Staff at Oracle, building ultra-low latency auth in Golang. Creator of The Weekly Golang Journal, turning system design theory into practical, high-performance code using language and SQL internals, cloud solutions.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1d3d8661-57d3-4b2a-8e42-f789caf14aa6_200x200.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-02-26T13:30:47.032Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/837caa73-ceef-4856-b9c0-49f815314471_1536x1024.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.substack.com/p/deep-engineering-36-archit-agarwal&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:188591840,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:3,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div><hr></div><h2>They go quiet when they are thinking</h2><p>Senior engineers are often comfortable sitting with a difficult problem for several minutes before speaking, and in a production context that is a perfectly reasonable way to work through something complex. In a system design interview it reads as disengagement, and the interviewer has no way to tell whether you are making progress or whether you are stuck. Agarwal uses a phrase that reframes what good communication looks like in this context: the interviewer needs to be able to follow your brain&#8217;s commit history. Every decision you make, every trade-off you consider and reject, every assumption you surface and then validate or invalidate, should be spoken out loud as you make it, not as a performance or a monologue but as a live narration of your actual reasoning as it happens.</p><p>This serves two distinct purposes. It gives the interviewer genuine insight into how you think rather than just what conclusion you eventually reached, which is what they are actually evaluating. And it forces you to be more precise about your own reasoning, because articulating a decision out loud surfaces the assumptions underneath it in a way that thinking silently does not. Agarwal&#8217;s observation is that engineers who think out loud often catch their own errors in real time and self-correct naturally, and that self-correction is not a weakness. It is exactly the kind of flexible, honest thinking the interviewer is looking for.</p><h2>They defend their design when constraints change</h2><p>Experienced engineers have ownership instincts built over years of shipping and defending decisions in production. When they have built something they defend it, and in most professional contexts that instinct is appropriate. In a system design interview it becomes a liability the moment the interviewer introduces a constraint change mid-session, which Agarwal says he genuinely enjoys doing precisely because it reveals something important about the candidate.</p><p>Changing constraints are the normal reality of production engineering. Requirements shift, scale changes, new compliance requirements appear, and the ability to absorb a change, restate it clearly to confirm alignment, identify which parts of the design need updating and which parts remain intact, and then restructure calmly is exactly the capability that distinguishes an engineer who can operate in a real production environment from one who can only design under controlled conditions. The engineers who struggle here are the ones who treat the curveball as an attack on their design and respond by defending the original rather than adapting to the new information. Agarwal&#8217;s point is unambiguous: the interviewer is not trying to invalidate your architecture. They are trying to see whether you can hold your design lightly enough to change it when the situation demands it, which is something you will be required to do repeatedly in any engineering role worth having.</p><h2>They use jargon to sound credible instead of clarity to be understood</h2><p>Senior engineers have large vocabularies built from years of working across complex systems. Distributed systems, eventual consistency, CQRS, saga pattern, two-phase commit. These are real concepts with real meanings and knowing them is genuinely useful. But using them in rapid succession without grounding them in the specific problem being discussed is a signal that the engineer is performing knowledge rather than applying it, and experienced interviewers recognise the difference immediately.</p><p>Agarwal&#8217;s standard for communication in a system design interview is demanding but correct: your explanation should be clear enough that even a junior engineer could follow the reasoning without needing to already know the answer. Not dumbed down, and not simplified to the point of inaccuracy, but clear enough that every choice is grounded in the specific requirements of the system being designed rather than in a general desire to demonstrate familiarity with advanced concepts. The engineers who stand out in Agarwal&#8217;s interviews are not the ones with the most impressive vocabulary. They are the ones who make him feel like he is sitting with another engineer genuinely working through a problem together, which is exactly what a system design interview is supposed to be.</p><div><hr></div><p>The full conversation with Archit Agarwal is now live on Deep Engineering. </p><div id="youtube2-rOLA2NpKPfM" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;rOLA2NpKPfM&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/rOLA2NpKPfM?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div><hr></div>]]></content:encoded></item><item><title><![CDATA[Rust Is Hard for the Engineers with the Most Experience]]></title><description><![CDATA[On the unlearning problem that trips up experienced engineers and the production payoff that makes it worth it]]></description><link>https://deepengineering.net/p/rust-is-hard-for-the-engineers-with-the-most-experience</link><guid isPermaLink="false">https://deepengineering.net/p/rust-is-hard-for-the-engineers-with-the-most-experience</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Mon, 18 May 2026 16:07:41 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f3cfa859-6cd7-4d3e-ab4b-31c7b0cd00c5_1440x660.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Rust, who would have thought, has ranked as the most loved programming language in the Stack Overflow developer survey for nine consecutive years. Honestly, I must admit this is an unusual kind of statistic because it measures not just adoption but retention. The engineers who use Rust want to keep using it, and that pattern has only deepened even as the language moved from systems programming curiosity to production infrastructure at companies including Amazon, Google, Meta, and Microsoft. Interestingly, the Linux kernel now carries Rust code without the experimental label it held for years. <a href="https://lists.debian.org/debian-devel/2025/11/msg00188.html">Debian&#8217;s APT package manager</a> is also introducing hard Rust dependencies this year.</p><p>The performance benchmarks and the memory safety arguments have been made, tested in production, and largely validated. But what the benchmarks do not explain is why so many experienced engineers find Rust genuinely difficult to work with, why the teams that adopt it often go through a period where velocity drops before it recovers, and what it actually takes to get good at it rather than just competent. These are the questions that sit behind the adoption numbers and they matter more than the numbers do for any engineer thinking seriously about where Rust fits in their work.</p><p><a href="https://www.linkedin.com/in/evan-williams-1512092">Evan Williams</a>, author of <em><a href="https://www.packtpub.com/en-ar/product/design-patterns-and-best-practices-in-rust-9781836209461">Design Patterns and Best Practices in Rust</a></em>, has been writing software for more than 40 years and came to Rust while building a hardware system that needed to be rock solid and run without access in a remote location. <a href="https://it.linkedin.com/in/francesco-ciulla-roma/en">Francesco Ciulla</a>, author of <em><a href="https://www.packtpub.com/en-us/product/the-rust-programming-handbook-9781836208860">The Rust Programming Handbook</a></em><a href="https://www.packtpub.com/en-us/product/the-rust-programming-handbook-9781836208860"> </a>and a Docker Captain who previously worked at the European Space Agency on the Copernicus project, started publishing Rust content in 2022 and 2023, earlier than most, and has since watched the language&#8217;s adoption from the inside. Their perspectives on Rust come from different parts of the stack and different kinds of work, but on the questions that actually trip engineers up, they are in close agreement.</p><blockquote><p><em>We interviewed Evan Williams and Francesco Ciulla separately for <a href="https://deepengineering.substack.com/s/newsletter-issues">Deep Engineering Newsletter</a> issues.</em></p></blockquote><h2>The engineer who struggles most is usually the most experienced one</h2><p>The reasonable assumption when a team introduces Rust is that the senior engineers will pick it up fastest. They have the most context, the most pattern recognition, and the most experience navigating unfamiliar codebases. In practice, the opposite tends to happen, and both Williams and Ciulla have seen it play out firsthand.</p><p>&#8220;The more experienced you are, the more years you have doing something in some other language, the more trouble you&#8217;re likely to have,&#8221; Williams says, &#8220;because you have patterns of thought that come from those languages that you don&#8217;t even realize are there.&#8221; The problem is not that experienced engineers consciously try to apply Java or C++ patterns to Rust. The problem is that those patterns are invisible to them, baked in over years of use until they no longer register as choices at all. The engineer is not making a decision when they reach for inheritance or shared mutable state. They are doing what has always worked, and Rust will not let them.</p><p>Ciulla put it more directly. &#8220;Even if you are a senior developer, even if you have twenty years of experience, if you want to try to learn Rust comparing it to other programming languages, you will fail, because it&#8217;s like learning something which is completely new.&#8221; He makes the case that this is not a reason to avoid Rust but a reason to go in with a specific kind of openness, one that experienced engineers often find harder to maintain than junior ones do precisely because they have more to unlearn. A developer learning Rust as their second or third language has no competing mental model to discard. A senior engineer with a decade long experience in Java has to dismantle instincts that have been reliable for years before they can build new ones, and that dismantling is the work that most people underestimate going in.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;ef4dcc9c-3432-4c97-99aa-2f125e7d6b37&quot;,&quot;caption&quot;:&quot;Francesco Ciulla has been building with Rust since 2022, has spoken about it internationally at conferences, and his perspective on Rust adoption is shaped less by enthusiasm for the language and more by a practitioner&#8217;s view of where it actually earns its place in a production system.&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #45: Francesco Ciulla on Building Production Systems in Rust Without the Expensive Rewrite &quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:11407185,&quot;name&quot;:&quot;Francesco Ciulla&quot;,&quot;bio&quot;:&quot;Developer Advocate at @dailydotdev\n&#183; Docker Captain &#128051;\n&#183; Public Speaker\n&#183; Building a 1 Million Community 22%&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b8c30606-10ba-4c87-a89b-af2f9dc27a01_400x400.jpeg&quot;,&quot;is_guest&quot;:true,&quot;bestseller_tier&quot;:null,&quot;primaryPublicationSubscribeUrl&quot;:&quot;https://francescociulla.substack.com/subscribe?&quot;,&quot;primaryPublicationUrl&quot;:&quot;https://francescociulla.substack.com&quot;,&quot;primaryPublicationName&quot;:&quot;Francesco's Newsletter&quot;,&quot;primaryPublicationId&quot;:1410908}],&quot;post_date&quot;:&quot;2026-04-30T16:32:00.000Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6e29b888-c64d-4029-96f3-08e45522d077_656x375.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.substack.com/p/issue-45-francesco-ciulla-building-production-systems-rust-without-rewrite&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:195995087,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:2,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>Trusting the compiler is not a beginner tip</h2><p>The first thing most engineers do when the borrow checker rejects their code is look for the minimum change that will make it compile. That is the right approach in almost every other language and the wrong one in Rust, and it is where a significant amount of early frustration comes from.</p><p>&#8220;The golden rule is to trust the compiler, especially at the beginning,&#8221; Ciulla says. What he means is not passive acceptance but active reading. The borrow checker is not producing noise. It is producing information about what the program&#8217;s structure requires, and engineers who learn to read it that way move through the learning curve faster than engineers who treat every error as an obstacle to clear. The difference is subtle at first and significant over time.</p><p>Williams argues that the borrow checker is doing something more useful than preventing bugs. &#8220;The borrow checker is your friend because it prevents you from making a messy design. It prevents you from making a broken design. It prevents you from writing whole classes of bugs that you will then spend many hours trying to find,&#8221; he explains. &#8220;I have found it to be an incredible partner in writing code that allows me to sleep at night.&#8221; The reason it works this way is that Rust&#8217;s ownership rules enforce a discipline that experienced engineers in other languages apply selectively and inconsistently because those languages do not require it. A value has one owner. References are either shared and immutable or exclusive and mutable, never both. The compiler will not proceed until the code is explicit about who owns what and when.</p><p>&#8220;The principles that the borrow checker forces you to adhere to in Rust are the exact principles that you should be using in every programming language,&#8221; Williams reasons. &#8220;But you don&#8217;t have to. So it&#8217;s very easy to not think about those things.&#8221; That observation reframes what the borrow checker is. It is not an imposed restriction. It is a discipline that good engineers apply in other languages by habit and judgment, made non-negotiable and automatic in Rust.</p><p>The discipline extends from individual functions to the shape of the whole system. A program that handles ownership correctly at the function level has to handle it correctly across modules, across threads, and across component boundaries, because the same rules apply everywhere. &#8220;You need to think about who controls what, how it is controlled, and you need to start from the very beginning thinking about the boundaries of your program and the system architecture, dividing things up into areas of responsibility,&#8221; Williams underscores. &#8220;Because unlike Python or Java, you can&#8217;t have links going all over the place. The borrow checker is never going to accept that.&#8221; The result is that well-written Rust systems tend toward a specific architectural shape: data flows in one direction, ownership chains move forward and do not loop back, and the behavior of the system is legible from its structure in a way that systems with shared mutable state often are not.</p><p>The most underutilized expression of what this makes possible is the typestate pattern. It uses the type system to encode the state of a value at compile time in a way that makes invalid state transitions not just errors but programs that cannot be compiled at all. Williams reflects on it with visible enthusiasm. &#8220;It&#8217;s a way of developing state machines and systems that have state that evolves where invalid state transitions aren&#8217;t just errors, they&#8217;re impossible to write. The compiler won&#8217;t compile them,&#8221; he says. &#8220;It represents a huge advance in the way that such systems are written because now instead of runtime errors, you have a state machine that is guaranteed to work because every transition either is a valid transition or it won&#8217;t even compile. That&#8217;s an amazing thing.&#8221; The pattern was not invented for Rust, but the language&#8217;s ownership system and type handling make it practical in a way that other languages do not, and for systems where invalid state transitions are genuinely dangerous rather than merely inconvenient, it is one of the most concrete expressions of what Rust makes possible.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;caa3ca39-6118-4ad9-9d55-53c3803aabaf&quot;,&quot;caption&quot;:&quot;Evan Williams shares engineering insights on the borrow checker as a design tool, the object-oriented trap, and why the engineers who struggle most with Rust are often the most experienced ones.&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #47: Evan Williams on Why Experienced Developers Have the Hardest Time Learning Rust&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-05-14T16:42:52.048Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/43a42d88-ec70-4d7f-8213-85796343b4f5_677x337.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.substack.com/p/deep-engineering-47-why-experienced-developers-hardest-time-learning-rust&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:197666671,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>What Rust actually gives you in production</h2><p>Ciulla&#8217;s case for Rust is grounded in things he measured when running a Rust web server on his own machine. It was consuming four megabytes at rest and five in production. &#8220;If you have a droplet with one gigabyte of RAM, you can have 200 plus services,&#8221; he notes, &#8220;of course in idle, but this proves that if you have a service that consumes a lot of RAM, it is worth thinking about.&#8221; For teams running infrastructure where memory costs money and density matters, the difference between a Rust service and an equivalent service in a garbage-collected language is not marginal.</p><p>The latency story is also specific. &#8220;By not having a garbage collector on the back end side, you basically have a flat latency,&#8221; Ciulla observes. &#8220;If a user makes an HTTP request when the garbage collector starts, it will experience a higher latency. Rust removes that problem entirely.&#8221; Go and Node.js both have garbage collectors that pause for collection cycles, and even short pauses measured in hundreds of milliseconds are enough to introduce latency spikes that affect users who are unlucky enough to hit the request at the wrong moment. Rust&#8217;s absence of a garbage collector means the latency profile is predictable rather than probabilistic, which matters significantly for services where consistency is as important as average throughput.</p><p>The deployment model is simpler than most engineers expect going in. A Rust project built with cargo produces a standalone binary for the target architecture, which packages cleanly into a container image. &#8220;If you build the executable when you build the Docker image, you have something which is just deployable everywhere,&#8221; Ciulla says. &#8220;A Linux executable running in a Docker container. That&#8217;s the dream.&#8221; The operational benefit is smaller images, faster startup, and a runtime with almost no overhead beyond the binary itself.</p><p>Williams approaches the production question from the correctness angle rather than the performance angle. The systems where Rust earns its place most clearly are the ones where failure has a real cost. &#8220;Systems that are mission critical in some way or other are really key Rust use cases,&#8221; he says. &#8220;All of these features combine into a whole that make Rust a really powerful language for doing things that have to work. Things where failure is monetarily or in human cost even a terrible problem.&#8221; The memory safety and the ownership model and the compile-time guarantees are not separate features. They are different expressions of the same underlying commitment: the program either demonstrates its correctness to the compiler or it does not compile.</p><p>Williams also reflects on something unexpected he discovered while writing the early chapters of his book, the ones covering what not to do in Rust. He went back and deliberately tried to write bad code, the kind of code that would illustrate the mistakes he was cautioning against, and found it harder than he expected. &#8220;When I went back and tried to write bad code in Rust, it was much harder than writing the good code,&#8221; he recalls. &#8220;That&#8217;s an interesting perspective that just didn&#8217;t even occur to me.&#8221; The language&#8217;s constraints push code toward a particular shape so consistently that departing from it requires actively working against the grain of the language rather than simply making a poor choice.</p><p>&#8220;The biggest benefit in Rust is about the lack of the debugging depth. You spend more time thinking up front, but you spend almost zero time chasing segfaults or memory leaks in production,&#8221; Ciulla remarks. &#8220;And we always underestimate this part. We always talk about the efficiency of the code, but if you need less time to debug your code, you&#8217;re basically writing more logic at the end of the day.&#8221; The upfront investment in getting the types and the ownership right is real, but the downstream debugging cost it removes is larger and does not diminish as the team becomes more experienced. It is simply gone.</p><h2>Where to start and where Rust is the wrong tool</h2><p>On the practical question of how to bring Rust into an existing codebase, both Williams and Ciulla give advice that converges almost exactly despite coming from different engineering contexts. Neither recommends starting with a rewrite.</p><p>&#8220;The best way to introduce Rust in a big project is to find that hard part that&#8217;s the bottleneck and try to write one single service in Rust,&#8221; Ciulla says. &#8220;And then you will see, probably slowly, Rust might take over your code base, but I mean this in a good sense.&#8221; Williams makes the same point with a specific warning about the temptation to go faster. &#8220;What you don&#8217;t want to do is jump into saying, we&#8217;re just going to rewrite our project in Rust now. Pick a small piece, focus on that, gain confidence and mastery of the language, and then use that to build upon it and start bringing in more things,&#8221; he says. Starting with a bounded, non-critical component gives the team room to move through the learning curve without the pressure of a production incident concentrating everyone&#8217;s attention on the wrong things.</p><p>Ciulla adds something worth noting about the AI-assisted workflow that is becoming standard for many engineers. &#8220;In this AI era, everyone is rushing stuff with AI, but you still need the validation,&#8221; he says. &#8220;Okay, AI wrote this Rust service, but now who decides if this is okay to put in production? Of course, you need the validation of an expert.&#8221; The Rust compiler catches a large class of errors automatically, but the errors that survive it, logic errors rather than memory errors, still require someone who understands the language well enough to see what the code is actually doing. Having at least one engineer on the team who knows Rust well enough to review AI-generated code is not optional.</p><p>Both are also direct about when Rust is not the right choice. Ciulla points to tight deadlines and fast prototyping as the clearest case against it. &#8220;If you need fast prototyping, you are familiar already with Java, JavaScript, why don&#8217;t you use it?&#8221; he says. &#8220;When the deadline is so close, probably it&#8217;s not the best way to try something new because something would go wrong, especially if you&#8217;re not an expert.&#8221; </p><p>Williams points to user interfaces as an area where the ecosystem is still catching up and the tooling gaps are large enough to make other languages more practical. &#8220;Doing a website in Rust is still kind of a feat,&#8221; he notes, &#8220;and it&#8217;s an awful lot easier to use the tools that everybody else is using to accomplish that goal.&#8221; The Python data science ecosystem is another area Ciulla names directly: the libraries are simply better established there, and using Rust for data science work means building against a thinner set of available tools than Python provides.</p><h2>Where Is the Ecosystem Headed</h2><p>Ciulla expects Rust to grow most significantly in the near term, and his prediction lands in a direction that surprises most of the Rust community. &#8220;I think the next big wave might be in web development,&#8221; he says, adding that he is aware this is an unpopular position in a community that still thinks of Rust primarily as a systems language. His reasoning is grounded in what he has been seeing directly: companies with hundreds of developers reaching out to tell him they are moving services to Rust for their web backends. &#8220;I get this news because I&#8217;m well known for talking about Rust and being quite vocal about it,&#8221; he observes. &#8220;I&#8217;m not talking about a person just doing this on a random Saturday night. I&#8217;m talking about companies that have hundreds of developers.&#8221; He points to Axum as the framework that has matured to the point where he would now use it in a production SaaS product, which he says was not true two years ago. Embedded systems, in his view, have already crossed the threshold where Rust&#8217;s place is settled.</p><p>Williams takes a longer view on how the ecosystem will evolve. &#8220;The ecosystem is going to get richer and people are going to be branching out in the set of use cases, hitting areas that right now Rust has relatively weak support for,&#8221; he says. &#8220;As larger and larger projects are built, there is going to be more refinement of the language itself, but more importantly, more refinement of the use of the language.&#8221; The patterns that make Rust work well at scale are still being discovered and codified. The language is stable, but the understanding of how to use it well is still developing, and that development is happening inside the teams building the largest Rust codebases.</p><p>What both conversations point toward is a language whose difficulty and whose value come from the same source. Rust is hard to learn for experienced engineers because it refuses to accommodate the habits that made them experienced. It is valuable in production because that same refusal, enforced by the compiler on every build, produces code whose behavior is predictable, whose data flows are legible, and whose failure modes are constrained to things the language cannot check rather than things the engineer forgot to check. The engineers who get the most out of it are the ones who stop trying to carry their existing instincts across and start letting the compiler teach them what the program actually needs.</p><div><hr></div><h4><em><strong>In case you missed</strong></em></h4><p><em>Here&#8217;s the full interview video featuring Evan Williams.</em></p><div id="youtube2--ElpmT7DCX4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;-ElpmT7DCX4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/-ElpmT7DCX4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div>]]></content:encoded></item><item><title><![CDATA[Clean Code Is a Trap, Decompose Instead for Physics and Performance]]></title><description><![CDATA[On cache locality, cognitive load, and the engineering habits that make fast development sustainable with Sam Morley and S&#225;ndor Darg&#243;]]></description><link>https://deepengineering.net/p/clean-code-trap-decompose-for-performance-physics</link><guid isPermaLink="false">https://deepengineering.net/p/clean-code-trap-decompose-for-performance-physics</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 23 Apr 2026 15:15:59 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e4a62837-3922-40ce-8859-73c783c89af9_822x371.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Engineering teams obsess over clean code because they want software to look organized and logical in the text editor. Principles like SOLID get followed strictly, and hours get spent debating folder structures, because it feels like the disciplined way to build software. But this desire for logical cleanliness often leads into a trap where teams build systems that are beautiful to read but terrible to run.</p><p>The most maintainable codebases are not the ones that adhere to a style guide. They are the ones that respect the physical and cognitive reality of the environment they live in.</p><p>We have spoken (interviewed separately) to two notable engineers who think about this from very different directions. <a href="https://www.linkedin.com/in/morleys90/">Sam Morley</a>, a mathematician and C++ researcher at the <a href="https://www.maths.ox.ac.uk/">University of Oxford</a>, approaches software from the ground up, where the cost of every abstraction shows up immediately in performance metrics. <a href="https://www.linkedin.com/in/sandor-dargo/">S&#225;ndor Darg&#243;</a>, a senior software engineer at <a href="https://engineering.atspotify.com/">Spotify</a> who works on large-scale C++ systems, approaches it from the maintainability side, where the cost of every abstraction shows up in the engineers who have to live with the code months or years later. </p><p>Both conversations happened at different points, on different topics, but arrived at the same conclusion that logical cleanliness is not the goal. But understanding what the machine and the team actually need is.</p><h2>Your CPU does not care how tidy your objects look</h2><p><strong>Sam Morley&#8217;s</strong> starting point is hardware, and his argument is that the way most engineers are taught to structure code works directly against the way processors are designed to access memory.</p><p>The instinct is to group data into objects because it models the real world. A Player class holds position, health, velocity, and inventory in one contiguous block, because those things belong together conceptually. But the CPU fetches data in contiguous blocks called cache lines, and if the object structure fills that cache line with data the processor does not need for the current operation, the application pays for it in cycles. The cost is invisible in code review but shows up immediately in a profiler under load.</p><p>Morley points to the Structure of Arrays pattern, common in game development, as the counterintuitive solution. Instead of an array of Player objects, you create separate arrays for positions, health values, and velocities. This looks messy to a developer trained in object-oriented design. It violates the instinct to keep related data together, and it produces code that does not map neatly onto the real-world entities it represents. But it allows the CPU to process data significantly faster because every byte in a fetched cache line is a byte the processor actually needs. Cache locality, not conceptual tidiness, determines throughput under real conditions.</p><p>Morley&#8217;s recommendation is direct: be willing to break clean object models when the hardware requires it. The machine is not going to adapt to the abstraction. The abstraction has to adapt to the machine. And this is not a concern limited to embedded engineers or game studios. It is a reality for any C++ system under sustained load, and the gap between what looks clean and what runs efficiently widens as the scale increases. Teams that do not understand this distinction tend to optimize the wrong things when performance problems eventually surface.</p><h2>Clever code is a debt that Future You will have to repay</h2><p>Morley&#8217;s second argument shifts from CPU cost to cognitive cost, and it is the more insidious of the two because it compounds slowly and invisibly until a maintenance crisis makes it visible all at once.</p><p>His framing here is precise. Future You is a completely different person who has lost all the context that made the current design feel obvious at the time it was written. The engineer writing the code holds the whole system in their head. The engineer returning to it six months later does not. And the engineer reading it for the first time never did. Every clever abstraction that felt natural in the moment of writing becomes a reconstruction problem for every reader who comes after.</p><p>Template-heavy code and metaprogramming are the most common form of what Morley calls Wizardry. The name is apt because Wizardry works by concealment. The complexity does not disappear when abstracted away. It becomes invisible until someone needs to debug or extend the system, at which point the engineer is starting from a significant disadvantage with no clear view of how data actually moves through the code. What Morley advocates instead is Process Awareness: code that exposes the data flow clearly rather than hiding it behind layers of indirection. Not short code or smart code. Code whose execution model is obvious to the next engineer who reads it, regardless of whether that engineer was involved in writing it.</p><p>The practical implication is to treat <strong>Future You</strong> as a first-class stakeholder in every design decision. And so, the documentation that explains what the code does is far less valuable than documentation that explains why it is structured the way it is, because the what is usually legible from the code itself. The why rarely is.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;b51368e0-7c19-47ba-bdef-b91f8ea7f718&quot;,&quot;caption&quot;:&quot;Template metaprogramming, cache-aware design, concurrency models, and why learning Rust might actually make you a better C++ programmer&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #41: Scaling C++ the Right Way with Sam Morley&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:480630041,&quot;name&quot;:&quot;Deepayan Bhattacharjee&quot;,&quot;bio&quot;:&quot;Content engineer at Packt&quot;,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!jegY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F385931a8-2eb0-4f8a-bd57-14713ff3988d_144x144.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:440051761,&quot;name&quot;:&quot;Sam Morley&quot;,&quot;bio&quot;:&quot;Research software engineer and mathematician on the DataSig project at the University of Oxford.&quot;,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!hfZA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b7dcfcf-a878-45d0-99e4-a8f2045dee3e_144x144.png&quot;,&quot;is_guest&quot;:true,&quot;bestseller_tier&quot;:null,&quot;primaryPublicationSubscribeUrl&quot;:&quot;https://sammorley.substack.com/subscribe?&quot;,&quot;primaryPublicationUrl&quot;:&quot;https://sammorley.substack.com&quot;,&quot;primaryPublicationName&quot;:&quot;Sam Morley&quot;,&quot;primaryPublicationId&quot;:7726502}],&quot;post_date&quot;:&quot;2026-04-02T15:16:15.996Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4017b543-27c2-4082-9f3b-1bd7abbbdd3a_1432x840.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.substack.com/p/deep-engineering-41-scaling-c-the&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:192949941,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:4,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>When cognitive load becomes your biggest bug</h2><p><strong>S&#225;ndor Darg&#243;</strong> approaches the same problem from a different direction but arrives at the same place. His work at Spotify on large-scale C++ systems has given him a practitioner&#8217;s view of what happens to codebases over time when cognitive cost is not treated as a first-class engineering concern from the start.</p><p>For Darg&#243;, the thread connecting clean code, binary size, undefined behavior, and C++ language evolution is a single idea: reducing complexity in real-world systems. Not as an aesthetic preference, but as a measurable engineering outcome with consequences for how fast teams can move, how safely they can refactor, and how much institutional knowledge survives when people leave. &#8220;If you think about clean code, it clearly reduces the cognitive load,&#8221; Darg&#243; said during a <a href="https://deepengineering.substack.com/p/clean-c-code-and-the-hidden-cost">recent Deep Engineering interview.</a> &#8220;If you think about binary size, it might reduce operational cost. New standards like C++23 and C++26 reduce boilerplate and enable safer, more readable abstractions. All of these topics make large C++ systems more maintainable and more evolvable.&#8221;</p><div id="youtube2-vsdeOS8snN0" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;vsdeOS8snN0&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/vsdeOS8snN0?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>The connection between these concerns is not accidental. Binary size reduction often leads teams toward simpler code as a side effect, because the practices that reduce binary size, avoiding unnecessary template instantiation, being deliberate about what gets inlined, minimizing heavy type erasure, also tend to reduce the number of moving parts an engineer has to hold in mind. The discipline required to keep a binary small and the discipline required to keep a codebase readable are more closely related than most teams realize until they have worked on both problems at the same time.</p><p>Darg&#243;&#8217;s warning is about the human cost of poor abstraction choices, and in his experience, teams routinely optimize the wrong things because they measure the wrong variables. The heap allocation is visible. The cost of a network request made inside a loop is harder to see until a profiler makes it undeniable. Darg&#243; during our interview cited Amdahl&#8217;s Law to make the point concrete: the overall performance improvement gained by optimizing a single part of a system is limited by the fraction of time that part is actually used. The engineers spending time on heap allocations while making network requests in a loop are not being careless. They are solving the problem they can see. The discipline is in learning to find the problem that actually matters, which requires measurement rather than intuition. &#8220;If your code takes a long time to execute due to network latency, then relatively speaking, the heap allocation is not so slow anymore,&#8221; Darg&#243; said. &#8220;Don&#8217;t worry about things that don&#8217;t really matter in a given environment.&#8221;</p><h2>Write it readable first, then measure, then and only then optimize</h2><p>Darg&#243;&#8217;s practical framework for navigating these trade-offs is structured around a clear hierarchy of defaults, and the first default is unambiguous: readable code comes first.</p><p>His reasoning is grounded in a simple observation that engineering culture tends to underweight. Engineers read code far more often than they write it. Every decision that makes code harder to read imposes a recurring cost on every future reader, and that cost accumulates over the lifetime of the codebase. Defaulting to readability is not a concession to comfort. It is an engineering position with compounding returns, because code that is easy to read is code that is easy to reason about, and code that is easy to reason about is code that is safer to change.</p><p>The second principle follows directly: if optimization is necessary, measure before touching anything. The trap is optimizing before a measurement has confirmed that the thing being optimized is the actual problem. This wastes time, introduces unnecessary complexity, and often leaves the real bottleneck untouched. Measure first, identify the hot path, and only then begin the optimization work. Once the hot path is identified, keep it isolated and document the reasoning behind every trade-off made there. Not documentation that explains what the code does, but documentation that explains why it is structured the way it is, so the next engineer understands what they would be giving up if they cleaned it up.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Packt Deep Engineering! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Darg&#243; has been on the receiving end of the alternative. He came into a codebase, saw code that looked wrong, began cleaning it up, and realized too late that the seemingly redundant choice was affecting binary size in a way that mattered for the system. Pull requests had already merged before the context became clear. &#8220;Make trade-offs conscious,&#8221; Darg&#243; said. &#8220;Make them explicit in code reviews, but also in the code itself. If you sacrifice the clarity you aim for, document why. Because otherwise someone later will come in and make it cleaner, unaware of why certain choices were made.&#8221;</p><p>And this principle has become more critical in the age of agent-assisted development. If engineers can miss the intent behind an undocumented trade-off, then AI agent working on the same codebase will miss it with far greater confidence. Agents read what is in the code. They do not have access to the Slack conversation where the binary size constraint was first discussed, or the code review thread that resolved and got deleted. The context has to be in the code, because that is the only place every future reader, human or agent, will reliably look.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;b266a406-3df1-466a-984c-0c85e1ba4344&quot;,&quot;caption&quot;:&quot;Building an AI-Powered Internal Developer Platform from Scratch&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #44: S&#225;ndor Darg&#243; on C++26, Adoption Traps, Compiler Gap, and Maintainability&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-04-23T16:31:24.008Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/efb9ba36-137d-4762-93d8-c394bc1fe4da_681x277.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.substack.com/p/issue44-cpp-26-adoption-traps-compiler-gaps-maintainability&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:195242079,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:7,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>The invisible tax that is getting harder to ignore</h2><p>Morley and Darg&#243; are describing the same underlying problem from different directions. Every time an engineer has to reconstruct context that was lost, the system has failed them. Morley calls it the Future You constraint. Darg&#243; calls it cognitive load. The mechanism is identical in both cases, and the cost is real even when it does not appear in any metric the team currently tracks.</p><p>This cost has become harder to ignore in the last year or two, and not only because systems have grown more complex. Darg&#243; observed during the same session that the shift to AI-assisted development has made context switching materially worse for most engineers, and the profession has not yet fully reckoned with what that means for how software gets built. Engineers are managing multiple agent sessions simultaneously, jumping between prompts and code reviews, moving from one incomplete task to another before any of them reach resolution. The flow state that reliable engineering has always depended on, the gradual accumulation of a mental model, the ability to hold a system&#8217;s behavior in mind long enough to reason about it clearly, gets interrupted more frequently and at shorter intervals than at any point in most engineers&#8217; careers.</p><p>&#8220;We became, often, just prompters,&#8221; Darg&#243; said. &#8220;Many of us complained even before that we are living in a world of constant context switching. But it just became even worse. You keep jumping from one window to another, from one meeting to another, because others are also moving faster. At least they think they move faster.&#8221;</p><p>The irony embedded in that observation is significant. The tools promising to accelerate delivery are simultaneously increasing the interruption rate that undermines the deep work required to produce reliable software. Speed and depth are being traded against each other, and the trade is often invisible until the consequences show up in the codebase months later.</p><p>Darg&#243; in our live interview also referenced a research finding that makes the dynamic concrete. Engineers who adopt AI-assisted workflows tend to ship more code early on, because the friction of writing has dropped. But code quality drops alongside it, and the initial speed advantage disappears within a few months as technical debt accumulates faster than it can be serviced. &#8220;In the beginning you ship more code, because it became so much easier. But you don&#8217;t just ship more code. You ship worse code. And that gain in speed is vanishing after a few months because you start accumulating technical debt at the same time. What first seemed faster becomes not faster, but the debt stays,&#8221; Darg&#243; said.</p><p>The answer is not to reject the tools or return to slower workflows. It is to be deliberate about what the tools are being used for and what gets left behind when they are used. Code that was generated quickly but carries no trace of why it is structured the way it is will cost someone considerably when the context is gone. The practices Morley and Darg&#243; both advocate, keeping the hot path isolated, documenting the reasoning behind trade-offs, defaulting to the readable option unless a measurement says otherwise, are not conservative instincts. They are the engineering habits that make fast development sustainable over time rather than just in the short sprint.</p><h2><strong>And so, what this actually adds up to</strong></h2><p>Morley and Darg&#243; are pointing toward the same conclusion from different vantage points: engineering quality cannot be measured by how organized the code looks in the editor.</p><p>Morley&#8217;s measure is hardware efficiency. Does the code respect the physical reality of how the processor accesses memory, and does it make the execution model visible to the next reader, or does it hide it behind abstractions that feel clever now but become maintenance burdens later? Darg&#243;&#8217;s measure is team sustainability. Does the code reduce the cognitive load of the people who maintain it over time, and does it make trade-offs explicit so future engineers and future agents can understand what they would be changing if they touched it?</p><p>Clean code is not a trap because readability is wrong. It is a trap because readability without an understanding of what matters in the specific environment produces systems optimized for the wrong audience. The abstractions that feel clean in the editor are often the ones costing the most in production. And the ones that look strange in a code review are often the ones that matter most to the system&#8217;s actual behavior.</p><p>Not whether it looks clean. But whether it helps the machine run correctly, and whether it helps the next engineer understand why it runs that way. Those two questions do not always have the same answer, but they are always worth asking together, and always worth asking before the code is written rather than after the pull request is merged.</p><p></p>]]></content:encoded></item><item><title><![CDATA[Agentic AI Is Redefining Edge Infrastructure]]></title><description><![CDATA[Agents can't wait. Neither can your infrastructure.]]></description><link>https://deepengineering.net/p/agentic-ai-is-redefining-edge-infrastructure</link><guid isPermaLink="false">https://deepengineering.net/p/agentic-ai-is-redefining-edge-infrastructure</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Wed, 25 Mar 2026 18:13:57 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/280cd0a1-f0d5-42aa-8e78-90ac44439a30_1200x628.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!01dX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!01dX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 424w, https://substackcdn.com/image/fetch/$s_!01dX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 848w, https://substackcdn.com/image/fetch/$s_!01dX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 1272w, https://substackcdn.com/image/fetch/$s_!01dX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!01dX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png" width="1200" height="628" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:628,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:175042,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://deepengineering.substack.com/i/192122268?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!01dX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 424w, https://substackcdn.com/image/fetch/$s_!01dX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 848w, https://substackcdn.com/image/fetch/$s_!01dX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 1272w, https://substackcdn.com/image/fetch/$s_!01dX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Artificial intelligence is entering a new phase with agentic AI, where autonomous systems perceive, decide, act, and learn without constant human oversight, operating independently across distributed environments while collaborating with other agents in real time.</p><p>This shift from centralized AI models to distributed, autonomous agents requires a fundamental rethinking of WAN infrastructure architecture. Previous AI patterns such as centralized training clusters, cloud-based inference, and hub-and-spoke data flows are inadequate for agentic systems that must operate at the edge with speed, autonomy, and resilience.</p><p>And in these environments, the WAN is no longer just a means of connecting branch sites to core data centers. It becomes the essential fabric enabling edge agents to synchronize data, share insights, and coordinate actions, making WAN performance, availability, and adaptability critical to agentic AI effectiveness.</p><h2>Distributed intelligence is edge-centric</h2><p><a href="https://www.linkedin.com/in/leejpeterson">Lee Peterson</a>, VP of Secure WAN Product Management at <a href="https://www.cisco.com/">Cisco</a>, explains where the pressure lands first. Edge environments routinely face unpredictable connectivity, and agents operating in those conditions cannot wait for centralized systems to respond.</p><p>Peterson points to concrete scenarios where this plays out, from autonomous vehicle navigation systems to intelligent manufacturing floors to retail environments where AI agents manage inventory, pricing, and customer experience simultaneously. In each of these cases, he reasons, the decisions that matter most are the ones that have to be made in milliseconds, based on local conditions, often where connectivity to centralized systems is intermittent or constrained.</p><p>But the connectivity assumption is where many organizations get it wrong. Peterson recommends designing for intermittent or constrained WAN conditions rather than treating reliable connectivity as a given, and ensuring real-time path selection for critical systems such as point-of-sale, inventory sync, and IoT devices so that agents can perform automatic remediation during WAN degradation without waiting on human intervention.</p><p>Unlike traditional AI models operating on data in controlled environments, he notes, agentic systems exist in the physical world where latency is measured in milliseconds and decisions have immediate consequences. Sending data hundreds of miles to a cloud data center for processing, Peterson argues, is structurally incompatible with the real-time autonomy these systems require, because the agent must process information, evaluate options, and act locally, right where the action is happening.</p><p>And the scale of coordination compounds this further. A smart city deployment might involve thousands of agents managing traffic flow, energy distribution, and public safety simultaneously, and Peterson underscores that these agents need to share insights and coordinate actions even when network connectivity degrades.</p><p>Organizations that continue to architect around centralized control will find their agentic deployments constrained at precisely the moments that matter most, because this distributed intelligence model is inherently edge-centric and the infrastructure needs to reflect that from the start.</p><h2>Compute at the edge: the foundation of agent autonomy</h2><p>Agentic AI requires compute resources co-located with data sources and decision points, which means deploying high-performance processing across thousands of distributed locations including retail, manufacturing, healthcare, and transportation.</p><p>The workload requirements are diverse and demanding, covering agents performing rapid inference on streaming data, conducting local model fine-tuning based on environmental feedback, and coordinating with peer agents across locations in real time. In retail, Peterson notes, this might translate to supporting smart shelves, computer-vision inventory systems, digital signage, loss-prevention analytics, and customer-flow optimization directly at each store location, which is a significant compute footprint by any measure.</p><p>But powerful edge compute alone cannot deliver the full potential of agentic AI, and Peterson is direct about why. Without equally sophisticated networking, autonomous agents remain isolated, unable to coordinate with peers, synchronize insights, or maintain collective intelligence across distributed environments. The two investments have to be planned together, not sequenced, because the value of edge compute depends almost entirely on the quality of the network that connects it.</p><h2>Networking at the edge: the nervous system of distributed intelligence</h2><p>Just as compute provides the processing foundation for autonomous decisions, networking forms the connective tissue enabling multi-agent coordination. Peterson is specific about what agentic AI requires from it. Low-latency communication between distributed agents, efficient data synchronization, security across untrusted environments, and effective network partitioning are not aspirational requirements but operational ones, and the gap between meeting them and not meeting them is the gap between a functioning agentic system and an isolated one.</p><p>Consider a manufacturing environment where dozens of AI agents coordinate production, where vision systems inspect components, robots adjust operations in real time, and predictive maintenance agents analyze telemetry from across the floor. Peterson uses this kind of environment to ground the networking argument, because these agents must communicate with millisecond latency and maintain coordinated operation even if connectivity to central systems is temporarily lost. His architectural recommendation is specific in that high-performance networking should be integrated directly into edge compute infrastructure to enable agent-to-agent communication with low latency and high bandwidth, rather than routing every interaction through distant aggregation points, because that approach where networking and compute are designed together is what makes real-time coordination possible.</p><p>On security, Peterson is equally precise and equally unambiguous. These systems require cryptographic identity for every agent, encrypted communication, hardware-based roots of trust, and zero-trust architectures designed into both layers from the ground up, ensuring the integrity of autonomous decisions affecting physical systems and human safety in critical infrastructures such as healthcare and transportation. Not as hardening added after deployment, but as a design constraint from day one.</p><h2>The convergence of compute and networking at the edge</h2><p>Peterson frames this moment as an inflection point for enterprise infrastructure strategy, and the practical implication is straightforward even if the work is not. Organizations cannot simply extend cloud architectures to edge locations and expect agentic systems to thrive, because the autonomous, distributed, real-time nature of these systems demands infrastructure where compute and networking are designed together to support local intelligence, agent coordination, and secure operation across thousands of diverse locations.</p><p>And there is a visibility dimension that Peterson adds, one that often gets missed in these conversations. As organizations deploy distributed AI agents across vast, heterogeneous environments, continuous visibility into WAN performance, network health, and application performance at each edge location becomes indispensable, because without it, blind spots undermine the autonomy and resilience that agentic AI requires and teams lose the ability to detect issues proactively, optimize operations, and assure reliable service delivery before degradation affects outcomes.</p><p>Of the choices organizations face right now, Peterson is clear about which ones carry the most weight. Infrastructure decisions made today will determine whether organizations lead this transformation or spend years retrofitting, and the convergence of compute and networking at the edge, he concludes, is the essential foundation upon which the next generation of autonomous, intelligent systems will be built.</p>]]></content:encoded></item><item><title><![CDATA[Benchmarks Are Making AI Coding Look Safer Than It Is ]]></title><description><![CDATA[Passing tests is not proof the code is safe or maintainable.]]></description><link>https://deepengineering.net/p/benchmarks-are-making-ai-coding-look</link><guid isPermaLink="false">https://deepengineering.net/p/benchmarks-are-making-ai-coding-look</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Wed, 04 Feb 2026 18:02:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!uBaj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uBaj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uBaj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 424w, https://substackcdn.com/image/fetch/$s_!uBaj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 848w, https://substackcdn.com/image/fetch/$s_!uBaj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 1272w, https://substackcdn.com/image/fetch/$s_!uBaj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uBaj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png" width="1200" height="628" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:628,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:280880,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://deepengineering.substack.com/i/186883869?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uBaj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 424w, https://substackcdn.com/image/fetch/$s_!uBaj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 848w, https://substackcdn.com/image/fetch/$s_!uBaj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 1272w, https://substackcdn.com/image/fetch/$s_!uBaj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Most technical leaders are optimizing for speed. AI agents now generate code fast enough to reshape how teams ship software. So teams contending with shorter deadlines and shrinking budgets are integrating them into delivery pipelines to increase velocity.</p><p>If you are an engineering leader, you have likely seen the SWE-bench leaderboard. It is the current industry standard for ranking AI coding agents. It scores agents based on whether they can produce a patch that passes a test suite. If it does, the agent gets a gold star.</p><p>But there is a deeper and often overlooked problem that creates a blind spot for enterprise teams.</p><p>Most teams treat these scores like a proxy for real engineering readiness. Speed is not the same as quality, but true velocity is speed plus quality. And passing tests is not the same as writing safe, maintainable code. This then shows up later as security debt, brittle systems, and review fatigue.</p><h2>The Pass/Fail Trap</h2><p>Benchmarks like SWE-bench are designed to test code generation rather than code quality. They ask if the agent can generate a solution that satisfies the immediate requirement.</p><p>They do not ask if the code is maintainable or if it introduces a hidden security vulnerability. They also ignore whether the new code breaks the architectural pattern of the rest of the application.</p><p><a href="https://www.linkedin.com/in/itamarf">Itamar Friedman</a> is the CEO and co-founder of <a href="https://www.qodo.ai">Qodo</a>, the AI Code review platform, says this creates a false sense of security for technical leaders.</p><p>&#8220;SWE-bench is a benchmark that is meant mostly to check code generation capabilities. You can get a really good grade with quite shitty code. It will pass because it implements the requirements and passes the test. But maybe the code is not maintainable. Maybe it includes a security issue.&#8221;</p><h2>The Illusion of Speed</h2><p>In the past, humans wrote code slowly and other humans reviewed it just as slowly. Now that AI agents are writing code at lightning speed, developers are opening two to five times more Pull Requests than they did a year ago.</p><p>This creates a phenomenon called quality rot. Even if AI generates code that is as good as a human&#8217;s, generating ten times more of it means you also generate ten times more bugs.</p><p>Friedman argues that relying on a &#8220;generation benchmark&#8221; to solve this is dangerous. He compares software development to accounting to show why the roles must be separate.</p><p>&#8220;You have bookkeeping and you have auditing. Ideally, you have two different people that are experts. One is doing the bookkeeping and the other is doing the auditing to verify the quality. Using the same agent to do both tasks is counterproductive.&#8221;</p><h2>The Hidden Risk of Review Fatigue</h2><p>When AI agents generate thousands of lines of code in minutes, human reviewers naturally get overwhelmed. They start skimming the code and often trust the AI simply because the test suite passed.</p><p>This is exactly where bugs slip in. A generalist model like GPT-5 might fix a logic bug but accidentally hardcode a credential or use a deprecated library.</p><p>If you rely on the same model to review the code it just wrote, you are essentially asking the fox to guard the hen house. A generalist model might be creative enough to solve the problem, but it lacks the rigid structure needed to audit safety.</p><h3>What You Should Do</h3><p>You need to stop obsessing over which model has the highest SWE-bench score and instead build a system of checks for your AI.</p><p>First, do not trust the generalist model to police itself. You should use specialized agents where one agent writes the code and a completely different agent reviews it against a strict policy.</p><p>Second, you should measure the number of valid bugs your AI catches in PRs rather than just how many PRs it opens.</p><p>Finally, you need to treat your AI pipeline like a government rather than a single employee. Friedman emphasizes that a single agent is never enough to ensure enterprise trust.</p><p>&#8220;You need a system. A system like a country. There are policies, rules, and a police.&#8221;</p><p>The future is not about faster coding but about smarter reviewing.</p>]]></content:encoded></item></channel></rss>