<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Packt Deep Engineering: Engineering Leadership]]></title><description><![CDATA[Deep Engineering talks to a lot of practitioners. Engineering Leadership is where we bring those conversations together to find the patterns that no single interview can surface. These are not interviews. They are the conclusions that the interviews make possible.]]></description><link>https://deepengineering.net/s/engineering-leadership</link><image><url>https://substackcdn.com/image/fetch/$s_!H5BJ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png</url><title>Packt Deep Engineering: Engineering Leadership</title><link>https://deepengineering.net/s/engineering-leadership</link></image><generator>Substack</generator><lastBuildDate>Wed, 12 Aug 2026 19:27:10 GMT</lastBuildDate><atom:link href="https://deepengineering.net/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Packt]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[deepengineering@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[deepengineering@substack.com]]></itunes:email><itunes:name><![CDATA[Packt]]></itunes:name></itunes:owner><itunes:author><![CDATA[Packt]]></itunes:author><googleplay:owner><![CDATA[deepengineering@substack.com]]></googleplay:owner><googleplay:email><![CDATA[deepengineering@substack.com]]></googleplay:email><googleplay:author><![CDATA[Packt]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Materials Science Gets Real Value From Quantum Before Anything Else]]></title><description><![CDATA[Why materials science and small molecule chemistry reach real quantum value first, what higher fidelity simulation changes for drug discovery, and why optimization and cryptography wait.]]></description><link>https://deepengineering.net/p/materials-science-first-real-quantum-value</link><guid isPermaLink="false">https://deepengineering.net/p/materials-science-first-real-quantum-value</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 06 Aug 2026 13:48:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b91e798a-acfd-477e-94d1-3eff12314f55_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><em><span>By </span><a href="https://www.linkedin.com/in/shassinger"><span>Sebastian Hassinger</span></a><span>, author of</span><a href="https://www.packtpub.com/en-it/product/the-new-quantum-era-9781807787370"><span> </span></a><a href="https://www.packtpub.com/en-us/product/the-new-quantum-era-9781807787370"><span>The New Quantum Era</span></a><span> and former quantum lead at </span><strong><span>AWS</span></strong><span> and </span><strong><span>IBM</span></strong><span> | This piece is adapted from his live Deep Engineering session, </span><a href="https://deepengineering.net/p/quantum-computing-beyond-the-hype-sebastian-hassinger"><span>Quantum Computing Beyond the Hype</span></a><span>. Edited by</span><a href="https://substack.com/@saqibjan"><span> Saqib Jan</span></a></em></p></blockquote><p><span>People list the same five domains whenever they ask where quantum computing will actually pay off, so chemistry, materials, optimization, cryptography, and machine learning. </span><strong><span>Of those, materials science is closest to real value, with small molecule chemistry a close second</span></strong><span>, and small molecule chemistry is sometimes treated as a type of materials science anyway. Optimization, cryptography, and machine learning all need thousands of logical qubits, so those are considerably further off.</span></p><p><span>The reason materials leads is that materials science is condensed matter physics. It is about how atoms pack together into a lattice, into a crystalline structure, and how the material behaves as a result of that packing. All of that comes down to physics calculations, which is precisely the kind of work a quantum computer should do well. There is already interesting research going on around battery technology, using quantum information and quantum computing to help with designing and inventing new battery chemistries.</span></p><p><span>What sits behind that work is more ambitious. The aspiration is that if we can simulate materials precisely enough at sufficient scale, we can start creating designer materials. Maybe something much lighter for the same strength, or much stronger for the same weight. Maybe photosynthesis built into the material itself, so you coat your car in something that generates electricity without any cells on the outside. The imagination gets stimulated by the idea of engineering the attributes of a new material at atomic scale, and that is one of the genuinely underestimated parts of this whole story.</span></p><p><span>Chemistry follows for the same underlying reason, because chemistry is quantum mechanical at its core. It is the way atoms interact with one another, and that interaction is a quantum mechanical phenomenon happening at scale, which is what makes it so difficult to simulate precisely. People talk about small molecule chemistry as the next frontier beyond materials, and there is a natural leap from small molecule chemistry to pharmaceuticals, which tend to be larger and more complex molecules. You need a bigger machine to simulate them.</span></p><p><span>Provided we can build one large enough, you can easily imagine drug discovery and research happening inside a very high fidelity simulation, with all the acceleration we are used to getting from computer simulation. The pipeline could compress quite a bit, because you might screen a whole set of drug candidates in simulation before you ever formulate anything, and by the time you are making the drug in the real world you already know how it will interact with the human system.</span></p><p><span>The word doing the work in all of this is </span><strong><span>fidelity</span></strong><span>. Classical simulation of chemistry is approximate by necessity. Methods like</span><a href="https://en.wikipedia.org/wiki/Density_matrix_renormalization_group"><span> DMRG</span></a><span> throw away a great deal of information about the reaction because there is simply too much to calculate, so we have devised ways to focus algorithmically on the small slice of the problem we hope matters most to the answer. The promise of quantum computing is not throwing away as much, and getting a simulation that captures more of the real dynamics of the system accurately.</span></p><p><span>So if you are working out where to point attention, watch the physical sciences rather than your own industry for the first real result. And treat any near term claim about quantum optimization, quantum machine learning, or breaking encryption with the logical qubit count in mind, because those applications are waiting on hardware that does not exist yet.</span></p><div><hr></div><p><strong><span>Read the full issue</span></strong></p><p><span>This piece comes from a longer conversation on how to read quantum progress honestly. The complete interview and the rest of this week&#8217;s </span><strong><span>Deep Engineering</span></strong><span> issue are</span><a href="https://deepengineering.net/s/newsletter-issues"><span> available here</span></a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Eighty Years of Boolean Thinking Is the Real Barrier to Quantum Adoption]]></title><description><![CDATA[Why building quantum intuition inside an engineering organization takes years, what the JPMorgan research model gets right, and where to put effort before the hardware matures.]]></description><link>https://deepengineering.net/p/boolean-thinking-barrier-quantum-adoption</link><guid isPermaLink="false">https://deepengineering.net/p/boolean-thinking-barrier-quantum-adoption</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 06 Aug 2026 13:42:31 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/db4c1548-4f53-4c80-bf42-35e9e3cb67c4_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><em><span>By </span><a href="https://www.linkedin.com/in/shassinger"><span>Sebastian Hassinger</span></a><span>, author of</span><a href="https://www.packtpub.com/en-it/product/the-new-quantum-era-9781807787370"><span> </span></a><a href="https://www.packtpub.com/en-us/product/the-new-quantum-era-9781807787370"><span>The New Quantum Era</span></a><span> and former quantum lead at </span><strong><span>AWS</span></strong><span> and </span><strong><span>IBM</span></strong><span> | This piece is adapted from his live Deep Engineering session, </span><a href="https://deepengineering.net/p/quantum-computing-beyond-the-hype-sebastian-hassinger"><span>Quantum Computing Beyond the Hype</span></a><span>. Edited by</span><a href="https://substack.com/@saqibjan"><span> Saqib Jan</span></a></em></p></blockquote><p><span>It is smart for enterprises to start investing in the skills you need for understanding quantum information, and in exploring potential algorithms now, even though the computers that could actually run anything useful do not exist yet. That sounds premature until you look at what the work actually involves, because it takes a great deal of effort to recast a business problem into quantum terms.</span></p><p><span>We have been thinking about these problems in classical terms, in Boolean logic terms, since the middle of the last century. That is seventy or eighty years now, and it will be a hundred before too long. So we are carrying a very deeply ingrained set of preconceptions about problem solving. We look at our world through the lens of how our laptop could fix the problem in front of us. We write code on those machines, and our minds run along the lines of decomposing the problem, working out what the dynamics are, and figuring out how to represent them efficiently. All of that rests on an invisible reliance on the assumption that at the ground level you are doing Boolean algebra to solve the problem.</span></p><p><span>I do not think many people are fully aware of how much of their problem solving thinking is rooted in the way classical computing works. And it is so different in quantum computing that </span><strong><span>the gap becomes the real obstacle</span></strong><span>. People talk about developing </span><strong><span>quantum intuition</span></strong><span>, which means building the habit of looking at the world through the lens of linear algebra and through the dynamics and capabilities of quantum information. That takes a lot of work, and it is not the kind of work you can compress once the hardware arrives.</span></p><p><span>The smartest enterprise approaches to quantum computing I have seen understand this. They hire a small number of strong people and run research alongside quantum hardware companies and academic researchers at the top of the field. JPMorgan does this well. You can imagine an organization that size has a very large set of challenges, algorithms, and tasks it has to carry out to operate as a financial entity, and the team there looks at theoretical problems that have some mapping to an aspect of those business processes, then does exploratory research against them. They</span><a href="https://www.jpmorganchase.com/about/technology/research/applied-research"><span> publish open science papers</span></a><span>, so everybody benefits from the work they put in.</span></p><p><span>What they are really doing is building the muscle. When quantum technologies mature to sufficient scale, JPMorgan will know how to use them, because the people there will already have years of thinking in the right terms behind them. </span><strong><span>That head start cannot be bought later.</span></strong><span> The hardware will become available to everyone at roughly the same moment, and the differentiator will be whether your organization has anyone who can look at a business problem and recognize the high dimensional structure inside it.</span></p><p><span>Two things follow from this if you are deciding where to put effort right now. Pick one or two people who are genuinely curious and give them real time to build quantum intuition, rather than sending the whole team to an introductory session that changes nothing. And point them at a problem you actually have, something with heavily interconnected variables, so the learning attaches to your business instead of staying abstract.</span></p><div><hr></div><p><strong><span>Read the full issue</span></strong></p><p><span>This piece comes from a longer conversation on how to read quantum progress honestly. The complete interview and the rest of this week&#8217;s </span><strong><span>Deep Engineering</span></strong><span> issue are</span><a href="https://deepengineering.net/s/newsletter-issues"><span> available here</span></a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Site Reliability Engineer (SRE) Interviews Measure a Depreciating Skill]]></title><description><![CDATA[Notable engineering leaders on motivation, ambiguity, tribal knowledge and the measurement problem behind every reliability hire that did not work out]]></description><link>https://deepengineering.net/p/sre-hiring-interviews-measure-depreciating-skills</link><guid isPermaLink="false">https://deepengineering.net/p/sre-hiring-interviews-measure-depreciating-skills</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Wed, 29 Jul 2026 16:08:58 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/eeacb1d8-a1ef-492f-b24c-3a1a426dbe4f_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Reliability teams are hiring against a job description that describes work their agents increasingly do. That mismatch surfaces as a strong candidate who cleared every round and then underperformed, as a loop where three finalists looked interchangeable and the decision came down to feel, or as a new hire who put their first year into building instrumentation nobody had told them was missing.</p><p>For most of the discipline&#8217;s history the binding constraint during an incident was human attention, because somebody had to correlate logs across services, hold a working model of the system, and form a hypothesis fast enough to matter. Interview loops were built to find that person and they were right to. Agents now take a growing share of that first pass, and the loop has not moved with them. A coding exercise, a debugging scenario, and a systems round all measure how fast a candidate converges on a cause. Those are the skills losing value fastest, while the ones that decide whether the hire works go untested.</p><p><a href="https://www.linkedin.com/in/reid-s/"><span>Reid Savage</span></a><span>, Senior Engineering Manager for the Site Reliability Engineering team at </span><a href="https://www.honeycomb.io/"><span>Honeycomb</span></a><span>, shares how he has repeatedly seen that gap widen across the hires he has made and overseen.</span></p><p><span>&#8220;Standard SRE hiring loops are far too skills-based,&#8221; he says. And his own team acted on that conclusion, removing coding exercises from the loop at Honeycomb&#8217;s current stage.</span></p><h2><span>Motivation decides the hire, not the skills matrix</span></h2><p><span>The failure pattern Savage describes does not look like a skills gap, which is what makes it so easy for a conventional loop to miss. &#8220;I&#8217;ve had excellent hires that have never used our cloud of choice, and some who didn&#8217;t work out that had nearly everything,&#8221; he says. The candidates who arrived with the full stack on their r&#233;sum&#233; were not reliably the ones who lasted, and the ones missing the specific platform experience frequently were. That inversion is not a story about credentials being worthless, and it is a story about credentials answering a question that turns out not to be decisive.</span></p><p><span>What Savage found underneath the pattern was a mismatch of a different kind. The problem in the hires that did not work out was &#8220;usually a mismatch between true motivation (which requires significant self-insight on their part) and what the team needs,&#8221; he explains. Motivation in this framing is not enthusiasm for the role or eagerness in the interview, both of which candidates supply readily and neither of which predicts much. It is the specific thing that makes the work satisfying to that person, which requires enough self-knowledge on their part to name it accurately. Savage anchors the distinction to a line he paraphrases from </span><em><span>First, Break All The Rules</span></em><span>, about hiring the accountant who cannot sleep until the books are balanced rather than the one who knows Excel.</span></p><p><span>The scale of Honeycomb&#8217;s problem space is part of why this holds. &#8220;At Honeycomb&#8217;s size, the skills required for each problem we have wouldn&#8217;t fit in 5 job descriptions,&#8221; Savage says, which means any specific skill the loop screens for covers a small fraction of what the hire will actually encounter. A loop optimized for skills coverage is therefore optimizing a variable that cannot be made to matter, because the surface area defeats it. The organizations most exposed to this are the ones whose reliability function touches many systems rather than one, which describes most companies past their first platform consolidation.</span></p><p><span>Honeycomb&#8217;s replacement for the coding exercise is the actionable part. The technical screen is now a pull request review, chosen because, in Savage&#8217;s words, &#8220;it covers the important things: reading between the lines, social skills, eagerness to help, mindset, and integrity.&#8221; A PR review works as a screen because it presents the candidate with someone else&#8217;s decisions rather than a blank editor, and reliability work consists largely of forming a fast, correct read on choices other people made under constraints that are no longer visible. Leaders who want to run this should use a real pull request from their own history, ideally one where the right call was contested, and score the review on what the candidate noticed and chose to raise rather than on whether they found a seeded bug.</span></p><p><span>Savage&#8217;s position is that the technical bar clears more reliably in the presence of those qualities than the reverse, because &#8220;if you have those, and they display the level of technical aptitude you need, they&#8217;ll be able to catch up.&#8221;</span></p><h2><span>Profile of the hire changed inside a year</span></h2><p><span>Fixing the screen addresses the process. The harder change is to the role itself, and it has moved faster than most loops have been revised. &#8220;The profile of a great SRE has fundamentally changed over the last year,&#8221; says </span><a href="https://www.linkedin.com/in/anish-agarwal-io"><span>Anish Agarwal</span></a><span>, co-founder and CEO of </span><a href="https://www.traversal.com/"><span>Traversal AI</span></a><span>, which builds AI systems for production incident investigation. His account of what companies previously optimized for matches what most loops still measure, which is the engineer who could reason through dashboards, logs, and metrics under pressure and arrive at a root cause. That was the correct thing to optimize while humans carried every investigation from alert to resolution.</span></p><p><span>Agarwal&#8217;s argument is that agentic systems now take on the most time-consuming parts of that work, and the consequence for hiring follows directly. &#8220;Raw troubleshooting ability is becoming less of the differentiator than it was even a year ago,&#8221; he says. A capability stops being a differentiator once it stops being scarce, and a loop weighted toward a non-scarce capability produces candidates who cluster indistinguishably at the top. Leaders tend to notice this as a symptom well before they diagnose it, in loops where several finalists all pass the debugging round cleanly and the final decision comes down to feel.</span></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Packt Deep Engineering! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>What replaces it, in Agarwal&#8217;s view, is the part of the role that always belonged there and rarely got attention. &#8220;What has become far more important is the ability to engineer resilient systems,&#8221; he says, being precise about why that capability went underexercised for so long. Most teams never found time for it because troubleshooting consumed their days, so resilience work stayed aspirational rather than scheduled. Two changes now arrive together, with AI removing much of the investigative burden while production systems grow more complex through new services, new dependencies, and new failure modes each quarter. His conclusion is that resilience engineering &#8220;is no longer the work SREs wish they had time for. It&#8217;s increasingly the work the role demands.&#8221;</span></p><p><span>That reframes what a strong candidate has to demonstrate in an interview. Agarwal describes the highest-leverage SREs as the ones &#8220;thinking about what the system will need a year from now, which SLAs actually matter, and which technologies best meet those SLAs given the tradeoffs between cost, reliability, and performance.&#8221; A loop that wants to surface this has to ask forward-looking questions rather than diagnostic ones, and the practical version is to hand the candidate a real architecture with real growth assumptions and ask what breaks first at ten times the load. The answer reveals whether the candidate reasons about failure modes that do not exist yet, which no incident replay exercise can establish. Agarwal adds one further capability he would weight more heavily, describing it as &#8220;AI fluency. Not just expertise, but genuine curiosity,&#8221; and locating its value in engineers who use agents to compress investigations so that the time freed goes into resilience work.</span></p><h2><span>Judgment is the scarce input, and it leaves when people do</span></h2><p>Hiring toward resilience assumes the judgment to do that work can be found and kept, and that assumption meets an organizational problem that no interview round surfaces.<span> </span><a href="https://www.ghantzaras.com/"><span>George Hantzaras</span></a><span>, Director of Engineering, Core Platforms at </span><a href="https://www.mongodb.com/"><span>MongoDB</span></a><span>, speaking at the AI-Powered Platform Engineering workshop we hosted some time back, argued that the problems facing platform and operations teams have stopped being problems of tooling or frameworks and have become problems of people, team dynamics, and how humans interact with the systems they built. His first category is the knowledge gap, where the answers engineers need lie buried across dozens of systems, dated documentation, and sprawling wikis. He described his own developers as having become archaeologists, running scavenger hunts through internal documentation to recover a single answer.</span></p><p><span>The part of that gap most relevant to hiring is the part nobody wrote down. Hantzaras underscored tribal knowledge as the real problem, meaning the unwritten rules and undocumented workarounds that exist only in the minds of the few people who have been at a company long enough to accumulate them. He described one team where new engineers took close to six months to reach full productivity, with the primary cause being the difficulty of navigating internal systems and discovering rules that appeared on no wiki. That knowledge carries a second liability beyond slow onboarding, because it leaves the organization when the people holding it leave, and nothing in a standard hiring process accounts for either effect.</span></p><p><span>Two further problems he named sharpen what the reliability hire is actually for. Golden paths handle roughly the first eighty percent of cases well and then become what he calls golden cages for the remaining twenty percent, which tend to be the most innovative work and the cases where a developer needs something the template will not express. He also described a platform lead whose team was putting more time into maintaining YAML templates for their portal than into building new platform capability, an overhead he calls the tooling tax. Both problems consume senior judgment on work that produces nothing durable, which is exactly the capacity a resilience-focused hire is supposed to create.</span></p><p><span>Hantzaras reaches the case for automating the first pass from the opposite direction, arriving through the operations burden rather than through the hiring loop. He pointed out that the industry has fixated on mean time to recovery while, in his reading, the largest component is mean time to identify, and that teams are drowning in logs, metrics, traces, and alerts without a good way to act on any of it. The resulting alert fatigue leads engineers to tune out noise until a critical alert has a high probability of being lost. His read on responsibility for that is unambiguous, because it is almost never a case of an engineer not working and almost always a case of a system not working as it should, and his conclusion is that filtering data at that volume &#8220;shouldn&#8217;t be a human job. This is a machine job.&#8221;</span></p><p><span>For leaders hiring into this, the implication is that the scarce input is accumulated judgment rather than throughput, and that judgment currently has no home outside individual heads. The practical response runs on two tracks. Weight the loop toward candidates who can reconstruct the reasoning behind decisions they did not make, since that is the skill tribal knowledge recovery actually requires, and a PR review from an unfamiliar codebase tests it directly. Then treat the documentation of undocumented judgment as an explicit deliverable in the first six months rather than as good citizenship, because the alternative is rehiring the same knowledge every time someone resigns.</span></p><h2><span>Foundations come before fluency</span></h2><p><span>Hiring for AI fluency assumes the environment can support it, and that assumption fails more often than the hiring conversation admits. </span><a href="https://www.linkedin.com/in/chankramath"><span>Ajay Chankramath</span></a><span>, CTO of </span><a href="https://platformengineering.org/"><span>Platform Engineering Community</span></a><span>, a member organization for DevOps and cloud native practitioners, led the platform engineering workshop sessions for Deep Engineering recently, where he described platform maturity as a progression through four states, moving from reactive to responsive, then to predictive, and finally to autonomous. His caution for teams reaching for the later states is that observability, runbooks, and service level objectives have to be in place before AI gets layered on top, because introducing AI into an environment without those foundations amplifies the existing chaos rather than resolving it. That ordering has a direct hiring consequence, because a candidate hired for AI fluency into a team without instrumented systems will put their first year into building the substrate instead.</span></p><p><span>Chankramath is equally clear that the tooling does not displace the role. His framing of AI in incident response is that it hands the engineer a capability rather than taking one away, reading the runbook on their behalf and proposing the next step, while the engineer remains the expert who decides whether that step gets taken. He grounds this in an observation most people on call will recognize, which is that nobody reads documentation during a live incident, because when the system is failing and customers are escalating, engineers act on what they already know rather than opening a markdown file. AI closes that gap by surfacing the relevant procedure at the moment of the decision, and the judgment about whether to follow it stays where it was.</span></p><p><span>He also names a failure mode that should change how leaders think about seniority in these hires. Platform capabilities frequently go unused by the most experienced engineers, who treat self-service tooling as something built for juniors who do not know how to do the work directly, and because those engineers are the ones whose opinions carry weight internally, their dismissal propagates. A reliability hire brought in specifically to raise AI adoption will therefore run into resistance from exactly the people whose endorsement determines whether the change holds. Leaders should test for this in the interview by asking how the candidate has previously won over a skeptical senior engineer, rather than assuming enthusiasm for the tooling transfers to the team.</span></p><p><span>The last piece of Chankramath&#8217;s argument concerns measurement, and it constrains how any of this gets evaluated. He treats the standard delivery metrics as lagging indicators, useful for confirming what already happened and poor at telling a team where things are heading, which makes running an organization solely on them a mistake. Leaders hiring for resilience need leading indicators to judge whether the hire is working, because a resilience improvement shows up as incidents that never occurred, and no lagging metric records an absence. </span>The workable substitutes are the observable precursors, meaning coverage of critical paths by service level objectives, the share of alerts that map to a documented response, and error budget burn rate read before the budget is exhausted rather than after.</p><h2><span>Ambiguity behavior predicts the outcome better than anything else</span></h2><p><span>Across everything Savage has observed, one behavior separates the hires that worked from the hires that did not, and it is not a technical one. &#8220;The largest determinant I see is what they do when faced with an ambiguous problem,&#8221; he says, and he lays out the range of possible responses as moving away from it, tagging in someone else, asking for help, owning the solution, helping someone else own it, or investigating why the problem was ambiguous in the first place. The question underneath all of those, in his phrasing, is whether the candidate accepts or rejects unclear problems. Reliability work arrives almost exclusively as unclear problems, which is what makes this behavior more predictive than any skill the loop can measure.</span></p><p><span>The cost of getting it wrong is not usually a visible failure, which is part of why it goes undetected. Savage describes one SRE who was passionate and knowledgeable, and who struggled to bring others into the work or carry projects across the finish line, producing a long trail of initiatives that started without much to show for them. Because SREs touch every part of the business, that pattern compounds in both directions, generating leverage when it works and opportunity cost plus real cloud spend when it does not. </span>His summary of the missing skill is that &#8220;SREs need to know when to open, peek inside, rearrange, and put back problems.&#8221;<span> The judgment is as much about which problems to close as which to open.</span></p><p><span>Testing for this requires giving candidates something genuinely underspecified rather than a puzzle with a hidden solution. The practical version is to present a real ambiguous situation from the team&#8217;s history, withhold the resolution, and pay attention to whether the candidate starts solving or starts questioning the framing. Both responses can be correct, and what matters is that the candidate demonstrates a deliberate choice between them rather than defaulting.</span></p><h2><span>So define the role before you post it</span></h2><p><span>None of the above helps if the role itself was never specified, which Savage treats as the prior question. &#8220;You need to know what you&#8217;re hiring the SRE for,&#8221; he says, and the diagnostic he runs through is a series of concrete organizational facts rather than an abstract job description. </span>Those facts are whether reliability genuinely matters to the business and to what degree, how many nines the company actually needs, whether anyone currently has responsibility for predicting what will fail a year out, whether developers are resisting on-call, and whether the team is drowning in toil.<span> Each of those describes a different job, and an SRE can address any of them, though not all of them at once.</span></p><p><span>The nines question is the one most organizations answer loosely, and it is the one that determines everything downstream. A target expressed as a service level objective with an error budget attached tells a candidate what the job actually involves, because it establishes how much unreliability the business has already agreed to tolerate and therefore where the engineering effort goes. A posting that claims high availability without naming a target is describing an aspiration rather than a role.</span></p><p><span>His standard for handling that honestly is where the recommendation lands. It is fine to need some of those things and not others, as long as the job posting states what the role is expected to be today and how that might change. Most mis-hires that look like candidate failures are role definition failures surfacing late, because a candidate optimized for toil reduction will underperform in a role that turns out to require reliability forecasting, and neither party did anything wrong in the interview. The discipline is to write the posting from the specific problem rather than from a template, and to name the parts of the role that are unsettled instead of leaving them out.</span></p><p><span>The same clarity should extend into the first ninety days, and Savage&#8217;s approach here inverts the usual instrument. The senior hires who succeeded did not treat his 30/60/90 as gospel, and read it instead as an aspirational list with goals attached, oriented toward making social connections, becoming familiar with the code, and starting the long climb through the tech stack. He asks new hires early on to disagree with him and to offer at least one piece of constructive feedback, which he uses to open room for them to work on problems he cannot see himself. That is a deliberate test of the same quality the loop was built to find, because an engineer who will surface an uncomfortable observation in week three is demonstrating the disposition toward unclear problems that predicts everything else.</span></p><p><span>What connects all four positions is that none of them argues for a lower technical bar. The bar holds, and the argument is that the loop currently puts most of its measurement on the part of the bar that agents are absorbing, while the parts that decide whether the hire succeeds go untested. The reliability hire that works out is the person who accepts unclear problems, reasons about failures that have not happened, carries judgment that was never documented, and improves a system nobody else wanted to own. A process built to find that person looks different from the one most companies are still running.</span></p>]]></content:encoded></item><item><title><![CDATA[System Design Hiring Is Really a Judgment Test]]></title><description><![CDATA[Some repeat architecture while others jump straight to scaling, but the ones who stand out reason through failure, cost, and change.]]></description><link>https://deepengineering.net/p/system-design-hiring-judgment-test</link><guid isPermaLink="false">https://deepengineering.net/p/system-design-hiring-judgment-test</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 09 Jul 2026 18:22:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f3576bfb-c703-4d4d-b8c2-6e0f48dea57f_2400x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Tech hiring is tighter, and AI has raised the bar for what companies expect from senior engineers. When companies do open senior engineering roles, they are paying closer attention to quality of hire, adding more signal to the process, and trusting the old formats less.</p><p>And the system design round is where that scrutiny concentrates, because it is the one round an AI assistant will not get you through. The system design interview may look the same as it did a couple of years ago, but the grading rubric underneath it has changed a lot.</p><p><span>Consider </span><a href="https://karat.com/engineering-interview-trends-2026/"><span>Karat&#8217;s 2026 survey</span></a><span> of 400 engineering leaders. It found that AI has widened the gap between strong and weaker engineers rather than closing it, and 73 percent of leaders now say a strong engineer is worth at least three times their total compensation. CoderPad&#8217;s </span><a href="https://coderpad.io/blog/hiring-developers/new-research-the-2026-state-of-tech-hiring-what-ai-means-for-developers-and-hiring-teams/"><span>State of Tech Hiring Report 2026</span></a><span> shows technical assessments up 48 percent globally since mid 2023, with 60 percent of talent leaders naming quality of hire as their top priority for the year.</span></p><p><span>We asked engineering leaders who personally conduct system design loops what they actually look for in a session. These insights are not usually found in prep guides and they converge on a shift most candidates have not yet noticed.</span></p><h2><span>Buzzword-first designs are a huge turn off</span></h2><p><a href="https://in.linkedin.com/in/architagarwal984"><span>Archit Agarwal</span></a><span>, Principal Member of Technical Staff at </span><a href="https://www.oracle.com/"><span>Oracle</span></a><span>, has in his engineering purview interviewed hundreds of engineers, especially for ultra-low-latency authorization work. And the failure he catches most often has nothing to do with AI. &#8220;I&#8217;ve seen a lot of engineers come in to a system design interview and, as soon as I give a problem, they start with &#8216;let&#8217;s use microservices,&#8217; and start using distributed cache,&#8221; he shares, and when he asks how many users they are planning for, the answer rarely matches the architecture they just proposed. &#8220;That is a key difference between any interview-ready engineer and a genuinely good engineer,&#8221; he explains, because &#8220;a genuinely good engineer would not want to implement everything up front.&#8221;</span></p><p><span>Agarwal&#8217;s baseline is quite blunt, that &#8220;if the problem isn&#8217;t complex yet, don&#8217;t overengineer it.&#8221; And he, like many notable leaders, holds the position that &#8220;microservices aren&#8217;t the magical fix that fixes bad architecture. They just distribute that over the network.&#8221; What he seeks instead is one to two minutes of genuine alignment at the start, functional requirements first to establish what is being built and what the user needs, then the nonfunctional requirements that set scale, consistency, and latency. &#8220;Nonfunctional requirements are the ones that decide the architecture&#8212;not the other way around.&#8221; And not every system earns planet-scale treatment, since a tool used only by a company&#8217;s own engineers needs no multi-region deployment, and proposing one tells him the candidate is performing rather than designing.</span></p><p><span>He also weighs cost. It is something that always runs through his evaluation the same way, and he shared with us a line that reframed his own career, &#8220;a good engineer would design for performance, but a great engineer would design for performance per dollar,&#8221; and he pushes candidates to weigh a latency win against the infrastructure bill it creates. That cost lens is exactly where the interview has changed most, because the most expensive and least predictable component on the whiteboard is now the model.</span></p><blockquote><p>Earlier this year, we spoke with Agarwal about <a href="https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews">trade-offs in modern system design</a>, and later published another piece featuring his candidate-side advice on <a href="https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews">why senior engineers fail system design interviews</a>.</p></blockquote><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Treat every AI component as one that will fail</span></h2><p><span>AI is on every whiteboard now, and the reliability assumptions underneath it are exactly what interviewers have started to stress test. </span><a href="https://www.linkedin.com/in/rohitrpoduval"><span>Rohit Poduval</span></a><span>, Senior Software Engineer on </span><strong><span>Prime Video Trust</span></strong><span> and </span><strong><span>Safety</span></strong><span> at </span><strong><span>Amazon</span></strong><span>, opens with a test that did not exist two years ago. &#8220;Does the candidate treat an AI component as an unreliable dependency or a magic box?&#8221; &#8220;I ask them to design a system that uses an LLM. Content classification, recommendation filtering, whatever fits the role,&#8221; he explains, and then he watches what they do with it. &#8220;Weak candidates draw a box labeled LLM and move on, like it&#8217;s a database that always returns the right answer,&#8221; while strong candidates, he shares, &#8220;immediately ask: What happens when it&#8217;s wrong? What&#8217;s the fallback? How do we know it&#8217;s drifting?&#8221;</span></p><p><span>That one reaction most often separates people who have shipped AI systems from people who have read about them. LLMs return different outputs for the same input, and Poduval points out that you can enforce structure on the output, valid JSON or responses from an allowed set, but you cannot assert that the decision itself is correct with a traditional pass and fail test. Candidates who have built these systems talk about evaluation suites, guardrails that constrain what the model can do regardless of its output, and monitoring for drift in production. But candidates who have not will handwave past every one of those concerns. And that handwaving is quite visible within minutes.</span></p><p><span>&#8220;The next thing I probe: how do they test it?&#8221; Poduval shares, because traditional assertions do not apply to AI components. &#8220;Strong candidates talk about benchmark evaluations: a defined set of inputs with expected behavioral boundaries (not exact outputs) that the system must satisfy before shipping.&#8221; They also talk about continuous auditing against real traffic after launch, not just checks before deployment, because a model that passed benchmarks last month might be drifting today. His deeper observation is that with AI systems you often cannot enumerate every right answer before launch, so the old playbook of define requirements, build to spec, and ship no longer closes. The engineers he rates highly design for safe iteration instead, with shadow mode, holdback groups, evaluation frameworks from day one, and human escalation paths, because they expect to keep refining the definition of correct after launch. &#8220;That&#8217;s exactly where the memorized answers fall apart.&#8221;</span></p><p><span>Taking into consideration the resource-intensive workloads and the shrinking budgets, cost has to be something a candidate can defend out loud. Poduval changes a constraint mid session and tells the candidate the system now needs to handle ten times the traffic. &#8220;Candidates who&#8217;ve never dealt with inference costs at scale just say &#8216;add more instances&#8217;. Candidates who&#8217;ve lived it start talking about which requests actually need an expensive model vs which can be routed to cheaper classifiers,&#8221; he shares. They design tiered systems with routing logic that decides in real time which compute budget each request deserves.</span></p><p><span>The practical takeaway for anyone walking into a loop this year is direct, so state the fallback, the evaluation plan, and the per-request cost tier for any AI component before the interviewer has to ask. Everything on that checklist has an architectural counterpart, and we have captured the nuances of each in </span><a href="https://in.linkedin.com/in/sampritimitra"><span>Sampriti Mitra&#8217;s</span></a><span> practical deep-dive on</span><a href="https://deepengineering.net/p/core-architectural-patterns-for-llm-system-design"><span> core architectural patterns for LLM system design</span></a><span>, which walks through the gateway, tiered fallback, model routing, and evaluation patterns behind it.</span></p><h2><span>Constraint changes reveal who patches and who rethinks</span></h2><p><a href="https://www.linkedin.com/in/chandu-p-a5a896118/"><span>Chandu Putta</span></a><span>, Senior Software Engineer at </span><strong><span>MissionSquare Retirement</span></strong><span>, in our email interview shared what earns his confidence. &#8220;The signal that I trust the most is not the architecture a candidate produces, it&#8217;s how they think if I change a constraint in the middle of the interview.&#8221; In a recent loop he had candidates design a document processing pipeline, and midway through he added an LLM call to the extraction step, with variable latency and non-deterministic output. The majority of candidates bolted a retry mechanism onto their existing design and moved on. One candidate stopped and asked, &#8220;What&#8217;s the acceptable failure mode &#8212; an explicit error to the engineer or a silent degradation?&#8221; he recalls.</span></p><p><span>That question, to his mind, is the level marker, because it treats the new component as a design problem rather than a network hiccup. Putta is blunt about how he sizes this up, as &#8220;rehearsed candidates patch the existing design, but the real senior engineers discard it and start fresh with this new anchor.&#8221; The engineers he considers qualified for the job also ask about context boundaries and cost at scale without being prompted, he adds, and they raise a question that prep articles don&#8217;t teach, which is who decides whether the model output is good enough, the technology team or the business team. An engineer who asks that has sat in the meeting where the answer was contested.</span></p><p><a href="https://www.linkedin.com/in/ed-tian"><span>Edward Tian</span></a><span>, creator of </span><strong><span>GPTZero</span></strong><span>, sees the same moment through a different lens. He shares with us how his interviews confirm the pattern from another direction. &#8220;When we see candidates redesign their system following the introduction of a new constraint, we see that they are able to quickly identify the most significant aspect of their original design that has become the bottleneck,&#8221; he explains, and they design new components to replace it.</span></p><p><span>But weak engineers make incremental changes to accommodate the new requirement without revisiting the larger trade-offs the original design was built on. His sharpest distinction is about posture rather than knowledge, because &#8220;most engineers will assume that their original design is now incorrect after the constraint has been added and defend the changes made to it,&#8221; he shares. &#8220;Senior engineers will simply continue to increase the rate of evolution of their design until it meets their new constraint.&#8221;</span></p><p><span>Tian also shares how he has noticed a newer failure mode that improved AI tooling has made common. &#8220;Many candidates can produce beautiful pipeline diagrams for their LLM, but when we dig through more complex failure modes, for example context window, they really struggle,&#8221; and they default to talking about scaling the model. The diagram generated fluency, but probing exposes it.</span></p><p><span>Agarwal looks for something similar when he changes the requirements mid session. The candidate who absorbs the change, restates it in their own words to confirm alignment, and then highlights which parts of the system must change and which remain intact is showing him a structured redesign instinct rather than attachment to a diagram. He also likes to give away something candidates rarely hear from the other side of the table. &#8220;The curveballs that the interviewer gives you will never be in a way that you will have to scrap the complete diagram, the complete architecture, unless you were already off the track.&#8221; A constraint change is a test of surgical judgment, so the working advice is to name out loud what survives and what dies before you touch the whiteboard.</span></p><h2><span>Can you defend the choice you just made?</span></h2><p><a href="https://www.linkedin.com/in/prakharchaube/"><span>Prakhar Chaube</span></a><span>, Senior Software Engineer at </span><a href="https://www.insurgrid.com/"><span>InsurGrid</span></a><span> who previously held the same level at </span><strong><span>Whatfix</span></strong><span> and has conducted more than 25 technical rounds, argues the design round should never hunt for a correct answer. He assesses the thought process and then questions the candidate on their component choices, because that questioning mirrors what happens inside real companies. &#8220;Most companies have some form of an Architectural Review Board, meaning it&#8217;s rare that a new implementation will not be questioned even if it is the right choice.&#8221; Candidates should know why the standard patterns work as standards rather than assuming a plug and play model, and the ones who can stand by their choice under pressure usually survive a real review. In his loops he is categorical about telemetry, since &#8220;if a candidate is not mentioning that in an interview it is a big no,&#8221; because &#8220;without proper observability the system is doomed to fail at scale and cause chaos during debugging.&#8221;</span></p><p><span>Chaube has also built a ladder for assessing AI fluency that moves well past the 2024 questions. &#8220;I first assess the basics like parameters, prompt structure, prompt security, hallucination,&#8221; he shares, then he wants to dig deeper into implementation scenarios drawn from real products, autonomous tasks orchestrated by a central model such as a code review agent or a login and scraping system. The final rung, he explains, is how candidates handle agent-guided coding, because &#8220;I look at their queries and flow&#8221; rather than their claimed experience. This method, he says, reliably separates candidates who prepared just for the interview from candidates who are genuinely strong. The preparation gap shows up in the follow-ups, never in the opening answer.</span></p><p><a href="http://fr.linkedin.com/in/sandor-dargo"><span>S&#225;ndor Darg&#243;</span></a><span>, senior software engineer at </span><a href="https://engineering.atspotify.com/about"><span>Spotify</span></a><span>, has observed the same gap from years of interviewing in the C++ world, and his experience reinforces Chaube&#8217;s from a different stack. He describes a depth gap first, because a language whose standard runs about 2,000 pages guarantees that even genuine experts have blind areas, and the honest move is admitting them.</span></p><p><span>Darg&#243; in </span><a href="https://deepengineering.net/p/clean-c-code-and-the-hidden-cost"><span>our live interview session</span></a><span> shared how he once told an interviewer he did not know a topic and did not want to guess. The interviewer replied that their team did not use it either, and he got hired. His second observation lands harder, as he has seen senior engineers who handle architectural questions with real thoughtfulness struggle to write simple algorithms live under pressure. Rehearse defending your choices against a hostile follow-up, not presenting them to a friendly one, and practice the small problems you assume you have outgrown.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Safety-critical systems raise the same bar higher</span></h2><p><span>With more than 15 years in secure embedded systems, </span><a href="https://www.linkedin.com/in/hareesha-rameshappa/"><span>Hareesha KoratikereRameshappa</span></a><span>, Senior Technical Product Manager at </span><a href="https://phantom.ai/"><span>Phantom AI</span></a><span>, likes to conduct design interviews where the system under discussion can kill someone. He evaluates whether a candidate can translate high-level product requirements into a structured, developable architecture without compromising safety, cost, accuracy, reliability, and power consumption. He looks for system-level thinking across hardware, firmware, application software, and flashing methodologies like OTA. And what he expects, he shares, is that &#8220;the candidate should understand the key elements of system design, such as system and operating modes, error and fault detection, active and history fault mode managements, transition to the safe mode, and recovery from the fault mode.&#8221; They must also supply the inputs verification depends on, KPIs, requirements, corner cases, and they must plan to monitor the system after deployment rather than treating launch as the finish line.</span></p><p><span>His trade-off vocabulary sounds exactly like the web scale interviewers we heard from earlier, and that essentially is the point. Rameshappa presses candidates on functionality versus performance and accuracy versus resource consumption, and he expects them to justify design decisions with data and adjust the design when another team&#8217;s feedback is valuable. At Phantom AI he shared how he struggled to find candidates who combined production experience with sensor selection, calibration integration, and corner case validation, which tells you how rare the full reasoning package is even among experienced engineers.</span></p><p><span>This means that what the interviewers are testing is not a property of a tech stack. It is engineering judgment, and it transfers from a recommendation pipeline to a multi-camera ADAS system without translation.</span></p><h2><strong><span>Trade-off reasoning still decides the level</span></strong></h2><p><span>Everything above rests on a foundation that has not moved, and </span><a href="https://www.linkedin.com/in/dhirendra-sinha"><span>Dhirendra Sinha</span></a><span>, Software Engineering Manager at </span><a href="https://about.google/"><span>Google</span></a><span>, described it </span><a href="https://deepengineering.net/p/designing-for-scale-and-resilience"><span>last year in one of our conversations</span></a><span> with him and Tejas Chopra. &#8220;It&#8217;s easy to choose between good and bad solutions,&#8221; Sinha told us, &#8220;but senior engineers often have to choose between two good options. I want to hear their reasoning.&#8221; He also carries a line from a chief architect at Yahoo that reframes how scale changes judgment, because when an engineer dismissed a corner case as one in a million, the architect replied, &#8220;One in a million happens every hour here.&#8221; Scale invalidates assumptions, and mostly every leader in this piece is testing whether a candidate&#8217;s assumptions survive contact with it.</span></p><p><span>Agarwal&#8217;s version of the same evaluation focuses on where a candidate chooses to go deep. When a candidate takes one area, distributed storage or authentication or performance engineering, and drives into real depth, he takes that as an engineer who &#8220;understands the gravity of things&#8221; rather than one collecting surface vocabulary. And he calibrates the questions a candidate asks against their level, because clarification questions are always welcome, but a flood of questions too basic for someone with a strong resume tells him the candidate has never actually thought about the system in front of them.</span></p><p><a href="https://www.linkedin.com/in/chopratejas"><span>Tejas Chopra</span></a><span>, Senior Engineer at </span><strong><span>Netflix</span></strong><span>, closes the loop between the old test and the new one with a single follow-up he has asked for years. When a candidate picks SQL over NoSQL or strong over eventual consistency, he asks what changes if the user base grows tenfold, and the AI era version of that question is exactly what some interviewers in technical rounds now ask about model cost and failure modes. The component changed, but the muscle is the same. For readers who want the underlying vocabulary these interviews keep invoking, consistency, availability, partition tolerance, and the read and write trade-offs that connect them, Sinha and Chopra&#8217;s practical deep-dive on</span><a href="https://deepengineering.net/p/distributed-system-attributes"><span> distributed system attributes</span></a><span> from their book </span><a href="https://www.packtpub.com/en-in/product/system-design-guide-for-software-professionals-9781805124993"><span>System Design Guide for Software Professionals</span></a><span> walks through all of it with worked examples.</span></p><h2><strong><span>What to do in your next loop</span></strong></h2><p><span>Ask for the acceptable failure mode before you design anything, because that single question outranks any diagram you will draw. Treat every AI box in your architecture as a component that will be wrong, and price its errors with a fallback, an evaluation plan, and a cost tier per request class. When a constraint changes mid session, say out loud which parts of your design survive and which parts die before you redraw anything. Say what you would monitor without being asked, and say I don&#8217;t know about the right things, because every interviewer quoted here treats honesty about limits as seniority rather than weakness.</span></p><p><span>The rehearsed candidate optimizes for the first answer, and the bar has moved to the third follow-up. That is the whole shift.</span></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for free</strong> to get expert-led deep dives, system design breakdowns, and full book chapters like this one every week.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p><strong><span>Contributors</span></strong></p><ul><li><p><strong><span>Archit Agarwal</span></strong><span>, Principal Member of Technical Staff, Oracle.</span><a href="https://in.linkedin.com/in/architagarwal984"><span> LinkedIn</span></a><span>. Read his candidate-side advice in</span><a href="https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews"><span> Why Senior Engineers Fail System Design Interviews</span></a><span>.</span></p></li><li><p><strong><span>Rohit Poduval</span></strong><span>, Senior Software Engineer, Amazon (Prime Video Trust &amp; Safety).</span><a href="https://www.linkedin.com/in/rohitrpoduval/"><span> LinkedIn</span></a></p></li><li><p><strong><span>Naga Chand Putta</span></strong><span>, Senior Software Engineer, MissionSquare Retirement. </span><a href="https://www.linkedin.com/in/chandu-p-a5a896118/"><span>LinkedIn</span></a></p></li><li><p><strong><span>Edward Tian</span></strong><span>, Founder, GPTZero. </span><a href="https://www.linkedin.com/in/ed-tian"><span>LinkedIn</span></a></p></li><li><p><strong><span>Prakhar Chaube</span></strong><span>, Senior Software Engineer (SDE-3), InsurGrid. </span><a href="https://www.linkedin.com/in/prakharchaube"><span>LinkedIn</span></a></p></li><li><p><strong><span>Hareesha KoratikereRameshappa</span></strong><span>, Sr. Technical Product Manager, Phantom AI. </span><a href="https://www.linkedin.com/in/hareesha-rameshappa/"><span>LinkedIn</span></a></p></li><li><p><strong><span>S&#225;ndor Darg&#243;</span></strong><span>, Senior Software Engineer, Spotify.</span><a href="https://sandordargo.com/"><span> sandordargo.com</span></a></p></li><li><p><strong><span>Dhirendra Sinha</span></strong><span> (Google) and </span><strong><span>Tejas Chopra</span></strong><span> (Netflix), authors of </span><em><span>System Design Guide for Software Professionals</span></em><span> (Packt). Read their</span><a href="https://deepengineering.net/p/distributed-system-attributes"><span> free chapter</span></a><span>.</span></p></li></ul><p><strong><span>Sources</span></strong></p><ul><li><p><span>Karat, &#8220;Engineering Interview Trends in 2026,&#8221; March 2026. </span><a href="https://karat.com/engineering-interview-trends-2026/"><span>https://karat.com/engineering-interview-trends-2026/</span></a></p></li><li><p><span>CoderPad, &#8220;State of Tech Hiring Report 2026,&#8221; March 2026. </span><a href="https://coderpad.io/blog/hiring-developers/new-research-the-2026-state-of-tech-hiring-what-ai-means-for-developers-and-hiring-teams/"><span>https://coderpad.io/blog/hiring-developers/new-research-the-2026-state-of-tech-hiring-what-ai-means-for-developers-and-hiring-teams/</span></a></p></li><li><p><span>Deep Engineering, &#8220;Why Senior Engineers Fail System Design Interviews,&#8221; May 2026. </span><a href="https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews"><span>https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews</span></a></p></li><li><p><span>Deep Engineering #36, Archit Agarwal on System Design Trade-offs, February 2026. </span><a href="https://deepengineering.substack.com/p/deep-engineering-36-archit-agarwal"><span>https://deepengineering.substack.com/p/deep-engineering-36-archit-agarwal</span></a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[How SUSE Runs AI Without Losing Control]]></title><description><![CDATA[Open source enterprises treat data sovereignty, MCP governance, and cost predictability as one connected problem]]></description><link>https://deepengineering.net/p/how-suse-runs-ai-without-losing-control</link><guid isPermaLink="false">https://deepengineering.net/p/how-suse-runs-ai-without-losing-control</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Wed, 24 Jun 2026 20:20:29 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c14a30e1-69ef-4264-9338-6a7b82cc9e59_2760x1240.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most companies adopting AI are making a quiet trade in exchange for speed. They accept that their data passes through systems they do not control, priced on terms they did not set, governed by rules someone else can change. For a consumer app that trade is often fine. But for an enterprise running infrastructure under compliance regimes, it is not. <a href="https://www.linkedin.com/in/rickspencer3">Rick Spencer</a>, General Manager for Technology and Product at <a href="https://www.suse.com/">SUSE</a>, has for the last two years been working out how an open source enterprise adopts AI at scale without making that trade, and his approach is a useful template for any organization that takes data sovereignty, auditability, and cost predictability seriously.</p><p>SUSE&#8217;s position is unusual in a way that sharpens the problem. &#8220;All the software that we write is open source,&#8221; Spencer explains. &#8220;We&#8217;re not worried about, oh, we leaked the code, we publish the code.&#8221; The concern is the data that belongs to others. When an engineer debugs a customer environment, the logs are not SUSE&#8217;s to hand to a third-party model. &#8220;They trust us to not do those kinds of things,&#8221; he says, and that trust is the thing the entire approach is built to protect.</p><div><hr></div><p><em>You can watch our full interview or read the <a href="https://deepengineering.substack.com/p/sovereign-ai-agentic-infrastructure-rick-spencer-suse">Q&amp;A article here</a>.</em></p><div id="youtube2-8PdtwqLL6YI" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;8PdtwqLL6YI&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/8PdtwqLL6YI?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div><hr></div><h2><strong>Sovereignty means how, not whether</strong></h2><p>The first instinct when AI collides with compliance is to draw a boundary around where AI is allowed to go. Spencer rejects that framing, saying &#8220;I don&#8217;t think it&#8217;s can or cannot. It&#8217;s more so how.&#8221; The distinction matters because a can-or-cannot policy ends with engineers either blocked from useful tools or quietly routing around the rules. A question of how keeps the capability available while controlling the conditions under which it runs.</p><p>The clearest illustration is SUSE&#8217;s build process. A lot of what the company ships is built in an internal instance of Open Build Service, and the defining property of those builds is that they happen offline. &#8220;All the builds are offline. They literally are not connected to the internet,&#8221; he points out. &#8220;This is super important because you need to be able to prove that nothing happened during the build process.&#8221; Proving a negative is far easier when there was no connection through which anything could have happened. It means doing things the hard way, making sure every source is present ahead of time because nothing can be pulled live during the build, but the payoff is provable integrity.</p><p>Applying AI inside that environment is where the how becomes concrete. The example Spencer gives is backporting a patch to previous stable releases, the kind of repetitive, knowledge-intensive work AI is well suited to. The question is whether it can be done in a sovereign way, and his answer is yes, on conditions. &#8220;As long as we are running AI in a way that it&#8217;s able to run disconnected from the internet, and we can have complete visibility into everything it&#8217;s doing.&#8221; SUSE goes further than running existing models in isolation. &#8220;In some cases we even train our own models to accomplish these things,&#8221; he shares, &#8220;and that way we know the model doesn&#8217;t have some naughty time bombs built into it.&#8221; Because owning the model end to end removes the last category of thing the enterprise would otherwise have to take on trust.</p><p>The foundation under all of this is SUSE AI, the company&#8217;s own stack for running AI workloads on private infrastructure, which it uses heavily internally and runs Llama on. &#8220;It&#8217;s all within our private infrastructure, so we make sure there&#8217;s no chance that any data can escape,&#8221; Spencer says. &#8220;We only use models which can be vetted effectively.&#8221; The principle is consistent. Keep the data inside the boundary, and only run models you can actually inspect.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Packt Deep Engineering! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><strong>MCP as the control layer, not just the connector</strong></h2><p>Most of the conversation about Model Context Protocol treats it as plumbing that turns a chatbot into an agent by giving it tools to act with. Spencer agrees that is what it does, but his more interesting argument is that MCP is where an enterprise regains the control that autonomous agents would otherwise erode. &#8220;MCP servers do another really important thing,&#8221; he explains, &#8220;which is provide a place where you can, as an enterprise, bring some sanity and control to the usage.&#8221;</p><p>The mechanism is straightforward once you see the server as infrastructure rather than glue. &#8220;If you have MCP servers running, they&#8217;re just servers,&#8221; Spencer says. &#8220;That means you can provide access ACLs to them.&#8221; A server can be told that a given user&#8217;s agent may use these tools and not those. The usage can be logged. Gateways can sit in front, and the company runs its own alongside a partnership with StackLok. The architectural rule that holds it together is that the language model never touches tools directly. &#8220;You don&#8217;t give the LLMs access directly to tools, only the MCP servers,&#8221; he reasons, &#8220;and then you can have that oversight, meet your compliance needs.&#8221;</p><p>He takes the containment idea down to the operating system. &#8220;You can put the MCP server, I call it, in jail,&#8221; Spencer says, describing a systemd process scoped to present only the compute resources the server actually needs. The reasoning is a security posture rather than a convenience. &#8220;For every MCP server you&#8217;re running, there&#8217;s an LLM out there that&#8217;s trying to use it, and who knows what kind of prompt injections people are running.&#8221; The same boundary defends against the model&#8217;s own failures, not only malicious input. An agent cannot delete a production server with a tool it was never given. &#8220;They guard against things like the AI hallucinating something and deleting your production server, because you simply don&#8217;t provide that tool to it,&#8221; he warns. Control here is not a policy document. It is the set of tools the agent is and is not handed.</p><p>There is a quality dimension to MCP that reinforces the control argument, because a well-built server encodes expert knowledge rather than leaving the model to guess. SUSE ships MCP servers with its products, crafted by the people who know those products best. &#8220;It would be like, instead of you sitting down in front of a chatbot saying I need to figure out how to use Rancher, you&#8217;re sitting down with the whole Rancher development team telling you how to prompt the chatbot,&#8221; Spencer says. An agent working against an expert-built server is not interpreting raw APIs and making guesses a human cannot easily validate. The encoded knowledge makes the agent both more capable and more predictable, which is control of a subtler kind.</p><h2>Cost sovereignty belongs in the same conversation</h2><p>The control story is incomplete if it stops at data and governance, because an enterprise that cannot predict its AI spend has lost a different kind of control. Spencer folds cost into the sovereignty argument directly. &#8220;Sometimes they call it cost sovereignty,&#8221; he says, &#8220;because no one can come back later and say, oh, by the way, we&#8217;re changing our model.&#8221; He has watched a supplier move developers from seat-based to usage-based pricing, a shift his teams did not control and could not prevent. Hosting your own infrastructure changes the nature of the exposure. &#8220;There&#8217;s a maximum cost there,&#8221; he notes of self-hosted AI, and the question shifts from whether you will overrun a variable bill to whether you have the observability to use a fixed capacity fully.</p><p>For the cases where variable cost is unavoidable, SUSE uses circuit breakers that cut off runaway spend in real time when usage spikes past a threshold in a given minute. Spencer is honest that this frustrates engineers who get rate limited mid-task, but the alternative is an autonomous agent running up cost with no human in the loop. The same discipline runs through the company&#8217;s use of frontier models, reserved for high-value strategic work rather than routine completion, and through the practice of using a frontier model once to build an agent that then runs on a cheaper model or a private one. Each of these is the same instinct expressed at the level of money, keeping the expensive capability available for the work that justifies it and putting hard limits around everything else.</p><p>What makes the SUSE approach instructive beyond its own walls is that none of it depends on slowing engineers down. The sovereignty, the MCP governance, and the cost controls exist so that engineers can move fast inside a boundary the enterprise can actually stand behind. &#8220;You don&#8217;t want to stop them from getting that 100X improvement,&#8221; Spencer says. &#8220;You need to give them the right tools for the job.&#8221; Control, in his telling, is not the opposite of speed. It is the thing that makes speed safe to allow.</p><p></p>]]></content:encoded></item><item><title><![CDATA[Compute Obsession Is Slowing Down AI Systems]]></title><description><![CDATA[Why data movement costs more than computation and what most engineers building AI systems are getting wrong because of it]]></description><link>https://deepengineering.net/p/compute-obsession-slowing-down-ai-systems</link><guid isPermaLink="false">https://deepengineering.net/p/compute-obsession-slowing-down-ai-systems</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Tue, 26 May 2026 05:30:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1d53f5cf-8656-4090-8fec-27c78e53b139_690x330.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Engineers building AI systems today tend to focus on compute first. It is typically about how many GPU cores, how many parameters, how much VRAM, and how to extract more from all of it. While the benchmarks are about throughput and inference speed, the infrastructure conversations are about scaling horizontally across more hardware.</p><p><a href="https://www.linkedin.com/in/jimledin">Jim Ledin</a>, a seasoned engineering leader, CEO of <a href="https://ledin.com/">Ledin Engineering</a> and author of <a href="https://www.packtpub.com/en-us/product/modern-computer-architecture-and-organization-9781806028023">Modern Computer Architecture and Organization</a> (third edition, Packt), thinks that framing misses the most important constraint in production AI systems. The bottleneck holding back real-world AI performance is not compute but data movement.</p><p>&#8220;Data movement can often be more expensive than the actual computation steps,&#8221; Ledin says. &#8220;The latency, especially moving large data structures across different levels of the memory hierarchy, can dominate and leave a lot of your compute bandwidth idle.&#8221; This is not a niche embedded systems concern. It is happening in the largest AI deployments in the world, and it is the reason hardware vendors like NVIDIA are designing systems the way they are today.</p><blockquote><p><strong>Continue reading</strong> or watch the full conversation with Jim Ledin below.</p></blockquote><div id="youtube2-Q21CJXfQLOk" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;Q21CJXfQLOk&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/Q21CJXfQLOk?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p></p><h2>Memory bandwidth is slowing your AI system more than your GPU is</h2><p>When a CPU or GPU requests data from memory and that data is not available in cache, the processor waits while the computation units sit idle. In a consumer application, that idle time seems like a minor inconvenience. But in an AI system processing large tensors continuously, it accumulates into a significant fraction of total runtime.</p><p>&#8220;AI workloads are becoming increasingly memory bandwidth limited,&#8221; Ledin shares, pointing to a dynamic that is reshaping how AI hardware gets built. &#8220;It is taking more time to bring data into the GPU or TPU memory than it is taking for the computation to take place on the data.&#8221; The raw ability to multiply matrices is no longer the binding constraint. But getting the data to the multipliers fast enough is.</p><p>This is exactly why high bandwidth memory exists. HBM modules are stacks of RAM chips built into a cube, physically close to the processing units, with far higher data transfer rates than conventional DRAM. &#8220;On a TPU card, you typically have several of these HBM modules,&#8221; Ledin explains, &#8220;and they have a far higher data rate for transferring data in and out of the GPU processing components than on a typical consumer grade GPU.&#8221; The engineering bet being made with systems like NVIDIA&#8217;s Blackwell architecture is that memory bandwidth is worth more than raw core count, because the cores are already faster than the data can reach them.</p><p>But there is a side effect that touches anyone buying consumer hardware. &#8220;A lot of the production capacity for memory is going into these high bandwidth memory modules, which cost a lot more for the purchaser and make a lot more money for the vendor,&#8221; Ledin observes. That is a direct reason DDR5 has been difficult to find and expensive when available. The memory fabs are prioritizing the more profitable HBM production, and consumer DRAM is downstream of that decision.</p><h2>The hardware cost your cloud bill is hiding</h2><p>Most software engineers, especially those working in cloud environments, treat the hardware as someone else&#8217;s concern. The abstraction is good enough, the managed services handle the infrastructure, and the code runs somewhere. Ledin&#8217;s argument is that this hands-off relationship with hardware has a real cost that shows up in performance and in cloud bills.</p><p>&#8220;If your code is accessing memory in inefficient patterns, if you are not using the cache memory within the processor in an effective manner, and if you are just moving data around more than is necessary, that can all have significant performance impacts,&#8221; he warns. The CPU requests data from memory, and if it is not in cache, it waits. &#8220;A lot of the time it is unavoidable, but the amount of latency can be minimized by different ways of optimizing algorithms.&#8221;</p><p>The mechanics are specific. When a modern CPU reads from DRAM, even a single byte triggers a 64-byte cache line transfer. The processor brings in a block of adjacent memory whether it needs all of it or not. If the algorithm then jumps to a different memory location, causes that block to be evicted from cache, and later needs it again, it has to re-read it from DRAM. That is wasted time. &#8220;For best efficiency, you would want your code to be working with data from that block before it moves on to something else,&#8221; Ledin explains, &#8220;rather than bouncing around to other memory locations.&#8221;</p><p>In a cloud environment, this inefficiency does not just slow things down. It costs money, and there is no incentive for cloud providers to surface it clearly. &#8220;You are paying for the usage of the system whether the CPU is actually crunching instructions or the CPU is idle waiting for a data item to come in from memory,&#8221; he points out. The cloud bill does not distinguish between productive cycles and stall cycles. Engineers who understand cache locality can write code that reduces stalls and therefore reduces cost, not just latency. Optimizing for cost comes down to understanding your memory access patterns and engineering around them, not just choosing the right managed tooling stack.</p><p>Drawing from his engineering work across embedded and production systems, Ledin shares a useful example. A Linux web server called Tux, which ran in kernel space to avoid user-to-kernel data transfers, developed a performance problem under high load because its per-request state data grew large enough to exceed the CPU&#8217;s level two cache. &#8220;Performance dropped off sharply,&#8221; he recalls. Engineers analyzed the cache behavior, restructured the data layout to keep per-request state smaller, and did the same for instruction caching by batching related processing together. &#8220;Fixes that they implemented increased the application performance by about 40%.&#8221; No new hardware, no architectural overhaul. Just understanding where the memory ceiling was and designing around it.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Deep Engineering publishes weekly expert-led insights on systems design, architecture, and real-world engineering. <strong>Subscribe for free!</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>GPUs are the right tool, but not always for the reason you think</h2><p>The assumption that GPUs are the correct architecture for AI workloads is not wrong, but it is incomplete in a way that matters for engineers making infrastructure decisions. Ledin draws a distinction that is often glossed over in the mainstream conversation about AI hardware.</p><p>&#8220;GPUs are probably the ideal architecture today for people and small companies that want to run language models locally,&#8221; he says, drawing from personal experience. He recently ran the Gemma 4 26-billion-parameter model on an NVIDIA RTX 4090, and for that use case the GPU is the right tool. But for larger-scale deployments running the much larger frontier models, the picture is different. &#8220;The trend there is for dedicated TPUs,&#8221; he notes.</p><p>The distinction matters because GPUs carry silicon dedicated to graphics work that has nothing to do with tensor operations. A consumer GPU has hardware for real-time video rendering, gaming pipelines, and display output. A TPU does not. &#8220;TPUs do not use up silicon for that purpose and focus everything on the tensor work,&#8221; Ledin explains. When you are running thousands of inference requests at scale, that difference in silicon allocation translates directly into efficiency at the workload that actually matters.</p><p>There is also the SIMT execution model to understand. Modern NVIDIA GPUs run 32 threads in lockstep, all executing the same instruction on different data streams simultaneously. This is efficient for linear, parallel workloads. When those threads hit a branch, a conditional where some threads take the if path and some take the else path, the hardware executes one side then goes back and executes the other. &#8220;You basically have effectively a pipeline stall where it has to go back and execute a different thread in that kind of situation,&#8221; Ledin highlights. The flexibility is there, but it comes at a cost. &#8220;Avoiding branching if possible can have a significant impact on performance.&#8221;</p><p>For engineers deciding where to run inference workloads, Ledin offers a practical heuristic. &#8220;The GPU only really becomes attractive when you have enough work for it to do that it can be parallelized and enough that it will amortize the costs associated with moving data onto the GPU, launching the kernels, and doing the management work to transfer data to and from the GPU.&#8221; If the workload is not large enough to keep the GPU busy, the CPU implementation may be faster because it avoids all that overhead entirely.</p><h2>Frameworks are hiding costs that engineers need to see</h2><p>Frameworks and libraries have made it possible to build sophisticated AI systems without ever thinking about what is happening in hardware. That is mostly a good thing. The abstraction accelerates development and reduces mistakes. But there is a point where abstraction stops being a benefit and starts hiding costs that need to be visible.</p><p>&#8220;Where it becomes dangerous to use too much abstraction is when it obscures what is happening with the data layout in memory and the execution patterns,&#8221; Ledin cautions. In performance-critical applications, the framework is making decisions about how data is structured and how the processor interacts with it. If the engineer does not know what those decisions are, they cannot tell when they are working against the hardware.</p><p>The practical approach Ledin recommends is a two-layer architecture. &#8220;Use the most expressive code at the edges of the system, and in the core, use more performance-aware code.&#8221; The boundary between those layers is not always obvious in advance, and finding it usually requires benchmarking rather than reasoning. But the principle is clear: abstractions are appropriate where they preserve meaning across the team, and they become a problem where they hide costs that affect the system&#8217;s ability to meet its requirements.</p><p>One specific pattern worth knowing is the array of structures versus structure of arrays tradeoff. A common data layout is an array of objects, where each object holds all the fields for one entity. For CPU cache efficiency, it can be significantly better to restructure this as a structure of arrays, where each field is stored as a separate array for all entities. &#8220;That might have a big impact on performance,&#8221; Ledin notes, because the CPU cache loads contiguous memory, and if the algorithm is operating on one field across many entities, the structure of arrays layout means each cache load is full of useful data rather than fields the algorithm is not touching.</p><h2>The skills that will matter when the hardware changes again</h2><p>The specific technologies that matter five years from now are difficult to predict, and Ledin is honest about that. &#8220;Four years ago when the previous version of my book came out, it was not at all clear to me, or I think a lot of people, what was going to be happening with AI in the coming years,&#8221; he says. Predicting which hardware architectures or AI frameworks will dominate is not the point. Building the mental model to understand them when they appear is.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.packtpub.com/en-us/product/modern-computer-architecture-and-organization-9781806028023" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qFZO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 424w, https://substackcdn.com/image/fetch/$s_!qFZO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 848w, https://substackcdn.com/image/fetch/$s_!qFZO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 1272w, https://substackcdn.com/image/fetch/$s_!qFZO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qFZO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775" width="336" height="414.46153846153845" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:336,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Modern Computer Architecture and Organization&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.packtpub.com/en-us/product/modern-computer-architecture-and-organization-9781806028023&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Modern Computer Architecture and Organization" title="Modern Computer Architecture and Organization" srcset="https://substackcdn.com/image/fetch/$s_!qFZO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 424w, https://substackcdn.com/image/fetch/$s_!qFZO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 848w, https://substackcdn.com/image/fetch/$s_!qFZO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 1272w, https://substackcdn.com/image/fetch/$s_!qFZO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae0cb75-2765-4af0-bdaa-f6514e4ac96c_2250x2775 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>3rd Edition</strong></figcaption></figure></div><p>The foundational skill is the ability to reason across abstraction layers. &#8220;The way to really understand the system requires the ability to reason across all of the abstraction layers from the software framework that you are working on at the top level, all the way down to the hardware that runs the code,&#8221; Ledin underscores. That does not mean reading assembly code for every application. It means understanding how pipelines and caches work and orienting code to work within those environments rather than against them.</p><p>The other shift is heterogeneous computing. Writing code that runs on a CPU is no longer sufficient context for many engineering problems. &#8220;It is also becoming more critical to understand heterogeneous computing environments,&#8221; Ledin says. &#8220;It is not just writing code that runs on a CPU. You might also have code that interacts with the GPU if you are running a parallelized algorithm on that, whether it is a language model or something else.&#8221; Domain-specific accelerators, TPUs, RISC-V implementations, and specialized inference chips are all becoming part of the environments that production engineers have to reason about. The engineers who will be most effective in that landscape are the ones who understand why those architectures make the tradeoffs they do, not just how to call their APIs.</p><div><hr></div><p><em>This article is based on <strong>Deep Engineering #46</strong>. You can read the full issue, including additional insights from Jim Ledin on modern computer architecture and AI infrastructure,</em></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;b15102d9-d90c-4b9c-a26f-9fb34ceb12fa&quot;,&quot;caption&quot;:&quot;Memory bandwidth, GPU trade-offs, and the infrastructure decisions that determine whether AI systems are resilient up in production.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #46: Jim Ledin on Modern Computer Architecture and the AI Infrastructure Layer&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:284339563,&quot;name&quot;:&quot;Jim Ledin&quot;,&quot;bio&quot;:&quot;Embedded system developer and cybersecurity tester. Author of \&quot;Modern Computer Architecture and Organization\&quot; and \&quot;Architecting High-Performance Embedded Systems.\&quot;&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d7d3f2aa-9d0a-426d-b514-634881eff942_144x144.png&quot;,&quot;is_guest&quot;:true,&quot;bestseller_tier&quot;:null,&quot;primaryPublicationSubscribeUrl&quot;:&quot;https://jimledin.substack.com/subscribe?&quot;,&quot;primaryPublicationUrl&quot;:&quot;https://jimledin.substack.com&quot;,&quot;primaryPublicationName&quot;:&quot;Jim Ledin&quot;,&quot;primaryPublicationId&quot;:8979199}],&quot;post_date&quot;:&quot;2026-05-07T15:03:11.734Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8adf3738-2898-4207-9a54-e3709b1e9c3c_850x400.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.substack.com/p/issue-46-jim-ledin-computer-architecture-ai-infrastructure-layer&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:196768803,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[Why Senior Engineers Fail System Design Interviews]]></title><description><![CDATA[Scoping, adaptability, and being exceedingly clear matter more than just knowing your tech stack in system design interviews.]]></description><link>https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews</link><guid isPermaLink="false">https://deepengineering.net/p/why-senior-engineers-fail-system-design-interviews</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Tue, 19 May 2026 20:19:10 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/bb2a37a3-e396-4fe6-8a55-f331e5572a15_1418x736.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most engineers presume that because they know their tech stack well enough, the system design interview will be easy. And why should they think any differently. They have shipped distributed systems at scale, debugged race conditions at 3am, and made the architectural calls that kept production stable under pressure. But then they walk into a system design interview with confidence and walk out having failed, often without understanding exactly why.</p><p><a href="https://in.linkedin.com/in/architagarwal984">Archit Agarwal</a>, Principal Member of Technical Staff at <a href="https://www.oracle.com/">Oracle </a>where he builds ultra-low-latency authorization services in Go, has interviewed hundreds of engineers. His observation about why experienced engineers fail is the most direct and honest assessment of this problem: they do not fail because they do not know what Kafka is or how DynamoDB handles consistency. They fail because of how they communicate. That single factor, how clearly and deliberately an engineer narrates their thinking, determines the outcome of most system design interviews more than any technical knowledge does.</p><h2>They jump to solutions before understanding the problem</h2><p>Agarwal described a pattern he sees play out repeatedly across interviews at every level of seniority. An interviewer gives a problem and within thirty seconds the candidate is already saying &#8220;I&#8217;ll use Redis, I&#8217;ll use Kafka, let&#8217;s go with microservices.&#8221; The interviewer has not said anything about scale. The candidate has not asked how many users the system needs to support, whether it is read-heavy or write-heavy, what the latency requirements are, or whether there are compliance constraints based on the geography of operation. They have skipped the part of the conversation that actually determines what should be built.</p><p>Those questions are not warm-up questions. They are the questions that drive the architecture. Nonfunctional requirements determine architecture, not the other way around. How many requests per second, what consistency model you need, whether you have a strict latency ceiling, these are the inputs. The architecture is the output. Engineers who skip to the output without gathering the inputs are designing in a vacuum, and the interviewer can see it the moment it happens. Agarwal&#8217;s recommendation is to spend the first one to two minutes of any system design interview doing nothing but alignment: gather functional requirements on what is being built and what the user actually needs, then gather nonfunctional requirements on scale, consistency, latency, and compliance. If you ask the right questions in those first two minutes, Agarwal says, you have already impressed the interviewer. They are listening properly now, engaged, and following where you are going rather than waiting for you to stumble.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Packt Deep Engineering! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2><strong>They design everything at Google scale</strong></h2><p>Senior engineers have worked on large systems and that experience is genuinely valuable, but it also creates a bias that hurts them in interviews: the instinct to design for the most demanding possible version of any problem, whether the problem actually requires it or not. Agarwal is direct about this. Not every system needs to scale to Google. If you are designing an internal tool that will only ever be used by your company&#8217;s engineers, you do not need multi-region deployment, and you do not even need cloud infrastructure. You could run it on a local area network and it would be perfectly adequate for the problem at hand. The engineer who reaches for global infrastructure for a problem that does not need it is demonstrating a failure of judgment, not a depth of knowledge.</p><p>Good system design is about matching the architecture to the requirements you gathered in those first two minutes, not about showcasing every pattern you have ever learned across a career. The interviewer is not evaluating whether you know how to design at Google scale. They are evaluating whether you understand when to use which level of complexity and why, and that distinction is entirely invisible if you default to maximum complexity regardless of the constraints in front of you.</p><div><hr></div><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;45f78df8-1ffd-4b85-a665-d571d1e41c5f&quot;,&quot;caption&quot;:&quot;From monolith-to-services signals to &#8220;performance per dollar&#8221; and practical resilience under real attacks&#8212;clear choices you can defend in production and interviews.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #36: Archit Agarwal on System Design Trade-offs &quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:140662997,&quot;name&quot;:&quot;Divya Anne Selvaraj&quot;,&quot;bio&quot;:&quot;Editor-in-Chief of Deep Engineering by Packt&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/309a6f07-27a6-40bf-ab99-d042556d816b_400x400.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:14853060,&quot;name&quot;:&quot;Archit Agarwal&quot;,&quot;bio&quot;:&quot;Principal Member of Technical Staff at Oracle, building ultra-low latency auth in Golang. Creator of The Weekly Golang Journal, turning system design theory into practical, high-performance code using language and SQL internals, cloud solutions.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1d3d8661-57d3-4b2a-8e42-f789caf14aa6_200x200.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-02-26T13:30:47.032Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/837caa73-ceef-4856-b9c0-49f815314471_1536x1024.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.substack.com/p/deep-engineering-36-archit-agarwal&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:188591840,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:3,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div><hr></div><h2>They go quiet when they are thinking</h2><p>Senior engineers are often comfortable sitting with a difficult problem for several minutes before speaking, and in a production context that is a perfectly reasonable way to work through something complex. In a system design interview it reads as disengagement, and the interviewer has no way to tell whether you are making progress or whether you are stuck. Agarwal uses a phrase that reframes what good communication looks like in this context: the interviewer needs to be able to follow your brain&#8217;s commit history. Every decision you make, every trade-off you consider and reject, every assumption you surface and then validate or invalidate, should be spoken out loud as you make it, not as a performance or a monologue but as a live narration of your actual reasoning as it happens.</p><p>This serves two distinct purposes. It gives the interviewer genuine insight into how you think rather than just what conclusion you eventually reached, which is what they are actually evaluating. And it forces you to be more precise about your own reasoning, because articulating a decision out loud surfaces the assumptions underneath it in a way that thinking silently does not. Agarwal&#8217;s observation is that engineers who think out loud often catch their own errors in real time and self-correct naturally, and that self-correction is not a weakness. It is exactly the kind of flexible, honest thinking the interviewer is looking for.</p><h2>They defend their design when constraints change</h2><p>Experienced engineers have ownership instincts built over years of shipping and defending decisions in production. When they have built something they defend it, and in most professional contexts that instinct is appropriate. In a system design interview it becomes a liability the moment the interviewer introduces a constraint change mid-session, which Agarwal says he genuinely enjoys doing precisely because it reveals something important about the candidate.</p><p>Changing constraints are the normal reality of production engineering. Requirements shift, scale changes, new compliance requirements appear, and the ability to absorb a change, restate it clearly to confirm alignment, identify which parts of the design need updating and which parts remain intact, and then restructure calmly is exactly the capability that distinguishes an engineer who can operate in a real production environment from one who can only design under controlled conditions. The engineers who struggle here are the ones who treat the curveball as an attack on their design and respond by defending the original rather than adapting to the new information. Agarwal&#8217;s point is unambiguous: the interviewer is not trying to invalidate your architecture. They are trying to see whether you can hold your design lightly enough to change it when the situation demands it, which is something you will be required to do repeatedly in any engineering role worth having.</p><h2>They use jargon to sound credible instead of clarity to be understood</h2><p>Senior engineers have large vocabularies built from years of working across complex systems. Distributed systems, eventual consistency, CQRS, saga pattern, two-phase commit. These are real concepts with real meanings and knowing them is genuinely useful. But using them in rapid succession without grounding them in the specific problem being discussed is a signal that the engineer is performing knowledge rather than applying it, and experienced interviewers recognise the difference immediately.</p><p>Agarwal&#8217;s standard for communication in a system design interview is demanding but correct: your explanation should be clear enough that even a junior engineer could follow the reasoning without needing to already know the answer. Not dumbed down, and not simplified to the point of inaccuracy, but clear enough that every choice is grounded in the specific requirements of the system being designed rather than in a general desire to demonstrate familiarity with advanced concepts. The engineers who stand out in Agarwal&#8217;s interviews are not the ones with the most impressive vocabulary. They are the ones who make him feel like he is sitting with another engineer genuinely working through a problem together, which is exactly what a system design interview is supposed to be.</p><div><hr></div><p>The full conversation with Archit Agarwal is now live on Deep Engineering. </p><div id="youtube2-rOLA2NpKPfM" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;rOLA2NpKPfM&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/rOLA2NpKPfM?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div><hr></div>]]></content:encoded></item><item><title><![CDATA[Rust Is Hard for the Engineers with the Most Experience]]></title><description><![CDATA[On the unlearning problem that trips up experienced engineers and the production payoff that makes it worth it]]></description><link>https://deepengineering.net/p/rust-is-hard-for-the-engineers-with-the-most-experience</link><guid isPermaLink="false">https://deepengineering.net/p/rust-is-hard-for-the-engineers-with-the-most-experience</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Mon, 18 May 2026 16:07:41 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f3cfa859-6cd7-4d3e-ab4b-31c7b0cd00c5_1440x660.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Rust, who would have thought, has ranked as the most loved programming language in the Stack Overflow developer survey for nine consecutive years. Honestly, I must admit this is an unusual kind of statistic because it measures not just adoption but retention. The engineers who use Rust want to keep using it, and that pattern has only deepened even as the language moved from systems programming curiosity to production infrastructure at companies including Amazon, Google, Meta, and Microsoft. Interestingly, the Linux kernel now carries Rust code without the experimental label it held for years. <a href="https://lists.debian.org/debian-devel/2025/11/msg00188.html">Debian&#8217;s APT package manager</a> is also introducing hard Rust dependencies this year.</p><p>The performance benchmarks and the memory safety arguments have been made, tested in production, and largely validated. But what the benchmarks do not explain is why so many experienced engineers find Rust genuinely difficult to work with, why the teams that adopt it often go through a period where velocity drops before it recovers, and what it actually takes to get good at it rather than just competent. These are the questions that sit behind the adoption numbers and they matter more than the numbers do for any engineer thinking seriously about where Rust fits in their work.</p><p><a href="https://www.linkedin.com/in/evan-williams-1512092">Evan Williams</a>, author of <em><a href="https://www.packtpub.com/en-ar/product/design-patterns-and-best-practices-in-rust-9781836209461">Design Patterns and Best Practices in Rust</a></em>, has been writing software for more than 40 years and came to Rust while building a hardware system that needed to be rock solid and run without access in a remote location. <a href="https://it.linkedin.com/in/francesco-ciulla-roma/en">Francesco Ciulla</a>, author of <em><a href="https://www.packtpub.com/en-us/product/the-rust-programming-handbook-9781836208860">The Rust Programming Handbook</a></em><a href="https://www.packtpub.com/en-us/product/the-rust-programming-handbook-9781836208860"> </a>and a Docker Captain who previously worked at the European Space Agency on the Copernicus project, started publishing Rust content in 2022 and 2023, earlier than most, and has since watched the language&#8217;s adoption from the inside. Their perspectives on Rust come from different parts of the stack and different kinds of work, but on the questions that actually trip engineers up, they are in close agreement.</p><blockquote><p><em>We interviewed Evan Williams and Francesco Ciulla separately for <a href="https://deepengineering.substack.com/s/newsletter-issues">Deep Engineering Newsletter</a> issues.</em></p></blockquote><h2>The engineer who struggles most is usually the most experienced one</h2><p>The reasonable assumption when a team introduces Rust is that the senior engineers will pick it up fastest. They have the most context, the most pattern recognition, and the most experience navigating unfamiliar codebases. In practice, the opposite tends to happen, and both Williams and Ciulla have seen it play out firsthand.</p><p>&#8220;The more experienced you are, the more years you have doing something in some other language, the more trouble you&#8217;re likely to have,&#8221; Williams says, &#8220;because you have patterns of thought that come from those languages that you don&#8217;t even realize are there.&#8221; The problem is not that experienced engineers consciously try to apply Java or C++ patterns to Rust. The problem is that those patterns are invisible to them, baked in over years of use until they no longer register as choices at all. The engineer is not making a decision when they reach for inheritance or shared mutable state. They are doing what has always worked, and Rust will not let them.</p><p>Ciulla put it more directly. &#8220;Even if you are a senior developer, even if you have twenty years of experience, if you want to try to learn Rust comparing it to other programming languages, you will fail, because it&#8217;s like learning something which is completely new.&#8221; He makes the case that this is not a reason to avoid Rust but a reason to go in with a specific kind of openness, one that experienced engineers often find harder to maintain than junior ones do precisely because they have more to unlearn. A developer learning Rust as their second or third language has no competing mental model to discard. A senior engineer with a decade long experience in Java has to dismantle instincts that have been reliable for years before they can build new ones, and that dismantling is the work that most people underestimate going in.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;ef4dcc9c-3432-4c97-99aa-2f125e7d6b37&quot;,&quot;caption&quot;:&quot;Francesco Ciulla has been building with Rust since 2022, has spoken about it internationally at conferences, and his perspective on Rust adoption is shaped less by enthusiasm for the language and more by a practitioner&#8217;s view of where it actually earns its place in a production system.&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #45: Francesco Ciulla on Building Production Systems in Rust Without the Expensive Rewrite &quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:11407185,&quot;name&quot;:&quot;Francesco Ciulla&quot;,&quot;bio&quot;:&quot;Developer Advocate at @dailydotdev\n&#183; Docker Captain &#128051;\n&#183; Public Speaker\n&#183; Building a 1 Million Community 22%&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b8c30606-10ba-4c87-a89b-af2f9dc27a01_400x400.jpeg&quot;,&quot;is_guest&quot;:true,&quot;bestseller_tier&quot;:null,&quot;primaryPublicationSubscribeUrl&quot;:&quot;https://francescociulla.substack.com/subscribe?&quot;,&quot;primaryPublicationUrl&quot;:&quot;https://francescociulla.substack.com&quot;,&quot;primaryPublicationName&quot;:&quot;Francesco's Newsletter&quot;,&quot;primaryPublicationId&quot;:1410908}],&quot;post_date&quot;:&quot;2026-04-30T16:32:00.000Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6e29b888-c64d-4029-96f3-08e45522d077_656x375.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.substack.com/p/issue-45-francesco-ciulla-building-production-systems-rust-without-rewrite&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:195995087,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:2,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>Trusting the compiler is not a beginner tip</h2><p>The first thing most engineers do when the borrow checker rejects their code is look for the minimum change that will make it compile. That is the right approach in almost every other language and the wrong one in Rust, and it is where a significant amount of early frustration comes from.</p><p>&#8220;The golden rule is to trust the compiler, especially at the beginning,&#8221; Ciulla says. What he means is not passive acceptance but active reading. The borrow checker is not producing noise. It is producing information about what the program&#8217;s structure requires, and engineers who learn to read it that way move through the learning curve faster than engineers who treat every error as an obstacle to clear. The difference is subtle at first and significant over time.</p><p>Williams argues that the borrow checker is doing something more useful than preventing bugs. &#8220;The borrow checker is your friend because it prevents you from making a messy design. It prevents you from making a broken design. It prevents you from writing whole classes of bugs that you will then spend many hours trying to find,&#8221; he explains. &#8220;I have found it to be an incredible partner in writing code that allows me to sleep at night.&#8221; The reason it works this way is that Rust&#8217;s ownership rules enforce a discipline that experienced engineers in other languages apply selectively and inconsistently because those languages do not require it. A value has one owner. References are either shared and immutable or exclusive and mutable, never both. The compiler will not proceed until the code is explicit about who owns what and when.</p><p>&#8220;The principles that the borrow checker forces you to adhere to in Rust are the exact principles that you should be using in every programming language,&#8221; Williams reasons. &#8220;But you don&#8217;t have to. So it&#8217;s very easy to not think about those things.&#8221; That observation reframes what the borrow checker is. It is not an imposed restriction. It is a discipline that good engineers apply in other languages by habit and judgment, made non-negotiable and automatic in Rust.</p><p>The discipline extends from individual functions to the shape of the whole system. A program that handles ownership correctly at the function level has to handle it correctly across modules, across threads, and across component boundaries, because the same rules apply everywhere. &#8220;You need to think about who controls what, how it is controlled, and you need to start from the very beginning thinking about the boundaries of your program and the system architecture, dividing things up into areas of responsibility,&#8221; Williams underscores. &#8220;Because unlike Python or Java, you can&#8217;t have links going all over the place. The borrow checker is never going to accept that.&#8221; The result is that well-written Rust systems tend toward a specific architectural shape: data flows in one direction, ownership chains move forward and do not loop back, and the behavior of the system is legible from its structure in a way that systems with shared mutable state often are not.</p><p>The most underutilized expression of what this makes possible is the typestate pattern. It uses the type system to encode the state of a value at compile time in a way that makes invalid state transitions not just errors but programs that cannot be compiled at all. Williams reflects on it with visible enthusiasm. &#8220;It&#8217;s a way of developing state machines and systems that have state that evolves where invalid state transitions aren&#8217;t just errors, they&#8217;re impossible to write. The compiler won&#8217;t compile them,&#8221; he says. &#8220;It represents a huge advance in the way that such systems are written because now instead of runtime errors, you have a state machine that is guaranteed to work because every transition either is a valid transition or it won&#8217;t even compile. That&#8217;s an amazing thing.&#8221; The pattern was not invented for Rust, but the language&#8217;s ownership system and type handling make it practical in a way that other languages do not, and for systems where invalid state transitions are genuinely dangerous rather than merely inconvenient, it is one of the most concrete expressions of what Rust makes possible.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;caa3ca39-6118-4ad9-9d55-53c3803aabaf&quot;,&quot;caption&quot;:&quot;Evan Williams shares engineering insights on the borrow checker as a design tool, the object-oriented trap, and why the engineers who struggle most with Rust are often the most experienced ones.&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #47: Evan Williams on Why Experienced Developers Have the Hardest Time Learning Rust&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-05-14T16:42:52.048Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/43a42d88-ec70-4d7f-8213-85796343b4f5_677x337.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.substack.com/p/deep-engineering-47-why-experienced-developers-hardest-time-learning-rust&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:197666671,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>What Rust actually gives you in production</h2><p>Ciulla&#8217;s case for Rust is grounded in things he measured when running a Rust web server on his own machine. It was consuming four megabytes at rest and five in production. &#8220;If you have a droplet with one gigabyte of RAM, you can have 200 plus services,&#8221; he notes, &#8220;of course in idle, but this proves that if you have a service that consumes a lot of RAM, it is worth thinking about.&#8221; For teams running infrastructure where memory costs money and density matters, the difference between a Rust service and an equivalent service in a garbage-collected language is not marginal.</p><p>The latency story is also specific. &#8220;By not having a garbage collector on the back end side, you basically have a flat latency,&#8221; Ciulla observes. &#8220;If a user makes an HTTP request when the garbage collector starts, it will experience a higher latency. Rust removes that problem entirely.&#8221; Go and Node.js both have garbage collectors that pause for collection cycles, and even short pauses measured in hundreds of milliseconds are enough to introduce latency spikes that affect users who are unlucky enough to hit the request at the wrong moment. Rust&#8217;s absence of a garbage collector means the latency profile is predictable rather than probabilistic, which matters significantly for services where consistency is as important as average throughput.</p><p>The deployment model is simpler than most engineers expect going in. A Rust project built with cargo produces a standalone binary for the target architecture, which packages cleanly into a container image. &#8220;If you build the executable when you build the Docker image, you have something which is just deployable everywhere,&#8221; Ciulla says. &#8220;A Linux executable running in a Docker container. That&#8217;s the dream.&#8221; The operational benefit is smaller images, faster startup, and a runtime with almost no overhead beyond the binary itself.</p><p>Williams approaches the production question from the correctness angle rather than the performance angle. The systems where Rust earns its place most clearly are the ones where failure has a real cost. &#8220;Systems that are mission critical in some way or other are really key Rust use cases,&#8221; he says. &#8220;All of these features combine into a whole that make Rust a really powerful language for doing things that have to work. Things where failure is monetarily or in human cost even a terrible problem.&#8221; The memory safety and the ownership model and the compile-time guarantees are not separate features. They are different expressions of the same underlying commitment: the program either demonstrates its correctness to the compiler or it does not compile.</p><p>Williams also reflects on something unexpected he discovered while writing the early chapters of his book, the ones covering what not to do in Rust. He went back and deliberately tried to write bad code, the kind of code that would illustrate the mistakes he was cautioning against, and found it harder than he expected. &#8220;When I went back and tried to write bad code in Rust, it was much harder than writing the good code,&#8221; he recalls. &#8220;That&#8217;s an interesting perspective that just didn&#8217;t even occur to me.&#8221; The language&#8217;s constraints push code toward a particular shape so consistently that departing from it requires actively working against the grain of the language rather than simply making a poor choice.</p><p>&#8220;The biggest benefit in Rust is about the lack of the debugging depth. You spend more time thinking up front, but you spend almost zero time chasing segfaults or memory leaks in production,&#8221; Ciulla remarks. &#8220;And we always underestimate this part. We always talk about the efficiency of the code, but if you need less time to debug your code, you&#8217;re basically writing more logic at the end of the day.&#8221; The upfront investment in getting the types and the ownership right is real, but the downstream debugging cost it removes is larger and does not diminish as the team becomes more experienced. It is simply gone.</p><h2>Where to start and where Rust is the wrong tool</h2><p>On the practical question of how to bring Rust into an existing codebase, both Williams and Ciulla give advice that converges almost exactly despite coming from different engineering contexts. Neither recommends starting with a rewrite.</p><p>&#8220;The best way to introduce Rust in a big project is to find that hard part that&#8217;s the bottleneck and try to write one single service in Rust,&#8221; Ciulla says. &#8220;And then you will see, probably slowly, Rust might take over your code base, but I mean this in a good sense.&#8221; Williams makes the same point with a specific warning about the temptation to go faster. &#8220;What you don&#8217;t want to do is jump into saying, we&#8217;re just going to rewrite our project in Rust now. Pick a small piece, focus on that, gain confidence and mastery of the language, and then use that to build upon it and start bringing in more things,&#8221; he says. Starting with a bounded, non-critical component gives the team room to move through the learning curve without the pressure of a production incident concentrating everyone&#8217;s attention on the wrong things.</p><p>Ciulla adds something worth noting about the AI-assisted workflow that is becoming standard for many engineers. &#8220;In this AI era, everyone is rushing stuff with AI, but you still need the validation,&#8221; he says. &#8220;Okay, AI wrote this Rust service, but now who decides if this is okay to put in production? Of course, you need the validation of an expert.&#8221; The Rust compiler catches a large class of errors automatically, but the errors that survive it, logic errors rather than memory errors, still require someone who understands the language well enough to see what the code is actually doing. Having at least one engineer on the team who knows Rust well enough to review AI-generated code is not optional.</p><p>Both are also direct about when Rust is not the right choice. Ciulla points to tight deadlines and fast prototyping as the clearest case against it. &#8220;If you need fast prototyping, you are familiar already with Java, JavaScript, why don&#8217;t you use it?&#8221; he says. &#8220;When the deadline is so close, probably it&#8217;s not the best way to try something new because something would go wrong, especially if you&#8217;re not an expert.&#8221; </p><p>Williams points to user interfaces as an area where the ecosystem is still catching up and the tooling gaps are large enough to make other languages more practical. &#8220;Doing a website in Rust is still kind of a feat,&#8221; he notes, &#8220;and it&#8217;s an awful lot easier to use the tools that everybody else is using to accomplish that goal.&#8221; The Python data science ecosystem is another area Ciulla names directly: the libraries are simply better established there, and using Rust for data science work means building against a thinner set of available tools than Python provides.</p><h2>Where Is the Ecosystem Headed</h2><p>Ciulla expects Rust to grow most significantly in the near term, and his prediction lands in a direction that surprises most of the Rust community. &#8220;I think the next big wave might be in web development,&#8221; he says, adding that he is aware this is an unpopular position in a community that still thinks of Rust primarily as a systems language. His reasoning is grounded in what he has been seeing directly: companies with hundreds of developers reaching out to tell him they are moving services to Rust for their web backends. &#8220;I get this news because I&#8217;m well known for talking about Rust and being quite vocal about it,&#8221; he observes. &#8220;I&#8217;m not talking about a person just doing this on a random Saturday night. I&#8217;m talking about companies that have hundreds of developers.&#8221; He points to Axum as the framework that has matured to the point where he would now use it in a production SaaS product, which he says was not true two years ago. Embedded systems, in his view, have already crossed the threshold where Rust&#8217;s place is settled.</p><p>Williams takes a longer view on how the ecosystem will evolve. &#8220;The ecosystem is going to get richer and people are going to be branching out in the set of use cases, hitting areas that right now Rust has relatively weak support for,&#8221; he says. &#8220;As larger and larger projects are built, there is going to be more refinement of the language itself, but more importantly, more refinement of the use of the language.&#8221; The patterns that make Rust work well at scale are still being discovered and codified. The language is stable, but the understanding of how to use it well is still developing, and that development is happening inside the teams building the largest Rust codebases.</p><p>What both conversations point toward is a language whose difficulty and whose value come from the same source. Rust is hard to learn for experienced engineers because it refuses to accommodate the habits that made them experienced. It is valuable in production because that same refusal, enforced by the compiler on every build, produces code whose behavior is predictable, whose data flows are legible, and whose failure modes are constrained to things the language cannot check rather than things the engineer forgot to check. The engineers who get the most out of it are the ones who stop trying to carry their existing instincts across and start letting the compiler teach them what the program actually needs.</p><div><hr></div><h4><em><strong>In case you missed</strong></em></h4><p><em>Here&#8217;s the full interview video featuring Evan Williams.</em></p><div id="youtube2--ElpmT7DCX4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;-ElpmT7DCX4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/-ElpmT7DCX4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div>]]></content:encoded></item><item><title><![CDATA[Clean Code Is a Trap, Decompose Instead for Physics and Performance]]></title><description><![CDATA[On cache locality, cognitive load, and the engineering habits that make fast development sustainable with Sam Morley and S&#225;ndor Darg&#243;]]></description><link>https://deepengineering.net/p/clean-code-trap-decompose-for-performance-physics</link><guid isPermaLink="false">https://deepengineering.net/p/clean-code-trap-decompose-for-performance-physics</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Thu, 23 Apr 2026 15:15:59 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e4a62837-3922-40ce-8859-73c783c89af9_822x371.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Engineering teams obsess over clean code because they want software to look organized and logical in the text editor. Principles like SOLID get followed strictly, and hours get spent debating folder structures, because it feels like the disciplined way to build software. But this desire for logical cleanliness often leads into a trap where teams build systems that are beautiful to read but terrible to run.</p><p>The most maintainable codebases are not the ones that adhere to a style guide. They are the ones that respect the physical and cognitive reality of the environment they live in.</p><p>We have spoken (interviewed separately) to two notable engineers who think about this from very different directions. <a href="https://www.linkedin.com/in/morleys90/">Sam Morley</a>, a mathematician and C++ researcher at the <a href="https://www.maths.ox.ac.uk/">University of Oxford</a>, approaches software from the ground up, where the cost of every abstraction shows up immediately in performance metrics. <a href="https://www.linkedin.com/in/sandor-dargo/">S&#225;ndor Darg&#243;</a>, a senior software engineer at <a href="https://engineering.atspotify.com/">Spotify</a> who works on large-scale C++ systems, approaches it from the maintainability side, where the cost of every abstraction shows up in the engineers who have to live with the code months or years later. </p><p>Both conversations happened at different points, on different topics, but arrived at the same conclusion that logical cleanliness is not the goal. But understanding what the machine and the team actually need is.</p><h2>Your CPU does not care how tidy your objects look</h2><p><strong>Sam Morley&#8217;s</strong> starting point is hardware, and his argument is that the way most engineers are taught to structure code works directly against the way processors are designed to access memory.</p><p>The instinct is to group data into objects because it models the real world. A Player class holds position, health, velocity, and inventory in one contiguous block, because those things belong together conceptually. But the CPU fetches data in contiguous blocks called cache lines, and if the object structure fills that cache line with data the processor does not need for the current operation, the application pays for it in cycles. The cost is invisible in code review but shows up immediately in a profiler under load.</p><p>Morley points to the Structure of Arrays pattern, common in game development, as the counterintuitive solution. Instead of an array of Player objects, you create separate arrays for positions, health values, and velocities. This looks messy to a developer trained in object-oriented design. It violates the instinct to keep related data together, and it produces code that does not map neatly onto the real-world entities it represents. But it allows the CPU to process data significantly faster because every byte in a fetched cache line is a byte the processor actually needs. Cache locality, not conceptual tidiness, determines throughput under real conditions.</p><p>Morley&#8217;s recommendation is direct: be willing to break clean object models when the hardware requires it. The machine is not going to adapt to the abstraction. The abstraction has to adapt to the machine. And this is not a concern limited to embedded engineers or game studios. It is a reality for any C++ system under sustained load, and the gap between what looks clean and what runs efficiently widens as the scale increases. Teams that do not understand this distinction tend to optimize the wrong things when performance problems eventually surface.</p><h2>Clever code is a debt that Future You will have to repay</h2><p>Morley&#8217;s second argument shifts from CPU cost to cognitive cost, and it is the more insidious of the two because it compounds slowly and invisibly until a maintenance crisis makes it visible all at once.</p><p>His framing here is precise. Future You is a completely different person who has lost all the context that made the current design feel obvious at the time it was written. The engineer writing the code holds the whole system in their head. The engineer returning to it six months later does not. And the engineer reading it for the first time never did. Every clever abstraction that felt natural in the moment of writing becomes a reconstruction problem for every reader who comes after.</p><p>Template-heavy code and metaprogramming are the most common form of what Morley calls Wizardry. The name is apt because Wizardry works by concealment. The complexity does not disappear when abstracted away. It becomes invisible until someone needs to debug or extend the system, at which point the engineer is starting from a significant disadvantage with no clear view of how data actually moves through the code. What Morley advocates instead is Process Awareness: code that exposes the data flow clearly rather than hiding it behind layers of indirection. Not short code or smart code. Code whose execution model is obvious to the next engineer who reads it, regardless of whether that engineer was involved in writing it.</p><p>The practical implication is to treat <strong>Future You</strong> as a first-class stakeholder in every design decision. And so, the documentation that explains what the code does is far less valuable than documentation that explains why it is structured the way it is, because the what is usually legible from the code itself. The why rarely is.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;b51368e0-7c19-47ba-bdef-b91f8ea7f718&quot;,&quot;caption&quot;:&quot;Template metaprogramming, cache-aware design, concurrency models, and why learning Rust might actually make you a better C++ programmer&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #41: Scaling C++ the Right Way with Sam Morley&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:480630041,&quot;name&quot;:&quot;Deepayan Bhattacharjee&quot;,&quot;bio&quot;:&quot;Content engineer at Packt&quot;,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!jegY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F385931a8-2eb0-4f8a-bd57-14713ff3988d_144x144.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:440051761,&quot;name&quot;:&quot;Sam Morley&quot;,&quot;bio&quot;:&quot;Research software engineer and mathematician on the DataSig project at the University of Oxford.&quot;,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!hfZA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b7dcfcf-a878-45d0-99e4-a8f2045dee3e_144x144.png&quot;,&quot;is_guest&quot;:true,&quot;bestseller_tier&quot;:null,&quot;primaryPublicationSubscribeUrl&quot;:&quot;https://sammorley.substack.com/subscribe?&quot;,&quot;primaryPublicationUrl&quot;:&quot;https://sammorley.substack.com&quot;,&quot;primaryPublicationName&quot;:&quot;Sam Morley&quot;,&quot;primaryPublicationId&quot;:7726502}],&quot;post_date&quot;:&quot;2026-04-02T15:16:15.996Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4017b543-27c2-4082-9f3b-1bd7abbbdd3a_1432x840.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.substack.com/p/deep-engineering-41-scaling-c-the&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:192949941,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:4,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>When cognitive load becomes your biggest bug</h2><p><strong>S&#225;ndor Darg&#243;</strong> approaches the same problem from a different direction but arrives at the same place. His work at Spotify on large-scale C++ systems has given him a practitioner&#8217;s view of what happens to codebases over time when cognitive cost is not treated as a first-class engineering concern from the start.</p><p>For Darg&#243;, the thread connecting clean code, binary size, undefined behavior, and C++ language evolution is a single idea: reducing complexity in real-world systems. Not as an aesthetic preference, but as a measurable engineering outcome with consequences for how fast teams can move, how safely they can refactor, and how much institutional knowledge survives when people leave. &#8220;If you think about clean code, it clearly reduces the cognitive load,&#8221; Darg&#243; said during a <a href="https://deepengineering.substack.com/p/clean-c-code-and-the-hidden-cost">recent Deep Engineering interview.</a> &#8220;If you think about binary size, it might reduce operational cost. New standards like C++23 and C++26 reduce boilerplate and enable safer, more readable abstractions. All of these topics make large C++ systems more maintainable and more evolvable.&#8221;</p><div id="youtube2-vsdeOS8snN0" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;vsdeOS8snN0&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/vsdeOS8snN0?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>The connection between these concerns is not accidental. Binary size reduction often leads teams toward simpler code as a side effect, because the practices that reduce binary size, avoiding unnecessary template instantiation, being deliberate about what gets inlined, minimizing heavy type erasure, also tend to reduce the number of moving parts an engineer has to hold in mind. The discipline required to keep a binary small and the discipline required to keep a codebase readable are more closely related than most teams realize until they have worked on both problems at the same time.</p><p>Darg&#243;&#8217;s warning is about the human cost of poor abstraction choices, and in his experience, teams routinely optimize the wrong things because they measure the wrong variables. The heap allocation is visible. The cost of a network request made inside a loop is harder to see until a profiler makes it undeniable. Darg&#243; during our interview cited Amdahl&#8217;s Law to make the point concrete: the overall performance improvement gained by optimizing a single part of a system is limited by the fraction of time that part is actually used. The engineers spending time on heap allocations while making network requests in a loop are not being careless. They are solving the problem they can see. The discipline is in learning to find the problem that actually matters, which requires measurement rather than intuition. &#8220;If your code takes a long time to execute due to network latency, then relatively speaking, the heap allocation is not so slow anymore,&#8221; Darg&#243; said. &#8220;Don&#8217;t worry about things that don&#8217;t really matter in a given environment.&#8221;</p><h2>Write it readable first, then measure, then and only then optimize</h2><p>Darg&#243;&#8217;s practical framework for navigating these trade-offs is structured around a clear hierarchy of defaults, and the first default is unambiguous: readable code comes first.</p><p>His reasoning is grounded in a simple observation that engineering culture tends to underweight. Engineers read code far more often than they write it. Every decision that makes code harder to read imposes a recurring cost on every future reader, and that cost accumulates over the lifetime of the codebase. Defaulting to readability is not a concession to comfort. It is an engineering position with compounding returns, because code that is easy to read is code that is easy to reason about, and code that is easy to reason about is code that is safer to change.</p><p>The second principle follows directly: if optimization is necessary, measure before touching anything. The trap is optimizing before a measurement has confirmed that the thing being optimized is the actual problem. This wastes time, introduces unnecessary complexity, and often leaves the real bottleneck untouched. Measure first, identify the hot path, and only then begin the optimization work. Once the hot path is identified, keep it isolated and document the reasoning behind every trade-off made there. Not documentation that explains what the code does, but documentation that explains why it is structured the way it is, so the next engineer understands what they would be giving up if they cleaned it up.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://deepengineering.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Packt Deep Engineering! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Darg&#243; has been on the receiving end of the alternative. He came into a codebase, saw code that looked wrong, began cleaning it up, and realized too late that the seemingly redundant choice was affecting binary size in a way that mattered for the system. Pull requests had already merged before the context became clear. &#8220;Make trade-offs conscious,&#8221; Darg&#243; said. &#8220;Make them explicit in code reviews, but also in the code itself. If you sacrifice the clarity you aim for, document why. Because otherwise someone later will come in and make it cleaner, unaware of why certain choices were made.&#8221;</p><p>And this principle has become more critical in the age of agent-assisted development. If engineers can miss the intent behind an undocumented trade-off, then AI agent working on the same codebase will miss it with far greater confidence. Agents read what is in the code. They do not have access to the Slack conversation where the binary size constraint was first discussed, or the code review thread that resolved and got deleted. The context has to be in the code, because that is the only place every future reader, human or agent, will reliably look.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;b266a406-3df1-466a-984c-0c85e1ba4344&quot;,&quot;caption&quot;:&quot;Building an AI-Powered Internal Developer Platform from Scratch&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Deep Engineering #44: S&#225;ndor Darg&#243; on C++26, Adoption Traps, Compiler Gap, and Maintainability&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:427210082,&quot;name&quot;:&quot;Saqib Jan&quot;,&quot;bio&quot;:&quot;/localhost&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/997a788a-cd78-4f84-9b3b-c72ab6dc0153_1008x1008.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-04-23T16:31:24.008Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/efb9ba36-137d-4762-93d8-c394bc1fe4da_681x277.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://deepengineering.substack.com/p/issue44-cpp-26-adoption-traps-compiler-gaps-maintainability&quot;,&quot;section_name&quot;:&quot;Newsletter Issues&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:195242079,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:7,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1729053,&quot;publication_name&quot;:&quot;Packt Deep Engineering&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!H5BJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F736bc1ee-d689-497e-83a8-7d9bf9022eb9_600x600.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>The invisible tax that is getting harder to ignore</h2><p>Morley and Darg&#243; are describing the same underlying problem from different directions. Every time an engineer has to reconstruct context that was lost, the system has failed them. Morley calls it the Future You constraint. Darg&#243; calls it cognitive load. The mechanism is identical in both cases, and the cost is real even when it does not appear in any metric the team currently tracks.</p><p>This cost has become harder to ignore in the last year or two, and not only because systems have grown more complex. Darg&#243; observed during the same session that the shift to AI-assisted development has made context switching materially worse for most engineers, and the profession has not yet fully reckoned with what that means for how software gets built. Engineers are managing multiple agent sessions simultaneously, jumping between prompts and code reviews, moving from one incomplete task to another before any of them reach resolution. The flow state that reliable engineering has always depended on, the gradual accumulation of a mental model, the ability to hold a system&#8217;s behavior in mind long enough to reason about it clearly, gets interrupted more frequently and at shorter intervals than at any point in most engineers&#8217; careers.</p><p>&#8220;We became, often, just prompters,&#8221; Darg&#243; said. &#8220;Many of us complained even before that we are living in a world of constant context switching. But it just became even worse. You keep jumping from one window to another, from one meeting to another, because others are also moving faster. At least they think they move faster.&#8221;</p><p>The irony embedded in that observation is significant. The tools promising to accelerate delivery are simultaneously increasing the interruption rate that undermines the deep work required to produce reliable software. Speed and depth are being traded against each other, and the trade is often invisible until the consequences show up in the codebase months later.</p><p>Darg&#243; in our live interview also referenced a research finding that makes the dynamic concrete. Engineers who adopt AI-assisted workflows tend to ship more code early on, because the friction of writing has dropped. But code quality drops alongside it, and the initial speed advantage disappears within a few months as technical debt accumulates faster than it can be serviced. &#8220;In the beginning you ship more code, because it became so much easier. But you don&#8217;t just ship more code. You ship worse code. And that gain in speed is vanishing after a few months because you start accumulating technical debt at the same time. What first seemed faster becomes not faster, but the debt stays,&#8221; Darg&#243; said.</p><p>The answer is not to reject the tools or return to slower workflows. It is to be deliberate about what the tools are being used for and what gets left behind when they are used. Code that was generated quickly but carries no trace of why it is structured the way it is will cost someone considerably when the context is gone. The practices Morley and Darg&#243; both advocate, keeping the hot path isolated, documenting the reasoning behind trade-offs, defaulting to the readable option unless a measurement says otherwise, are not conservative instincts. They are the engineering habits that make fast development sustainable over time rather than just in the short sprint.</p><h2><strong>And so, what this actually adds up to</strong></h2><p>Morley and Darg&#243; are pointing toward the same conclusion from different vantage points: engineering quality cannot be measured by how organized the code looks in the editor.</p><p>Morley&#8217;s measure is hardware efficiency. Does the code respect the physical reality of how the processor accesses memory, and does it make the execution model visible to the next reader, or does it hide it behind abstractions that feel clever now but become maintenance burdens later? Darg&#243;&#8217;s measure is team sustainability. Does the code reduce the cognitive load of the people who maintain it over time, and does it make trade-offs explicit so future engineers and future agents can understand what they would be changing if they touched it?</p><p>Clean code is not a trap because readability is wrong. It is a trap because readability without an understanding of what matters in the specific environment produces systems optimized for the wrong audience. The abstractions that feel clean in the editor are often the ones costing the most in production. And the ones that look strange in a code review are often the ones that matter most to the system&#8217;s actual behavior.</p><p>Not whether it looks clean. But whether it helps the machine run correctly, and whether it helps the next engineer understand why it runs that way. Those two questions do not always have the same answer, but they are always worth asking together, and always worth asking before the code is written rather than after the pull request is merged.</p><p></p>]]></content:encoded></item><item><title><![CDATA[Agentic AI Is Redefining Edge Infrastructure]]></title><description><![CDATA[Agents can't wait. Neither can your infrastructure.]]></description><link>https://deepengineering.net/p/agentic-ai-is-redefining-edge-infrastructure</link><guid isPermaLink="false">https://deepengineering.net/p/agentic-ai-is-redefining-edge-infrastructure</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Wed, 25 Mar 2026 18:13:57 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/280cd0a1-f0d5-42aa-8e78-90ac44439a30_1200x628.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!01dX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!01dX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 424w, https://substackcdn.com/image/fetch/$s_!01dX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 848w, https://substackcdn.com/image/fetch/$s_!01dX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 1272w, https://substackcdn.com/image/fetch/$s_!01dX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!01dX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png" width="1200" height="628" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:628,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:175042,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://deepengineering.substack.com/i/192122268?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!01dX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 424w, https://substackcdn.com/image/fetch/$s_!01dX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 848w, https://substackcdn.com/image/fetch/$s_!01dX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 1272w, https://substackcdn.com/image/fetch/$s_!01dX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a72e38c-5ace-48be-8131-562b93c393dd_1200x628.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Artificial intelligence is entering a new phase with agentic AI, where autonomous systems perceive, decide, act, and learn without constant human oversight, operating independently across distributed environments while collaborating with other agents in real time.</p><p>This shift from centralized AI models to distributed, autonomous agents requires a fundamental rethinking of WAN infrastructure architecture. Previous AI patterns such as centralized training clusters, cloud-based inference, and hub-and-spoke data flows are inadequate for agentic systems that must operate at the edge with speed, autonomy, and resilience.</p><p>And in these environments, the WAN is no longer just a means of connecting branch sites to core data centers. It becomes the essential fabric enabling edge agents to synchronize data, share insights, and coordinate actions, making WAN performance, availability, and adaptability critical to agentic AI effectiveness.</p><h2>Distributed intelligence is edge-centric</h2><p><a href="https://www.linkedin.com/in/leejpeterson">Lee Peterson</a>, VP of Secure WAN Product Management at <a href="https://www.cisco.com/">Cisco</a>, explains where the pressure lands first. Edge environments routinely face unpredictable connectivity, and agents operating in those conditions cannot wait for centralized systems to respond.</p><p>Peterson points to concrete scenarios where this plays out, from autonomous vehicle navigation systems to intelligent manufacturing floors to retail environments where AI agents manage inventory, pricing, and customer experience simultaneously. In each of these cases, he reasons, the decisions that matter most are the ones that have to be made in milliseconds, based on local conditions, often where connectivity to centralized systems is intermittent or constrained.</p><p>But the connectivity assumption is where many organizations get it wrong. Peterson recommends designing for intermittent or constrained WAN conditions rather than treating reliable connectivity as a given, and ensuring real-time path selection for critical systems such as point-of-sale, inventory sync, and IoT devices so that agents can perform automatic remediation during WAN degradation without waiting on human intervention.</p><p>Unlike traditional AI models operating on data in controlled environments, he notes, agentic systems exist in the physical world where latency is measured in milliseconds and decisions have immediate consequences. Sending data hundreds of miles to a cloud data center for processing, Peterson argues, is structurally incompatible with the real-time autonomy these systems require, because the agent must process information, evaluate options, and act locally, right where the action is happening.</p><p>And the scale of coordination compounds this further. A smart city deployment might involve thousands of agents managing traffic flow, energy distribution, and public safety simultaneously, and Peterson underscores that these agents need to share insights and coordinate actions even when network connectivity degrades.</p><p>Organizations that continue to architect around centralized control will find their agentic deployments constrained at precisely the moments that matter most, because this distributed intelligence model is inherently edge-centric and the infrastructure needs to reflect that from the start.</p><h2>Compute at the edge: the foundation of agent autonomy</h2><p>Agentic AI requires compute resources co-located with data sources and decision points, which means deploying high-performance processing across thousands of distributed locations including retail, manufacturing, healthcare, and transportation.</p><p>The workload requirements are diverse and demanding, covering agents performing rapid inference on streaming data, conducting local model fine-tuning based on environmental feedback, and coordinating with peer agents across locations in real time. In retail, Peterson notes, this might translate to supporting smart shelves, computer-vision inventory systems, digital signage, loss-prevention analytics, and customer-flow optimization directly at each store location, which is a significant compute footprint by any measure.</p><p>But powerful edge compute alone cannot deliver the full potential of agentic AI, and Peterson is direct about why. Without equally sophisticated networking, autonomous agents remain isolated, unable to coordinate with peers, synchronize insights, or maintain collective intelligence across distributed environments. The two investments have to be planned together, not sequenced, because the value of edge compute depends almost entirely on the quality of the network that connects it.</p><h2>Networking at the edge: the nervous system of distributed intelligence</h2><p>Just as compute provides the processing foundation for autonomous decisions, networking forms the connective tissue enabling multi-agent coordination. Peterson is specific about what agentic AI requires from it. Low-latency communication between distributed agents, efficient data synchronization, security across untrusted environments, and effective network partitioning are not aspirational requirements but operational ones, and the gap between meeting them and not meeting them is the gap between a functioning agentic system and an isolated one.</p><p>Consider a manufacturing environment where dozens of AI agents coordinate production, where vision systems inspect components, robots adjust operations in real time, and predictive maintenance agents analyze telemetry from across the floor. Peterson uses this kind of environment to ground the networking argument, because these agents must communicate with millisecond latency and maintain coordinated operation even if connectivity to central systems is temporarily lost. His architectural recommendation is specific in that high-performance networking should be integrated directly into edge compute infrastructure to enable agent-to-agent communication with low latency and high bandwidth, rather than routing every interaction through distant aggregation points, because that approach where networking and compute are designed together is what makes real-time coordination possible.</p><p>On security, Peterson is equally precise and equally unambiguous. These systems require cryptographic identity for every agent, encrypted communication, hardware-based roots of trust, and zero-trust architectures designed into both layers from the ground up, ensuring the integrity of autonomous decisions affecting physical systems and human safety in critical infrastructures such as healthcare and transportation. Not as hardening added after deployment, but as a design constraint from day one.</p><h2>The convergence of compute and networking at the edge</h2><p>Peterson frames this moment as an inflection point for enterprise infrastructure strategy, and the practical implication is straightforward even if the work is not. Organizations cannot simply extend cloud architectures to edge locations and expect agentic systems to thrive, because the autonomous, distributed, real-time nature of these systems demands infrastructure where compute and networking are designed together to support local intelligence, agent coordination, and secure operation across thousands of diverse locations.</p><p>And there is a visibility dimension that Peterson adds, one that often gets missed in these conversations. As organizations deploy distributed AI agents across vast, heterogeneous environments, continuous visibility into WAN performance, network health, and application performance at each edge location becomes indispensable, because without it, blind spots undermine the autonomy and resilience that agentic AI requires and teams lose the ability to detect issues proactively, optimize operations, and assure reliable service delivery before degradation affects outcomes.</p><p>Of the choices organizations face right now, Peterson is clear about which ones carry the most weight. Infrastructure decisions made today will determine whether organizations lead this transformation or spend years retrofitting, and the convergence of compute and networking at the edge, he concludes, is the essential foundation upon which the next generation of autonomous, intelligent systems will be built.</p>]]></content:encoded></item><item><title><![CDATA[Benchmarks Are Making AI Coding Look Safer Than It Is ]]></title><description><![CDATA[Passing tests is not proof the code is safe or maintainable.]]></description><link>https://deepengineering.net/p/benchmarks-are-making-ai-coding-look</link><guid isPermaLink="false">https://deepengineering.net/p/benchmarks-are-making-ai-coding-look</guid><dc:creator><![CDATA[Saqib Jan]]></dc:creator><pubDate>Wed, 04 Feb 2026 18:02:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!uBaj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uBaj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uBaj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 424w, https://substackcdn.com/image/fetch/$s_!uBaj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 848w, https://substackcdn.com/image/fetch/$s_!uBaj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 1272w, https://substackcdn.com/image/fetch/$s_!uBaj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uBaj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png" width="1200" height="628" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:628,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:280880,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://deepengineering.substack.com/i/186883869?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uBaj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 424w, https://substackcdn.com/image/fetch/$s_!uBaj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 848w, https://substackcdn.com/image/fetch/$s_!uBaj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 1272w, https://substackcdn.com/image/fetch/$s_!uBaj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2556e4bd-217a-4b38-a98d-9b5dd7300e82_1200x628.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Most technical leaders are optimizing for speed. AI agents now generate code fast enough to reshape how teams ship software. So teams contending with shorter deadlines and shrinking budgets are integrating them into delivery pipelines to increase velocity.</p><p>If you are an engineering leader, you have likely seen the SWE-bench leaderboard. It is the current industry standard for ranking AI coding agents. It scores agents based on whether they can produce a patch that passes a test suite. If it does, the agent gets a gold star.</p><p>But there is a deeper and often overlooked problem that creates a blind spot for enterprise teams.</p><p>Most teams treat these scores like a proxy for real engineering readiness. Speed is not the same as quality, but true velocity is speed plus quality. And passing tests is not the same as writing safe, maintainable code. This then shows up later as security debt, brittle systems, and review fatigue.</p><h2>The Pass/Fail Trap</h2><p>Benchmarks like SWE-bench are designed to test code generation rather than code quality. They ask if the agent can generate a solution that satisfies the immediate requirement.</p><p>They do not ask if the code is maintainable or if it introduces a hidden security vulnerability. They also ignore whether the new code breaks the architectural pattern of the rest of the application.</p><p><a href="https://www.linkedin.com/in/itamarf">Itamar Friedman</a> is the CEO and co-founder of <a href="https://www.qodo.ai">Qodo</a>, the AI Code review platform, says this creates a false sense of security for technical leaders.</p><p>&#8220;SWE-bench is a benchmark that is meant mostly to check code generation capabilities. You can get a really good grade with quite shitty code. It will pass because it implements the requirements and passes the test. But maybe the code is not maintainable. Maybe it includes a security issue.&#8221;</p><h2>The Illusion of Speed</h2><p>In the past, humans wrote code slowly and other humans reviewed it just as slowly. Now that AI agents are writing code at lightning speed, developers are opening two to five times more Pull Requests than they did a year ago.</p><p>This creates a phenomenon called quality rot. Even if AI generates code that is as good as a human&#8217;s, generating ten times more of it means you also generate ten times more bugs.</p><p>Friedman argues that relying on a &#8220;generation benchmark&#8221; to solve this is dangerous. He compares software development to accounting to show why the roles must be separate.</p><p>&#8220;You have bookkeeping and you have auditing. Ideally, you have two different people that are experts. One is doing the bookkeeping and the other is doing the auditing to verify the quality. Using the same agent to do both tasks is counterproductive.&#8221;</p><h2>The Hidden Risk of Review Fatigue</h2><p>When AI agents generate thousands of lines of code in minutes, human reviewers naturally get overwhelmed. They start skimming the code and often trust the AI simply because the test suite passed.</p><p>This is exactly where bugs slip in. A generalist model like GPT-5 might fix a logic bug but accidentally hardcode a credential or use a deprecated library.</p><p>If you rely on the same model to review the code it just wrote, you are essentially asking the fox to guard the hen house. A generalist model might be creative enough to solve the problem, but it lacks the rigid structure needed to audit safety.</p><h3>What You Should Do</h3><p>You need to stop obsessing over which model has the highest SWE-bench score and instead build a system of checks for your AI.</p><p>First, do not trust the generalist model to police itself. You should use specialized agents where one agent writes the code and a completely different agent reviews it against a strict policy.</p><p>Second, you should measure the number of valid bugs your AI catches in PRs rather than just how many PRs it opens.</p><p>Finally, you need to treat your AI pipeline like a government rather than a single employee. Friedman emphasizes that a single agent is never enough to ensure enterprise trust.</p><p>&#8220;You need a system. A system like a country. There are policies, rules, and a police.&#8221;</p><p>The future is not about faster coding but about smarter reviewing.</p>]]></content:encoded></item></channel></rss>