<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>JeffOps — Jeff. Tech. Dev. Ops.</title>
  <link>https://jeffops.com/</link>
  <atom:link href="https://jeffops.com/rss.xml" rel="self" type="application/rss+xml"/>
  <description>Practical AI-tomation, platform engineering and enterprise ops from Jeff Wouters. Written from inside a 25,000-user environment, not from a vendor deck.</description>
  <language>en</language>
  <lastBuildDate>Tue, 18 Aug 2026 09:00:00 +0000</lastBuildDate>
  <item>
    <title>The human in the loop: safeguard, or blame sponge?</title>
    <link>https://jeffops.com/newsletter/human-in-the-loop/</link>
    <guid isPermaLink="true">https://jeffops.com/newsletter/human-in-the-loop/</guid>
    <pubDate>Tue, 18 Aug 2026 09:00:00 +0000</pubDate>
    <description>A human in the loop is worth exactly as much as their ability to catch the error. The four ways that ability collapses to nothing, and the four questions that tell you whether you built a safeguard or a moral crumple zone.</description>
    <category>Newsletter</category>
    <content:encoded><![CDATA[<p>Twice in this series I have told you to keep a human in the loop. Once in the piece on plotting AI to your business, and again last edition, where I handed you four rules for naming a human owner and called the whole thing accountability.</p>
<p>Both times I gave it to you like it was a solution. It wasn’t. Here is the line this whole edition hangs on: <strong>a human in the loop is worth exactly as much as their ability to catch the error, and not one cent more.</strong></p>
<p>So, this one is for anyone with an AI workflow running right now that has a name written next to it. And for the person whose name that is, who may not know what they are holding.</p>
<p>I ended the last edition on a cliffhanger, and this is me paying it off. Right after telling you to make a named human accountable for every AI outcome, I admitted the catch: that only works if the human can actually catch the error. Sometimes the person in the loop is there to absorb the blame rather than prevent the failure. Safeguard, or blame sponge? I said we would find out.</p>
<p>At best that advice was raw material. At worst, and it is the worst far more often than anyone admits, it is the exact mechanism by which a business manufactures a scapegoat and calls it governance.</p>
<h2 id="what-the-human-in-the-loop-is-actually-for">What the human in the loop is actually for</h2>
<p>Strip away the comfort and a human in the loop has exactly one job: to catch the error. To look at what the machine produced, know whether it is right, and stop it when it is wrong. That is it. That is the whole value. Not to be present. Not to be named in a policy. To catch it.</p>
<p>And here is the part that should make you uncomfortable about your own setup. That ability is often zero. Sometimes zero by accident, sometimes zero by design, and in almost every case the organisation has no idea it is zero, because the human is sitting right there in the diagram, looking for all the world like a safeguard.</p>
<p>There are four ways that ability collapses to nothing.</p>
<p><img alt="Card: four ways a human in the loop is worth nothing. They can’t verify it, they can’t keep up, they won’t push back, they aren’t equipped." src="https://jeffops.com/newsletter/human-in-the-loop/four-ways.png" /></p>
<h2 id="1-they-cant-verify-it">1. They can’t verify it</h2>
<p>Some tasks cannot be checked. Not “are hard to check”. Cannot, in principle, be verified quickly enough to matter.</p>
<p>I made this the third question of my own AI Fit Test: can you verify the output? And I will be honest about what I quietly glossed over when I wrote it. I treated verifiability as a property of the task. It is not. It is a property of the task <strong>and</strong> the human you assigned to it.</p>
<p>A verifiable task with a reviewer who cannot perform the verification is an unverifiable task wearing a hi-vis vest.</p>
<p>If the human in the loop cannot independently tell right from wrong on this specific output, because it needs judgement, or context they do not have, or a level of scrutiny the workflow does not allow, then they are not verifying anything. They are witnessing. And a witness is not a control.</p>
<h2 id="2-they-cant-keep-up">2. They can’t keep up</h2>
<p>Say the task is verifiable and the human genuinely could catch the error. Once. Carefully. With coffee and time. Now give them five hundred of those tasks a day.</p>
<p>This is the failure nobody wants to cost out, because it is the one that makes the business case work. The whole pitch of AI is volume: more, faster, cheaper. But the human in the loop does not get faster.</p>
<p>So, what actually happens is that the drowning reviewer stops reviewing and starts rubber-stamping. They approve on vibes. They spot-check one in fifty and trust the rest. “Human in the loop” quietly degrades into “human near the loop, occasionally glancing at it”. The overseer is technically present and functionally gone, and the throughput numbers look fantastic right up until the one that matters sails through unread.</p>
<h2 id="3-they-wont-push-back">3. They won’t push back</h2>
<p>This is the ugly one, because it is not about capacity or skill. It is about wiring. Yours.</p>
<p>There is a well-documented phenomenon called automation bias: humans tend to favour what an automated system tells them over the other information sitting in front of them, even when that other information is correct. Kathleen Mosier and Linda Skitka named its two flavours in the nineties, and both should worry you. Omission, where the automation says nothing about a problem and you do not act, because it never flagged it. Commission, where you act on what the automation tells you without cross-checking what is right next to it. Note that neither of those requires you to notice anything.</p>
<p>That is the part people get wrong. They picture a reviewer wrestling with doubt and losing. Most of the time there is no wrestle, because there was no second look. The harder version is real, where you see the contradiction and defer anyway, and it is much rarer than the version where you simply never checked. Which is the trap inside the trap. You put the human there to be a check, and the psychology quietly re-points their instincts to defend the machine instead.</p>
<p>Now connect it to the thing I hammered in the very first edition: the illusion of competence. AI fails convincingly. Its wrong answers are as fluent and as confident as its right ones. And here is the finding that ties the bow. Studies show that when you display the automation’s confidence, people become more likely to go along with the high-confidence output, with no actual improvement in accuracy.</p>
<p>So the confident tone does not merely fail to help your reviewer. It disarms them. The very thing that makes AI dangerous is the very thing that makes your safeguard defer to it. Your human in the loop is fighting their own neurology and losing.</p>
<h2 id="4-they-arent-equipped">4. They aren’t equipped</h2>
<p>The quietest failure. You needed a specialist in the loop and you put a generalist there, because “there’s a human checking it” felt like enough and the specialist was (of course) expensive.</p>
<p>There is a human checking it. They cannot tell a subtle-but-catastrophic answer from a correct one, because the task demands expertise they do not have. So they check the things they can see, being spelling, tone and format, and wave through the thing they cannot.</p>
<p>The loop is staffed. The loop is useless. And on paper it looks identical to a loop that works.</p>
<h2 id="the-moral-crumple-zone">The moral crumple zone</h2>
<p>Here is where I stop being a smart-ass and hand you somebody else’s concept, because it will change how you see every one of these setups.</p>
<p>The cultural anthropologist named Madeleine Clare Elish studied what happens to accountability in heavily automated systems, and gave the pattern a name: the moral crumple zone. A car’s crumple zone is engineered to absorb the force of a crash, protecting the passengers by sacrificing itself. A moral crumple zone does the same thing with blame. When a complex automated system fails, the human operator nearest to it absorbs the impact, the responsibility, the liability, the headlines. And in doing so protects the integrity of the system. Next to that, it protects the company behind it, and the vendor who sold it.</p>
<p>The human’s real function in that design is not to prevent the failure. It is to absorb it. A face to point at when things go wrong, with no realistic power to make things go right.</p>
<p><img alt="Card: the moral crumple zone. When a complex automated system fails, the human nearest to it absorbs the impact, protecting the integrity of the system, the company behind it and the vendor who sold it. Concept: Madeleine Clare Elish." src="https://jeffops.com/newsletter/human-in-the-loop/moral-crumple-zone.png" /></p>
<p>Elish builds the concept on two accidents, and then points it at a third that happened while she was writing.</p>
<p><strong>Three Mile Island, 1979.</strong> Partial nuclear meltdown. The Kemeny Commission that investigated it ranged wide: it demanded fundamental change in the attitudes of the regulator and the industry, called operator training greatly deficient, found procedures that could be read as leading operators to the wrong action, and described a control room with hundreds of alarms and no way to suppress the unimportant ones. In the newspapers it became operator error. “Nuclear Accident Blamed Primarily on Human Error”, ran the Los Angeles Times. The operators were the crumple zone, and it was the coverage that put them in it.</p>
<p><strong>Air France 447, 2009.</strong> 228 people dead in the Atlantic. The pitot tubes iced over, the airspeed readings disagreed, and the autopilot disconnected. Which is exactly what it was built to do. Elish’s line on it is the one that should stop you cold: “because the autopilot did not malfunction in a way recognized through its certification process, the only possible malfunction, systemically, is the human pilot.” Note that this is a claim about certification logic rather than about the investigation, which named the pitot probes, the stall warning, the airspeed display and high-altitude handling training as well. But inside the framework that decides what is allowed to count as a fault, the human was the only thing that could be one. That is a moral crumple zone by design, not by accident.</p>
<p>And the third, which opens her paper rather than featuring as a case study in it:</p>
<p><strong>Uber, Tempe, 2018.</strong> Elaine Herzberg, the first pedestrian killed by a self-driving car. The safety driver was positioned to supervise a system she had almost no realistic way to meaningfully supervise in that split second, and she became the locus of responsibility. She pleaded guilty to endangerment and got probation. Prosecutors declined to charge Uber at all, in a case where the NTSB had found the factory emergency braking disabled, the company’s own software suppressing braking for a full second while it decided, and no mechanism anywhere for managing the operator complacency the whole arrangement depended on.</p>
<p>There is a tell that the industry knows this problem is real. Google’s self-driving programme concluded it could not reliably solve the handoff, passing control back to a human for exactly the rarest, hardest, most dangerous moments, and changed its whole approach because of it.</p>
<p>Think about what that means for your office workflow. A company with self-driving-car money looked at “human catches the edge cases the machine can’t” and decided it does not work well enough to bet lives on. And you are relying on it to catch a hallucinated figure in a report nobody has time to read!</p>
<p>Now hold that concept up against my own advice. “Name a human owner for every AI outcome.” I wrote that. I meant it as accountability. But look how cleanly it can be perverted. The dishonest way to satisfy “name an owner” is to name someone who cannot do the job, with no time, no expertise, no authority and no real ability to catch the error, and let them hold the responsibility anyway.</p>
<p>You will have followed my advice to the letter. You will have built a moral crumple zone and filed the paperwork that proves it.</p>
<p>I told you to name a defendant and I called it governance. That one is on me, and I am correcting it here.</p>
<h2 id="four-questions-and-answer-them-about-the-real-setup">Four questions, and answer them about the real setup</h2>
<p>You have a human in the loop somewhere in your business right now. Is it a safeguard or a sponge? Answer these honestly, about what actually happens, not about the diagram.</p>
<p><img alt="Card: four questions to answer about the real setup, not the diagram. Time, expertise, authority, detection. All four yes is a circuit breaker, any one no is a crumple zone." src="https://jeffops.com/newsletter/human-in-the-loop/four-questions.png" /></p>
<p><strong>Do they have the time?</strong> Real minutes per item to check properly, or a queue that guarantees rubber-stamping?</p>
<p><strong>Do they have the expertise?</strong> Can this person genuinely tell a good output from a subtly broken one on this task, or only judge the surface?</p>
<p><strong>Do they have the authority?</strong> When they say “no, this is wrong”, does it stop? Or do they get overruled by a deadline, a manager, or a throughput target?</p>
<p><strong>Can they detect the error at all?</strong> Is the failure the kind a human can catch, or the silent, fluent, convincing kind that slips past everyone, including them?</p>
<p>Those four should look familiar. They are last edition’s rules for naming an owner, flipped from instructions into a test.</p>
<p>Any “no” and you do not have oversight. You have theatre. All the world's a stage, and your reviewer is merely a player. A person positioned to look like a control while being structurally incapable of acting as one.</p>
<p>And the cruel part is that theatre is worse than an honest absence. An empty loop makes you nervous enough to check. A staffed-but-useless loop makes you feel safe. It does not just fail to catch the error. It removes the fear that would have made you catch it yourself.</p>
<h2 id="where-that-leaves-keep-a-human-in-the-loop">Where that leaves “keep a human in the loop”</h2>
<p>Demoted. From a solution to a precondition that has preconditions of its own.</p>
<p>Adding a human is not a safeguard. It is the raw material for one. It becomes a safeguard only when that human clears all four criteria above, and it reverts to a crumple zone the moment they miss even one. Meet them and the loop is a circuit breaker, built to stop the failure. Miss them and it exists only to absorb the blame for a failure the safeguard was never able to prevent.</p>
<p>And here is what I most need you to take from all of this, because it follows from everything above and it is harder than anything I have said in this series so far. If you cannot meet those preconditions for a given use case, if there is no human who can actually catch the error at the speed and the scale you need, then the answer is not “ship it with a human in the loop anyway”. The answer is don’t ship it.</p>
<p>Adding an overwhelmed, underqualified, unheard human to an unsafe system does not make it safe. It just decides, in advance, who takes the fall.</p>
<p>So don’t put a human in the loop so that you can <em>feel</em> safe. Put one there who can make you safe, or don’t automate that thing at all. A tool to enhance the judgement of someone who has some, not to replace it with a signature.</p>
<p>A human in the loop is a safeguard. A human on the hook is a sponge. Most businesses cannot tell which one they have built right up until it fails, and a real person who never stood a chance of catching the fail gets handed the blame the system was quietly designed to shed.</p>
<p>Make sure you know which one you built. Before it matters 😉</p>]]></content:encoded>
  </item>
  <item>
    <title>AIccountability: you can automate the work, you can&#x27;t automate the blame</title>
    <link>https://jeffops.com/newsletter/aiccountability/</link>
    <guid isPermaLink="true">https://jeffops.com/newsletter/aiccountability/</guid>
    <pubDate>Tue, 04 Aug 2026 09:00:00 +0000</pubDate>
    <description>Accountability is not a capability, and no model is coming that can hold it. Why &quot;the AI did it&quot; explains nothing, and the four rules for designing accountability in on purpose.</description>
    <category>Newsletter</category>
    <content:encoded><![CDATA[<p><strong>You can automate the work. You can’t automate the blame.</strong></p>
<p>Read that again, because it’s the whole article and it’s the part everyone rushing to bolt AI onto their business is quietly pretending isn’t true.</p>
<p>I’ve written four pieces now on how AI works, where it fits, how to learn with it, and how to plot it against your business. Accountability has been standing in the corner of every single one of them, unaddressed. “The AI isn’t accountable, so the responsibility falls on you.” “No agent owns the outcome.” “Responsibility-heavy tasks — bad fit.” I kept pointing at it and walking past. This time we stop and take a deeper look at it directly.</p>
<p>So let me be blunt for the length of one article. If you’re deploying AI and you haven’t answered the question “when this is wrong, who answers for it?” you haven’t deployed a tool. You’ve deployed a liability. And you’ve probably handed it to someone who doesn’t know they’re holding it.</p>
<h2 id="why-ai-can-never-be-accountable">Why AI can never be accountable</h2>
<p>Let’s kill the comfortable idea some people have first: that accountability is a feature the next model will ship.</p>
<p>It won’t. Ever. Not because the models aren’t good enough, but because accountability isn’t a capability. It’s not something you get better at by adding parameters. Ask yourself what being accountable actually means. It means there are consequences that land on you. You can be fired. Sued. Fined. Struck off. Made to pay something back. Made to explain yourself to someone who’s furious. Made to lie awake at night knowing you got it wrong.</p>
<p>Now point any of that at a model. Fire it? Sue it? Fine it? There’s nothing there to punish and nothing there that cares. It has no license to lose, no reputation to protect, no skin in the game, or just no game at all. It produces an output and moves on the same whether it just saved you an afternoon or torched a client relationship.</p>
<p>That’s not a bug in this generation of AI. It’s the definition of a tool. A hammer isn’t accountable for the wall. A spreadsheet isn’t accountable for the forecast. And an AI isn’t accountable for the decision no matter how much it sounds like it made one. Stop waiting for the version that can hold responsibility. There isn’t one coming.</p>
<h2 id="the-accountability-gap">The accountability gap</h2>
<p>Here’s the part people don’t think thoroughly enough about. Accountability doesn’t vanish when you hand the work to something that can’t hold it. It’s afantasy that automating the task also automates away the responsibility for it. It doesn’t. The work moved, but the blame didn’t. It’s still sitting there, fully intact, looking for a human to land on.</p>
<p>And it always finds one. The developer who shipped the feature. The employee who copy-pasted the output into the email. The manager who signed off on the rollout. The founder who decided the whole thing should run on agents. Somebody in that chain is going to answer for it when it breaks. And the further from the keyboard you get and let the AI do the work on your behalf, the more likely it’s you, not the model, whose name is on it, while you didn’t even actually do it.</p>
<p>This is the accountability gap: the distance between the thing that did the work and the person who owns the outcome. AI doesn’t close that gap. It widens it and then quietly drops the whole weight of it onto whoever was standing closest.</p>
<h2 id="the-ai-did-it">“The AI did it”</h2>
<p>Which brings us to the sentence I want you to ban from your vocabulary.</p>
<p>“The AI did it.”</p>
<p>It feels like an explanation, but it’s a reason and never an excuse. It functions like nothing. It is, precisely, a bad workman blaming his tools and you already know how that story ends. Nobody sympathizes with the workman. Nobody says “ah, well, the chisel slipped, not your fault.” They held him responsible, because he’s the one who picked up the chisel, and he’s the one who was supposed to know how to use it.</p>
<p>“The AI generated the wrong number” is not a defense. You chose to use AI for that task. You chose not to check it. You chose to put it in front of a customer. Every one of those was a human decision, and the model didn’t make a single one of them. Hiding behind “the AI did it” doesn’t transfer the blame, it just advertises that you didn’t understand where the blame was sitting the whole time.</p>
<p>No matter how the failure happens, no matter how many layers of automation it passes through, the blame lands on a person. Every time. The only real question — the one you should be asking before you deploy anything — is which person, and whether they signed up for it.</p>
<h2 id="the-disclaimer-shield">The disclaimer shield</h2>
<p>Here’s a small five-word sentence that tells you the entire industry already knows this.</p>
<p>Look under the prompt box in Microsoft’s Copilot: <em>AI-generated content may be incorrect.</em></p>
<p>Sit with that for a second, because it’s not a friendly heads-up. It’s a legal position. The company that built the tool, sold you the tool, and profits from the tool is telling you, in writing, that it will not stand behind what the tool produces. Read the terms of service on basically any of these products and you’ll find the same thing dressed up in more words: the provider disclaims responsibility, and the liability for what you do with the output is yours.</p>
<p>Think about what that means for the chain of blame. The AI can’t be accountable, we covered that, and now the vendor has explicitly, contractually stepped out of the way too. So when the output is wrong and something breaks, walk the line back: not the model, not the company that made the model. Who’s left standing?</p>
<p>You are. You were always the one left standing. The disclaimer isn’t a warning label, but a transfer of custody, and you accepted it the moment you hit enter or clicked ‘accept’ without reading page upon page of legalese writing just to use a simple tool.</p>
<h2 id="when-the-chain-erases-accountability">When the chain erases accountability</h2>
<p>Now make it worse, the way real businesses make it worse: chain the systems together.</p>
<p>I wrote about this in the first article — the multi-agent pipeline where one model’s output feeds the next, and a subtle error five handoffs upstream travels downstream, compounds, and lands in the final result with no stack trace to trace it back. Confident output at every step. No error thrown anywhere. Just a broken outcome and no obvious place to point.</p>
<p>Watch what that does to accountability. In a single system, at least you know where the failure happened. Distribute the work across a chain of agents and you distribute the failure and blame too, and although distributed failure doesn’t automatically becomes distributed blame, it does have a nasty habit of rounding down to nobody’s blame. “It wasn’t my agent, mine worked fine.” “The input it got was already wrong.” Everyone in the chain is technically correct and the outcome is still broken and somehow no single component owns it. You’ll end up with a situation I’d describe as: “The operation was a great success, but the patient died.”.</p>
<p>But the drumbeat holds: the blame didn’t disappear just because you can’t locate it. The customer who got burned holds your business responsible, not your fifth-agent-in-the-chain.</p>
<h2 id="you-own-the-pager">You own the pager</h2>
<p>If you’ve done any ops work, this whole article is already familiar, because you’ve probably lived it.</p>
<p>Your service goes down at 3am. The root cause is a dependency you didn’t write — some library maintainer’s bug, some upstream provider’s outage. And you know exactly how much that matters when the pager goes off: not at all. You don’t get to reply to the incident with “not my code.” You own the service. You’re on-call for it. The fact that someone else’s component failed doesn’t move the responsibility one inch. It was your job to know your dependencies could fail and to build accordingly.</p>
<p>For some AI is becoming a dependency. The most confident, most convincing, least accountable dependency you’ve ever used. And it does not come with its own on-call rotation. When its output breaks something, the pager goes off on your phone.</p>
<p>That’s the reframe I want you to leave with. Accountability was never about who typed the line or generated the paragraph. It was always about who answers when it breaks. And a tool — any tool, including this one — cannot answer. It can’t hold the pager. Someone with a pulse has to, and that someone is you, whether you planned for it or not.</p>
<h2 id="so-how-do-you-design-it-in">So how do you design it in?</h2>
<p>Everything above is the diagnosis. Here’s the practical part, because “keep a human accountable” is easy to nod along to and easy to get wrong. If accountability has to be built in on purpose, this is what “on purpose” actually looks like:</p>
<ul>
<li><strong>Name the owner before you ship, not after it breaks:</strong> every AI-touched outcome has one human whose name is on it. If you’re assigning the owner during the post-mortem, you’ve already lost.</li>
<li><strong>Give that owner real authority:</strong> the power to catch the error and overrule the output. Accountability without authority isn’t a safeguard; it’s a scapegoat with a job title.</li>
<li><strong>Match the checker to the stakes:</strong> the more expensive being wrong is, the more the human in the loop has to actually verify the output, not rubber-stamp it. Low-stakes, let it ride. High-stakes, someone competent reads every word. Note: Watch out for the things that seem low-stake, and are (or become) high-stake 😉</li>
<li><strong>No owner, no use case:</strong> if you can’t name the person who answers for a given use case, you haven’t found a use case an unassigned liability waiting for the worst possible moment to find an owner. Because when something has become a problem/liability, who actually <em>wants</em> to be its owner?</li>
</ul>
<h2 id="conclusion-you-dont-offload-accountability-you-design-to-keep-it">Conclusion — you don’t offload accountability, you design to keep it</h2>
<p>So here’s the shift, and it’s the same shape as everything I’ve argued across this series. You don’t “adopt AI” and let responsibility sort itself out, because the one thing you cannot automate is the answering-for-it. You design it in, on purpose, or it lands on someone by accident.</p>
<p>Underneath the four rules above sits a single one: no outcome your business ships is allowed to be owned by “the AI.” Ever. That’s the truth the disclaimer told you and the pager taught you. Use AI for everything it’s brilliant at, being the structured, the repeatable, the verifiable, just never let it convince you it’s holding a responsibility it fundamentally can’t.</p>
<p>You can automate the work. You can’t automate the blame. A tool to enhance, not to replace and never a name to hide behind.</p>
<p>But here’s the catch, and it’s the one that’s been quietly undermining half of what I just told you. “Keep a human in the loop” (the answer I keep reaching for) only works if that human can actually catch the error. And thanks to everything I’ve written about the illusion of competence, that’s a much bigger <em>if</em> than it looks. Sometimes the human in the loop isn’t a safeguard at all. Sometimes they’re just there to absorb the blame for a system nobody could realistically supervise.</p>
<p>That’s the next edition. Safeguard, or blame sponge? We’ll find out.</p>]]></content:encoded>
  </item>
  <item>
    <title>MTA-STS on Microsoft 365, with the policy hosted on GitHub Pages</title>
    <link>https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/</link>
    <guid isPermaLink="true">https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/</guid>
    <pubDate>Mon, 03 Aug 2026 22:00:00 +0000</pubDate>
    <description>Your inbound mail already uses TLS. That is not the same as requiring it. Setting up MTA-STS for Microsoft 365, with the policy hosted free on GitHub Pages, in the order that avoids the mistakes. Works whether or not Microsoft hosts your DNS.</description>
    <category>email</category><category>security</category><category>mta-sts</category><category>dns</category><category>microsoft-365</category><category>github-pages</category><category>cloudflare</category><category>tls</category>
    <content:encoded><![CDATA[<p>The <code>id</code> on my own <code>_mta-sts</code> record reads <code>20260804045609Z</code>. And I published the first version of that record before there was anything behind it for a sender to fetch, so for a while jeffops.com was announcing a policy that did not exist.</p>
<p>Nothing broke. That is the problem with getting this wrong: it looks exactly like getting it right.</p>
<p>So this one is for anyone running mail on Microsoft 365 who would rather inbound TLS was required than hoped for. You do not need to be a mail specialist and you do not need a server. Twenty minutes, three DNS records and a free GitHub repository.</p>
<p>Here is the line the whole thing hangs on: <strong>your inbound mail already uses TLS. That is not the same as requiring it.</strong></p>
<h2 id="but-my-mail-is-already-encrypted">But my mail is already encrypted?</h2>
<p>It is. Every mail server that delivers to you already tries TLS. It sends <code>STARTTLS</code>, your server agrees, the session is encrypted, and everyone feels fine about it.</p>
<p>Now think about what happens when the handshake does not work. The sending server shrugs and delivers in plaintext instead, because that is what the protocol tells it to do. Nobody anywhere gets told.</p>
<p>That is the whole attack. Sit between two mail servers, strip <code>STARTTLS</code> out of the greeting, and the message arrives in the clear. No certificate warning, no bounce, no log entry that looks like anything.</p>
<p>It is the difference between asking for a signature and requiring one. A courier who is supposed to get a signature, finds nobody home, and puts the parcel through the letterbox anyway has still delivered it. You even get your parcel. You just never find out it spent the afternoon on the mat, and neither does the person who sent it.</p>
<p>MTA-STS is the signature requirement. It is a published promise, fetched over HTTPS, that says <em>this domain requires TLS, and here are the servers allowed to accept it</em>. A sender that reads that promise will not silently downgrade.</p>
<h2 id="what-you-are-actually-building">What you are actually building</h2>
<p>Three things, and only three:</p>
<ol>
<li>A TXT record at <code>_mta-sts.yourdomain</code> saying a policy exists.</li>
<li>A policy file at <code>https://mta-sts.yourdomain/.well-known/mta-sts.txt</code>.</li>
<li>A certificate valid for <code>mta-sts.yourdomain</code> exactly, serving that file.</li>
</ol>
<p>Point three is the one that catches people. The policy cannot live on your website. It has to be served from the <code>mta-sts.</code> subdomain, on a certificate for that host, over HTTPS. Which sounds like it needs a server. It does not. GitHub Pages will do it for nothing, and it will provision the certificate for you.</p>
<p>Next to that there is a fourth thing, not part of the standard and worth having anyway: TLS-RPT, a separate TXT record telling senders where to post reports when TLS fails. MTA-STS without TLS-RPT is a policy that never tells you it is breaking.</p>
<h2 id="first-find-out-where-your-dns-actually-lives">First, find out where your DNS actually lives</h2>
<p>I host my DNS at Microsoft 365 itself, so the screenshots here are the Microsoft 365 admin centre. Worth stating up front, because that is a different thing from having Microsoft 365 mail, and plenty of people have the second without the first. Your mailboxes can live in Exchange Online while your zone sits at your registrar, at Cloudflare, at Route 53, or on a domain controller in a cupboard.</p>
<p>Two ways to check. In the admin centre, go to <strong>Settings → Domains → your domain → DNS records</strong>. If there is an <strong>Add record</strong> button, Microsoft hosts your DNS. If instead the page names your DNS hosting provider and sends you there, Microsoft does not, and the records listed are only Microsoft telling you what it would like you to create elsewhere.</p>
<p>Faster, from a terminal:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code>nslookup -type=NS yourdomain.com
</code></pre>
<p>Microsoft-hosted zones answer with <code>ns1.bdm.microsoftonline.com</code> through <code>ns4</code>. Anything else and you are somewhere else.</p>
<p><strong>If Microsoft hosts your DNS</strong>, the screenshots below are what you will see, and you inherit a ceiling. Hold that thought until the last section.</p>
<p><strong>If it does not</strong>, nothing here changes except where you type. The three records are identical wherever they live: a CNAME at <code>mta-sts</code>, a TXT at <code>_mta-sts</code>, a TXT at <code>_smtp._tls</code>. But three things do differ by provider, and all three are worth knowing before you start.</p>
<p><strong>Relative names versus fully qualified ones.</strong> Some control panels want <code>_mta-sts</code> and append the domain for you. Others want <code>_mta-sts.yourdomain.com</code> in full. Give the wrong one and you quietly create <code>_mta-sts.yourdomain.com.yourdomain.com</code>, which resolves for nobody and looks perfectly fine in the interface. So query the record after saving it, always, rather than trusting the panel.</p>
<p><strong>Cloudflare, specifically: the <code>mta-sts</code> record must be DNS only. Grey cloud, not orange.</strong> Proxy it with your SSL mode on Full or Full (Strict) and GitHub cannot renew the certificate, because Cloudflare refuses to connect to an origin whose certificate has expired and the renewal is exactly what needs that connection. The failure mode is nasty. It works perfectly for ninety days, then breaks… and then keeps breaking every ninety days. Not really a Cloudflare bug so much as what happens with any CDN sitting between your domain and GitHub Pages.</p>
<p><strong>Underscore labels.</strong> Most of the times you will never meet this one, but a few older registrar panels still refuse a record whose name begins with an underscore. If yours is one of them, MTA-STS is the least of it, because DMARC, DKIM and TLS-RPT all need the same thing. Move your DNS.</p>
<p>And there is one consolation for hosting elsewhere, a real one. Come back to the last section.</p>
<h2 id="do-it-in-this-order">Do it in this order</h2>
<p>The order matters, and I got it wrong the first time, which is where this article started. Harmless in testing mode, and every checker simply reports "no policy", but it is backwards.</p>
<p>Build it so each step is only taken once the thing it points at is real:</p>
<ol>
<li>The CNAME, so GitHub can verify the hostname and issue a certificate.</li>
<li>The repository and the policy file.</li>
<li>Pages, the custom domain, and HTTPS.</li>
<li>The TXT record, last, once there is genuinely something to fetch.</li>
</ol>
<h2 id="step-1-the-cname">Step 1: the CNAME</h2>
<p>In the Microsoft 365 admin centre: <strong>Settings → Domains → your domain → DNS records → Add record</strong>.</p>
<p>Before you fill anything in, open the Type dropdown and look at it, because it tells you something about what you can and cannot do here later:</p>
<p><img alt="The Microsoft 365 add-record dialog with the Type dropdown open, showing TXT, CNAME, A, AAAA, SRV and MX" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/01-record-types.jpg" /></p>
<p>Six types. TXT, CNAME, A, AAAA, SRV, MX. That is the whole list. Note what is missing, because we come back to it at the end.</p>
<p>Add a CNAME. Host name <code>mta-sts</code>, pointing at your GitHub Pages host, which is <code>yourusername.github.io</code>:</p>
<p><img alt="Adding a CNAME record named mta-sts pointing at jeffwouters.github.io" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/02-cname.jpg" /></p>
<p>Save it, and let it settle. The TTL here is an hour, so if you race ahead to GitHub in the next thirty seconds you will sit staring at "DNS check in progress" for no reason at all.</p>
<h2 id="step-2-a-repository-for-the-policy">Step 2: a repository for the policy</h2>
<p>The policy needs its own repository. Not a folder in your website repo, its own, because GitHub Pages allows exactly one custom domain per repo and yours is already spoken for by your site.</p>
<p>Create it public. Pages does not serve private repositories on a free account, and there is nothing secret in here anyway. The entire contents will be public DNS policy that you are actively trying to get strangers to read.</p>
<p><img alt="Creating a new public GitHub repository called mta-sts" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/03-new-repo.jpg" /></p>
<p>Then upload four files:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code>CNAME        mta-sts.yourdomain.com
.nojekyll    (empty)
index.html   (optional, a page for anyone who visits directly)
README.md    (optional, but future-you will want it)
</code></pre>
<p><img alt="Uploading CNAME, .nojekyll, index.html and README.md to the repository" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/04-upload-files.jpg" /></p>
<p><strong><code>.nojekyll</code> is not optional and it is the single most likely thing to silently break this.</strong> GitHub Pages runs Jekyll by default. Jekyll ignores every file and directory whose name begins with a dot. And the policy is required to live at <code>.well-known/mta-sts.txt</code>.</p>
<p>So without <code>.nojekyll</code> that path is never published, the URL returns 404, and every validator you try reports "no policy found" with no hint as to why. You will check your DNS four times before you think of Jekyll. I did… four times.</p>
<h2 id="step-3-the-policy-file">Step 3: the policy file</h2>
<p>Create <code>.well-known/mta-sts.txt</code>. In the GitHub web editor you can type the whole path into the filename box and it makes the directory for you.</p>
<p><img alt="Creating .well-known/mta-sts.txt in the GitHub web editor" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/05-policy-file.jpg" /></p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code>version: STSv1
mode: testing
mx: *.mail.protection.outlook.com
max_age: 604800
</code></pre>
<p>Four lines, and three of them deserve a sentence.</p>
<p><strong><code>mode: testing</code></strong> means a sending server that cannot establish valid TLS reports the failure and delivers the mail anyway. Nothing can be lost while this is set. The alternative, <code>enforce</code>, means it refuses to deliver instead. So start in testing. Always.</p>
<p><strong><code>mx:</code></strong> must match your real MX records. For Microsoft 365 that is <code>yourdomain-com.mail.protection.outlook.com</code>, and the wildcard <code>*.mail.protection.outlook.com</code> covers it, because <code>*</code> matches the leftmost label only and that is one label. Get this wrong and, once you move to enforce, senders will start refusing to deliver to you. So check it against your actual MX record rather than trusting a template… including this one.</p>
<p><strong><code>max_age: 604800</code></strong> is seven days, which is how long a sender caches the policy. Short while testing. Longer once you are confident.</p>
<p>The repository should now look like this:</p>
<p><img alt="The mta-sts repository containing .well-known, .nojekyll, CNAME, README.md and index.html" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/06-repo-contents.jpg" /></p>
<h2 id="step-4-turn-on-pages">Step 4: turn on Pages</h2>
<p><strong>Settings → Pages → Source: Deploy from a branch</strong>, branch <code>main</code>, folder <code>/ (root)</code>.</p>
<p><img alt="Selecting the main branch as the GitHub Pages source" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/07-pages-branch.jpg" /></p>
<p>The custom domain fills itself in from your CNAME file. Because you added the DNS record first, the check passes immediately rather than sitting in "in progress":</p>
<p><img alt="GitHub Pages showing the custom domain mta-sts.jeffops.com with DNS check successful" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/08-dns-check.jpg" /></p>
<p>Give it a couple of minutes for the certificate, then tick <strong>Enforce HTTPS</strong>:</p>
<p><img alt="Enforce HTTPS enabled on the GitHub Pages settings" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/09-enforce-https.jpg" /></p>
<p>Now confirm the policy is actually being served, over HTTPS, with a valid certificate. Load the URL in a browser. If you typed <code>http://</code> and it upgraded itself to <code>https://</code> without a warning, that is the certificate working:</p>
<p><img alt="The policy file served over HTTPS at mta-sts.jeffops.com" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/10-policy-live.jpg" /></p>
<h2 id="step-5-the-txt-record-now-that-it-points-at-something">Step 5: the TXT record, now that it points at something</h2>
<p>Back to the admin centre. Add a TXT record, name <code>_mta-sts</code>, value:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code>v=STSv1; id=20260804045609Z
</code></pre>
<p><img alt="Adding the _mta-sts TXT record in the Microsoft 365 admin centre" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/11-txt-record.jpg" /></p>
<p>The <code>id</code> is an opaque string of at most 32 alphanumeric characters. It has no meaning to anything. A UTC timestamp is the convention because it makes the last change self-documenting, which is the only reason you know when I did mine.</p>
<p>And it carries the second trap. <strong>Senders cache the policy and only re-fetch it when the <code>id</code> changes.</strong> So edit <code>mta-sts.txt</code> without bumping the <code>id</code> and you have changed nothing at all, for up to <code>max_age</code>. Every time you touch the policy file, change the id in DNS.</p>
<h2 id="step-6-tls-rpt">Step 6: TLS-RPT</h2>
<p>One more TXT record, at <code>_smtp._tls</code>:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code>v=TLSRPTv1; rua=mailto:you@yourdomain.com
</code></pre>
<p>That is the bit that tells you when TLS to your domain is failing. It has no enforcement behaviour of its own and no way to break anything. There is no good reason to run MTA-STS without it.</p>
<h2 id="step-7-check-it-from-outside">Step 7: check it from outside</h2>
<p>Do not take your own word for it, and do not take mine. Run it through an external validator:</p>
<p><img alt="Mailhardener MTA-STS validator reporting the domain is set up correctly" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/12-validator.jpg" /></p>
<p>The detail is where the useful confirmation lives. Policy fetched, HTTP 200, correct content type, certificate valid, and the policy service marked valid:</p>
<p><img alt="The validator's detailed report showing policy service valid and certificate valid" src="https://jeffops.com/posts/2026/mta-sts-microsoft-365-github-pages/13-validator-detail.jpg" /></p>
<p>Worth reading that panel properly rather than stopping at the green box. It tells you the resolved address of your policy host, the CAA records it inherits, the raw policy, the line ending style, and the certificate expiry. If anything is going to be quietly wrong, it will be visible there.</p>
<h2 id="moving-to-enforce">Moving to enforce</h2>
<p>Leave it in testing for a few weeks and read the TLS-RPT reports. When they are quiet, change <code>mode: testing</code> to <code>mode: enforce</code>, bump the <code>id</code>, and consider raising <code>max_age</code>.</p>
<p>Enforce is the point of the exercise. It is also the mode where a wrong <code>mx:</code> line stops mail reaching you, which is why the MX pattern got its own paragraph earlier and why testing exists at all.</p>
<h2 id="the-ceiling-if-microsoft-hosts-your-dns">The ceiling, if Microsoft hosts your DNS</h2>
<p>Go back to that dropdown from step one. Six record types, and no CAA.</p>
<p>CAA is the record that says which certificate authorities are allowed to issue certificates for your domain. Without one, any CA on earth can. Microsoft's DNS hosting does not offer the record type, so on a domain hosted there you simply cannot have one. Nor DNSSEC, which is not on that list either.</p>
<p>There is a small irony in what you have just built. <code>mta-sts.yourdomain</code> ends up better protected than your apex, because it is a CNAME to <code>github.io</code> and inherits GitHub's CAA records through it. The validator screenshot above shows exactly that. Your main domain has nothing equivalent.</p>
<p>Your policy host is safer than the domain it exists to protect!</p>
<p><strong>All of you hosting DNS somewhere other than Microsoft:</strong> this is the consolation I promised. At Cloudflare, Route 53, or most registrars worth the name, CAA is a record type like any other and DNSSEC is a toggle. Add both while you are in there. It is ten minutes and it closes a gap this article cannot.</p>
<p>And if you are on Microsoft's DNS and either matters to you, the fix is not a setting. It is moving your zone somewhere that supports them, which is a (much) bigger decision than an afternoon of MTA-STS and not one to make on the way past.</p>
<h2 id="what-you-end-up-with">What you end up with</h2>
<p>Three DNS records, one small repository, and about twenty minutes. Inbound mail that used to encrypt itself when convenient now encrypts itself because your domain says so, and tells you when it cannot.</p>
<p>MTA-STS is a tool to enhance the TLS you already had, not to replace it.</p>
<p>It just stops it being optional.</p>]]></content:encoded>
  </item>
  <item>
    <title>Use AI to learn, but don&#x27;t let it be your teacher</title>
    <link>https://jeffops.com/newsletter/use-ai-to-learn/</link>
    <guid isPermaLink="true">https://jeffops.com/newsletter/use-ai-to-learn/</guid>
    <pubDate>Mon, 20 Jul 2026 09:00:00 +0000</pubDate>
    <description>AI answers questions. Learning requires asking the right ones. How to use it to build understanding instead of quietly skipping past it.</description>
    <category>Newsletter</category>
    <content:encoded><![CDATA[<p>I've spent the last two articles on where AI fits and where it doesn't, as a business tool and as a coding tool. This one is more personal, because it's about you (and me) using AI to actually learn something.</p>
<p>And I'll start with the line this whole edition hangs on: AI answers questions. Learning requires asking the right ones.</p>
<h2 id="why-ai-feels-like-a-great-teacher">Why AI feels like a great teacher</h2>
<p>On the surface, AI looks like the perfect teacher. It never tires of your questions. It explains anything, instantly, in clean and patient language. Ask it to go simpler and it goes simpler. Ask again and it tries a different angle. No judgement, no waiting, available at 2am. If you'd sat down to design a tutor from scratch, it might look a lot like this.</p>
<p>These models were trained on an enormous pile of human explanation. Documentation, tutorials, textbooks, Q&amp;A threads, endless "explain like I'm five" posts. Explaining things is one of the most common tasks in the whole training set. But here's the tension sitting underneath all of it: AI doesn't guide your learning, it responds to your input. Those are very different things.</p>
<p>A real teacher steers. They decide what you should look at next, notice when you've skipped something important, and drag you back to it. AI doesn't steer. It reacts. It answers the question you asked, at the level you asked it, and then sits there waiting for the next one. Which means the quality of your learning is capped by the quality of your questions. If you don't know what to ask, it isn't going to tell you.</p>
<h2 id="the-missing-piece-it-doesnt-understand-you">The missing piece: it doesn't understand you</h2>
<p>Picture a good teacher in a room with you. They watch your face and they catch the exact moment your eyes glaze over, they hear the hesitation in a question and realize you've misunderstood something three steps back, they notice when you're bored and push harder, and when you're drowning and ease off. Half of good teaching is reading the student, not reciting the material.</p>
<p>Now look at what AI has to work with. It sees your question. That's the whole of it. It can't see your confusion. It has no way of knowing you've quietly misunderstood the premise of your own question, so it will cheerfully answer the wrong question you didn't realize you were asking. The model has no real contextual awareness. It doesn't see you or your confusion. It just sees the words about your problem.</p>
<p>Which gives us the sharpest version of the point: the model doesn't know what you don't know. And in learning, that gap is the whole game.</p>
<h2 id="the-illusion-of-learning">The illusion of learning</h2>
<p>This one is the illusion of competence wearing a different hat.</p>
<p>Understanding feels instant with AI, and that's the trap, because real understanding almost never is. You get a smooth, clean answer that sounds right. What you don't get is the slower, gnarlier process of an idea taking root, the part that would leave you able to judge whether the answer is right in the first place. The model can hand you the feeling of learning with none of the substance. And because it fails convincingly, you won't notice the gap until you go to use the thing you thought you'd learned.</p>
<h2 id="the-learning-trap">The learning trap</h2>
<p>AI removes friction, but friction isn't always the enemy. The being stuck, the working through it, the hitting a wall and clawing your way around it is the process that actually builds understanding. That's the part your brain keeps.</p>
<p>AI lets you skip all of it. You get the answer before you've done any of the wrestling. You read it, you move on, faster than ever, and it feels productive as anything. But you've quietly traded the thing that makes learning stick for the feeling of actually having learned something new.</p>
<p>Here's the test I keep coming back to. You read the explanation and it makes total sense. Could you reproduce it an hour later or explain it to someone else? Can you solve the next problem without going back to ask? If reading it felt effortless but reproducing it is impossible, you didn't learn it.</p>
<h2 id="where-ai-works-well-as-a-learning-tool">Where AI works well as a learning tool</h2>
<p>So let's be concrete. AI is genuinely strong at:</p>
<ul>
<li>Explanations: taking a concept and laying it out clearly</li>
<li>Alternative perspectives: the same idea from a new angle when the first one didn't land</li>
<li>Breaking down concepts: splitting something big into pieces you can actually hold</li>
<li>Generating examples: as many worked examples as you care to ask for</li>
<li>Answering follow-up questions: the endless "but why?" chain, without ever getting impatient</li>
</ul>
<p>The pattern: AI is strong at expanding your understanding once you're already pointed in a sensible direction. It's a phenomenal amplifier for exploration.</p>
<h2 id="where-it-breaks-down">Where it breaks down</h2>
<p>And the other column. AI is weak at:</p>
<ul>
<li>Identifying your misunderstandings: it can't see the gap that you can't see either</li>
<li>Building deep intuition: it can hand you the explanation, not the wrestling that turns into instinct</li>
<li>Ensuring correctness: a confident explanation is not a verified one</li>
<li>Guiding a long-term path: it reacts to the question in front of it, with no map of where you should go next</li>
</ul>
<p>Same core issue as always. It explains, but it doesn't teach in the full sense, because teaching means knowing the student and steering the journey, and the model does neither.</p>
<h2 id="how-to-actually-use-ai-to-learn">How to actually use AI to learn</h2>
<p>None of this means don't use it. It means use it on purpose. The tool is strong, the default way most people reach for it is weak. Here's how to flip that.</p>
<p>Move from "what is it?" to "why does it work?" The first hands you a fact to copy down. The second forces the model to show its reasoning, which is the part you actually needed. Answers you can look up any time. Understanding you have to build.</p>
<h3 id="force-depth-and-variation">Force depth and variation</h3>
<p>Don't take the first explanation and run. Push it:</p>
<ul>
<li>"Explain this in three different ways."</li>
<li>"Give me a real-world analogy."</li>
<li>"What are the common mistakes people make here?"</li>
<li>"Where does this break down?"</li>
</ul>
<p>Variation is how you find out whether you actually understand the thing or just got comfortable with one particular way of phrasing it.</p>
<h3 id="use-it-as-a-reviewer-not-a-teacher">Use it as a reviewer, not a teacher</h3>
<p>Start saying "here's what I tried, what's wrong with it?"</p>
<p>Bring your own attempt: your code, your explanation, your reasoning, however rough. Let the model react to your thinking instead of replacing your thinking for you. Now you're getting feedback on something real rather than passively sitting through a lecture. And learning happens in the feedback, not the consumption.</p>
<h3 id="use-it-to-test-yourself">Use it to test yourself</h3>
<p>Flip the model from answer-machine to examiner. Ask it to generate questions on the topic. Ask for the edge cases you probably missed. Ask for the pros and cons, and let it explain. Ask it to throw a scenario at you and then check your answer against it. Discuss the answers and challenge the AI on what it tells you but instruct the AI to also do this with you. That's the shift from passive to active, from reading about the thing to being put on the spot about it. Passive feels nice. Active is where it actually sticks.</p>
<h3 id="verify-when-it-matters">Verify when it matters</h3>
<p>If correctness matters, don't trust a single explanation. The model can be fluent and wrong in the very same breath. For idle curiosity, fine, let it ride. For something you're about to build on, act on, or teach to someone else, check it against a second source. Being confident is not the same as being right.</p>
<p>So why is any of this a good idea, when the warning I keep coming back to is that AI fails convincingly? Fair question, and it's the right one to end on. The answer is that none of this asks you to trust it. The danger I keep flagging is AI as the final word, unchecked, with nobody able to catch it. Learning flips that around: you are the one becoming able to catch it. A critique is easier to check than an answer is to produce, so when it says "this is wrong," that's a lead you go and verify, not a verdict you swallow. And you keep a referee that doesn't care how confident it sounded: the compiler, the tests, the docs, a second source. Can it be trusted to overturn a correct answer? No. If you're right and it confidently says otherwise, and you can't yet tell, it will talk you out of it, and that is exactly when you know the least. So, it never gets the final word. It raises the question; you and your referee settle it.</p>
<h2 id="example-learning-to-code">Example: learning to code</h2>
<p>Let's make it concrete with coding.</p>
<p>The good kind of use, the kind that builds a developer: paste the error you're stuck on and ask why it happened. Ask for two other ways to approach the problem and what each one costs you. Have it break down the concept you keep tripping over. Every one of those leaves you more capable than you were before.</p>
<p>The bad kind, the kind that quietly hollows you out: "write the whole thing for me." It'll work. The code will run. And you'll have learned precisely nothing, because working code is not a learned skill. Do it enough and you end up with a project full of solutions you can't reproduce or debug on your own. That's not a developer getting better. That's a dependency getting deeper.</p>
<h2 id="conclusion">Conclusion</h2>
<p>AI can be a genuinely powerful learning tool. Probably the most powerful one most of us have ever had our hands on. But only if you use it correctly, and "correctly" happens to be the opposite of how it's easiest to use.</p>
<p>AI doesn't replace learning. It amplifies how you already learn. Hand it good questions, your own attempts, and a habit of checking, and it will accelerate you enormously. Hand it "just give me the answer," and it will accelerate that too, straight past the point where learning happens. That's the part worth sitting with. AI scales whatever you bring to it.</p>
<p>It’s a tool to enhance, so let the thing it's enhancing be you.</p>]]></content:encoded>
  </item>
  <item>
    <title>Azure Deployment Environments and Dev Box are in maintenance mode</title>
    <link>https://jeffops.com/posts/2026/azure-platform-engineering-maintenance-mode/</link>
    <guid isPermaLink="true">https://jeffops.com/posts/2026/azure-platform-engineering-maintenance-mode/</guid>
    <pubDate>Wed, 15 Jul 2026 22:00:00 +0000</pubDate>
    <description>The two Azure services most commonly recommended as the foundation of an internal developer platform are both in maintenance mode with no further features planned. Here is what Microsoft actually said, and what is left standing.</description>
    <category>platform-engineering</category><category>azure</category><category>idp</category><category>enterprise</category><category>ops</category>
    <content:encoded><![CDATA[<p>If you are part-way through a design that assumes either of these, this is worth knowing this week rather than next year.</p>
<p><strong>Azure Deployment Environments and Microsoft Dev Box are both in maintenance mode.</strong> Most of the "build an internal developer platform on Azure" writing currently in circulation nominates one or both as the foundation. That advice has not caught up.</p>
<p>Microsoft's <a href="https://learn.microsoft.com/azure/deployment-environments/maintenance-mode">wording for Azure Deployment Environments</a> is short enough to quote in full: "Azure Deployment Environments is in maintenance mode, with no additional features planned. Existing capabilities remain available."</p>
<p><a href="https://learn.microsoft.com/azure/dev-box/dev-box-windows-365-announcement">For Dev Box</a> it is more pointed: "Dev Box is now in maintenance mode, with no additional features planned", and "Customers should consider Windows 365 as the recommended path forward."</p>
<p>Several announced Dev Box features were cancelled before they shipped, including on-behalf-of creation, firewall service tags, developer offboarding, single sign-on for existing machines, auto-remediation and the latency improvements. There is no migration guide yet. The documentation says guidance will be provided.</p>
<p>Maintenance mode is not end of life, and nothing you are running today stops working. But it changes the calculation. A service with no further features is a service you should not be building a five-year platform strategy on top of, and it is a service whose gaps will stay gaps.</p>
<h2 id="what-has-not-gone-into-maintenance">What has not gone into maintenance</h2>
<p>The awkward part is that the two services now parked are the ones that felt most like a platform, in the sense of a thing a developer touches. What is left is less exciting and considerably more durable, because it is the substrate rather than the surface.</p>
<p><strong><a href="https://learn.microsoft.com/azure/cloud-adoption-framework/ready/landing-zone/design-area/subscription-vending">Landing zones and subscription vending</a>.</strong> This is the genuinely platform-shaped primitive in Azure: a request produces a subscription placed in the right management group, with networking, RBAC and identity already wired. It is the paved road, expressed as infrastructure rather than as a portal.</p>
<p>Read Microsoft's own framing carefully, because it cuts against the way golden paths are usually described. It is not "pick one blessed route and make everyone take it". Microsoft says that organisations "that only provide a 'one size fits all' approach to subscription vending often limit their internal customers' flexibility" and that "platform teams need to provide various product lines to cater to their organization's needs", noting that customers typically start with three: sandbox, corp connected, and online.</p>
<p>A small catalogue of paved roads, not one. That is a more useful model than the singular golden path, and it is the difference between a platform people route around and one they use.</p>
<p><strong><a href="https://azure.github.io/Azure-Verified-Modules/">Azure Verified Modules</a></strong> as the module layer. Read the support statement before you depend on it. The commitment is a meaningful response toward resolution within five business days, and Microsoft is explicit that this is "not for a fix within these durations". Modules also carry lifecycle states, including "orphaned" for when maintainers stop responding. That is fine, and it is a great deal better than writing your own modules from scratch. It is not the same as a supported product, and the difference matters when you are telling your organisation that this is the supported way to build.</p>
<p><strong>Azure Policy</strong> as the governance layer. Microsoft's framing here is the useful bit: <a href="https://learn.microsoft.com/platform-engineering/application-platform">start right and stay right</a>, meaning catalogued templates get you compliant at creation and policy keeps you compliant afterwards. Note the warning that comes attached, because it is the thing that will bite you: while policies "can enforce compliance, they can also break applications unexpectedly". Roll them out in rings, the same way you would roll out anything else that can take production down.</p>
<h2 id="on-portals-microsofts-advice-runs-against-the-instinct">On portals, Microsoft's advice runs against the instinct</h2>
<p>Worth quoting because it is the opposite of where most platform projects start. A developer portal, Microsoft says, is <a href="https://learn.microsoft.com/platform-engineering/developer-self-service">"a destination rather than a starting point"</a>. Before building somewhere new for people to go, extend the surfaces they already use: the IDE, the CLI, the DevOps tooling, chat.</p>
<p>That ordering is well supported by evidence from outside Microsoft. Spotify, who built the most popular developer portal there is, <a href="https://www.techtarget.com/searchitoperations/news/366558592/Behind-the-scenes-Spotify-Backstage-a-work-in-progress">put average Backstage adoption inside adopting organisations at around 10%</a>. Building a second front door and then asking people to learn it is a hard way to earn attention you could have had for free.</p>
<h2 id="one-caveat-if-you-standardise-on-azd-templates">One caveat if you standardise on azd templates</h2>
<p>Azure Developer CLI templates are a genuinely good accelerator, and Microsoft states plainly that the templates, "including those provided from Microsoft, are not supported by any Microsoft support program or service". They are provided as is.</p>
<p>That is fine for a starting point and a problem for a paved road. The moment you tell your organisation "this is the supported way to deploy", you have quietly assumed the support obligation yourself. Decide that deliberately rather than discovering it during an incident.</p>
<h2 id="what-i-would-take-from-this">What I would take from this</h2>
<p>The services that went into maintenance are the ones that looked like a product. The ones still standing are the ones that look like plumbing. That is not a coincidence, and there is a lesson in it about which layer is worth building your platform on.</p>
<p>More generally, it is a small live demonstration of a bigger point: <a href="/posts/2026/internal-developer-platform-expiry-date/">an internal platform is a bet on a gap, and the gap moves</a>. Sometimes it moves because the hyperscalers close it. Sometimes, as here, it moves because the hyperscaler you built on quietly stops developing the thing you built on. Either way, the assumption is worth re-pricing on a schedule rather than at the point where someone sends you a documentation link.</p>]]></content:encoded>
  </item>
  <item>
    <title>Why AI Works So Well for Coding, and Where It Actually Fits</title>
    <link>https://jeffops.com/newsletter/why-ai-works-for-coding/</link>
    <guid isPermaLink="true">https://jeffops.com/newsletter/why-ai-works-for-coding/</guid>
    <pubDate>Mon, 06 Jul 2026 09:00:00 +0000</pubDate>
    <description>Coding is not the exception to AI&#x27;s limits. It is the one place those limits hurt least, because code tells you loudly when it is wrong.</description>
    <category>Newsletter</category>
    <content:encoded><![CDATA[<p>In my previous newsletter edition I made a fairly blunt argument: AI doesn't think, it predicts. It samples plausible-sounding text, it's inconsistent by nature, and it fails convincingly rather than loudly. That's why I don't trust it with legal interpretation, risk decisions, or anything where judgment matters more than pattern-matching.</p>
<p>So here's the obvious question I promised to answer. If all of that is true, why does coding feel like the exception?</p>
<p>Because it kind of is. And the reason isn't the one most people assume.</p>
<h2 id="coding-feels-different-and-theres-a-reason">Coding feels different, and there's a reason</h2>
<p>Ask AI to write you a marketing strategy and you get a confident answer you can't fully trust. Ask it to write a function that parses a date string, and… it usually just works. Same model, same token-prediction engine underneath, and yet the reliability is night and day.</p>
<p>The easy conclusion is "coding must be the thing AI is genuinely good at." I think that's backwards. AI isn't good at coding because it understands code. It's good at coding because coding happens to fit, almost perfectly, the way these models actually work. Every one of the limitations I keep banging on about is still there. Coding is just the one place where those limitations hurt the least.</p>
<p>Let me walk through why.</p>
<h3 id="code-is-predictable-by-nature">Code is predictable by nature</h3>
<p>Human language is ambiguous. A single sentence can carry three different meanings depending on tone, context, and who's reading it. Code doesn't get to do that. It has strict syntax, defined keywords, and rules that don't bend. <code>if</code> means <code>if</code>. A missing semicolon isn't a stylistic choice you get to defend in review.</p>
<p>The model, remember, is a next-token predictor. The less ambiguous the "next token" is, the better it does. Language hands it a wide, fuzzy field of plausible options. Code narrows that field hard, because a lot of the time there's genuinely only one correct next token. You've basically handed the prediction engine the easiest version of its own job.</p>
<h3 id="the-entire-industry-became-training-data">The entire industry became training data</h3>
<p>Think about what these models were trained on. Decades of open source. Every framework's documentation. Millions of Stack Overflow answers, GitHub more or less in its entirety, plus the endless tutorials, blog posts, code reviews and bug reports on top.</p>
<p>Software engineering might be the most thoroughly documented human activity on the internet. And it's documented in exactly the shape the model learns best from: here's a problem, here's the working solution, here's why it works. We didn't just give it examples. We handed it the answer key.</p>
<h3 id="most-code-is-variations-of-the-same-patterns">Most code is variations of the same patterns</h3>
<p>Here's the slightly uncomfortable truth most of us already know. We don't write that much genuinely original code. We wire up an API client. We loop over a collection. We validate some input. We map one shape of data onto another. Every one of us has written the same try/catch a few hundred times.</p>
<p>That's not an insult, by the way. Reuse and composition are good engineering. But it does mean most of what we produce is a variation on a pattern that already exists a million times over in the training data. And "give me a variation on a well-established pattern" is about the most on-target request you can hand one of these models.</p>
<h3 id="code-tells-you-when-its-wrong">Code tells you when it's wrong</h3>
<p>This is the big one. Honestly, if you take one thing from this piece, take this.</p>
<p>The real danger with AI is that it fails silently. A wrong business decision looks exactly like a right one: polished, confident, no red flags anywhere. There's no stack trace for a bad marketing call. You find out it was wrong three months later, if you ever find out at all.</p>
<p>Code is the opposite. Code has a feedback loop baked into it. The compiler rejects it. The linter starts complaining. The test suite goes red. The thing falls over at runtime and hands you an actual error message. The illusion of competence, which is the model's most dangerous trait everywhere else, gets punctured almost immediately here, because the code has to actually run.</p>
<p>For once the failures are loud. Something red, something you can see. And that one property changes almost everything about how safely you can lean on it.</p>
<h3 id="and-thats-how-it-kept-getting-better">And that's how it kept getting better</h3>
<p>That same property, code being checkable, did something bigger than help you catch mistakes. It's a large part of why these models improved at coding faster than at almost anything else they do.</p>
<p>To make a model better at something, you have to be able to tell it when it got something right and when it got it wrong. For most work that's hard and subjective. What's a "good" marketing email, or a "good" legal argument? Ask ten experts and you'll get ten answers, none of which compile. Code is different. Does it run? Do the tests pass? Those are yes/no questions a machine can grade millions of times over without tiring or having an opinion. So the model could be trained against a real answer key: reward the code that works, penalise the code that doesn't, and repeat at a scale no human review could ever touch. It was effectively taught what "good" and "bad" code look like by something that could check every single time, and that objective, automatic grader is a big reason coding ability raced ahead while the fuzzier, unverifiable work stayed stuck.</p>
<h3 id="almost-correct-is-still-extremely-useful">"Almost correct" is still extremely useful</h3>
<p>With something like legal work, being wrong 10% of the time is a liability nightmare at scale. Twenty wrong words in a 200-word paragraph, and any one of them could be the expensive one.</p>
<p>Coding flips that math on its head. If AI gets you 70% of the way to a working function, that isn't a 30% failure. It's a 70% head start. You take the scaffolding, spot the gaps (the compiler is right there helping you spot them), and finish the job. "Almost correct" in a contract is dangerous. "Almost correct" in code is just a normal Tuesday, and most days it's genuinely useful.</p>
<h3 id="most-problems-are-smaller-than-they-look">Most problems are smaller than they look</h3>
<p>A lot of programming, when you actually watch yourself do it, is a long series of small, local, self-contained problems. Parse this. Transform that. Sort the list. Format the output. These are the bite-sized, well-bounded tasks the model is good at, precisely because none of them ask it to hold the whole system in its head at once. The trouble starts when the problem isn't local. Hold that thought, because that's where this whole thing turns.</p>
<h2 id="what-this-really-changes-from-writing-code-to-evaluating-it">What this really changes: from writing code to evaluating it</h2>
<p>Put all of that together and something quietly shifts under your feet. The value isn't in writing the code anymore. It's in knowing whether the code is right. The model can produce the function. It can produce ten versions of the function before you've finished your coffee. What it can't do is tell you which one belongs in your system, handles your edge cases, and won't quietly fall over in production six weeks from now. That judgment is the actual job now.</p>
<h2 id="who-this-actually-works-well-for">Who this actually works well for</h2>
<p>The same tool lands completely differently depending on whose hands it's in.</p>
<h3 id="experienced-developers-faster-not-replaced">Experienced developers: faster, not replaced</h3>
<p>If you already know what good looks like, AI is a force multiplier. You read its output the way you'd read a junior's pull request. You catch the mistakes, refine it, wire it in properly, and move on. You're not trusting it, you're supervising it. And because you can actually evaluate the result, you get the speed without inheriting the risk.</p>
<h3 id="generalists-filling-the-gaps">Generalists: filling the gaps</h3>
<p>If you work across a lot of tools and languages, AI smooths over the friction. Less time lost to unfamiliar syntax, less time digging for the magic incantation in a language you touch twice a year. It won't make you an expert in everything, but it clears out a lot of the small stumbles that used to slow you to a crawl.</p>
<h3 id="prototypers-from-idea-to-code-instantly">Prototypers: from idea to code instantly</h3>
<p>If your goal is to get from idea to working demo as fast as humanly possible, this is where AI really shines. Iteration speed goes through the roof. You can try five approaches in the time it used to take to wire up one. And for a prototype that's exactly the right trade, because the goal is to learn something, not to ship something bulletproof.</p>
<h3 id="beginners-learning-tool-or-crutch">Beginners: learning tool or crutch</h3>
<p>Here's where I get cautious. For a beginner, AI is genuinely powerful and quietly dangerous at the same time. It'll hand you working code before you have any idea why it works. Point it at "why did that actually fix it?" and it's a phenomenal tutor. Use it to skip the understanding entirely and you're building a dependency on a tool that fails convincingly. You end up unable to tell when it's wrong, which is the one skill that turns out to matter most.</p>
<h3 id="specialists-where-the-model-runs-out-of-data">Specialists: where the model runs out of data</h3>
<p>Out at the deep end, the model starts to thin out. Novel problems, deep systems work, the genuinely weird edge cases: the places where there isn't much training data, because not many people have solved this before. That's exactly where pattern-matching has nothing to match against, and the confident-but-wrong behaviour comes creeping back. The further you get from the well-trodden path, the less the model has to offer you.</p>
<h2 id="what-actually-matters-now">What actually matters now</h2>
<p>If writing code isn't the bottleneck anymore, the questions worth asking are about you, not the tool. Four of them.</p>
<p><strong>Can you describe the problem clearly?</strong> The model is only ever as good as the problem you hand it. Vague in, vague out. Thinking clearly about what you actually need is now half the work.</p>
<p><strong>Can you tell when something is wrong?</strong> This is the successor to the illusion of competence. The output looks finished. Whether it is finished is your call, and you can only make that call if you know what wrong looks like.</p>
<p><strong>Can you tell a good solution from a bad one?</strong> Working isn't the same as good. Plenty of AI-generated code runs perfectly and is still a maintenance liability waiting to happen. Spotting that difference is judgment, not pattern-matching.</p>
<p><strong>Do you understand the system beyond the code?</strong> The model sees the snippet. It doesn't see your architecture, your constraints, your history, or the three other services this quietly touches. All of that context lives with you.</p>
<p>None of those are really coding skills. They're engineering skills. The typing got automated. The thinking didn't.</p>
<h2 id="where-ai-works-well-in-coding">Where AI works well in coding</h2>
<p>Let's get concrete. AI is genuinely strong at:</p>
<ul>
<li>Boilerplate and scaffolding: the repetitive setup nobody enjoys writing</li>
<li>Writing tests: structured, pattern-heavy, and easy to verify</li>
<li>Refactoring: reshaping code that already works into something cleaner</li>
<li>Explaining code: a fast way into an unfamiliar codebase</li>
<li>Translating between languages: porting a known solution from one syntax to another</li>
</ul>
<p>See the shape of it? Structured, repeatable, verifiable. It's the same profile as checking a contract against a fixed checklist. You're asking it to work inside known patterns rather than invent something new, and you can check the result when it's done.</p>
<h2 id="where-it-breaks-down">Where it breaks down</h2>
<p>And here's the other column. AI struggles badly with:</p>
<ul>
<li>Architecture decisions: trade-offs that hinge on where the business is heading</li>
<li>Cross-system reasoning: problems that span services, teams, and boundaries</li>
<li>Long-term maintainability: choices whose real cost only shows up months later</li>
<li>Edge cases: the rare, weird, high-stakes paths that judgment exists for in the first place</li>
</ul>
<p>Same core issue every time: the moment the problem stops being local and turns into judgment, the model runs out of road.</p>
<h2 id="the-same-problem-just-in-a-different-form">The same problem, just in a different form</h2>
<p>All of the above leads to the following. This is the part that I really want to land.</p>
<p>Nothing about the model actually changed between "don't use it for legal" and "it's great for coding." It's the same randomness, the same illusion of competence, the same engine sampling plausible tokens with no idea whether any of them are true. Coding didn't fix a single bit of that. It just happens to come with guardrails that business decisions don't have: strict syntax, a mountain of training data, and above all a feedback loop that makes the failures loud instead of silent. Strip those guardrails away, push it toward architecture and edge cases and cross-system judgment, and the exact same problems come marching straight back.</p>
<p>The reliability was never in the model. It was in the environment around it.</p>
<h2 id="conclusion">Conclusion</h2>
<p>So use it. For scaffolding, tests, refactoring, translation, and getting yourself unstuck, it's a real multiplier. Just keep doing the part that was always the actual work: deciding what to build, and knowing whether what came back is any good.</p>
<p>A tool to enhance, not to replace. Same as it ever was. Just this time with a stack trace.</p>]]></content:encoded>
  </item>
  <item>
    <title>Your internal developer platform has an expiry date</title>
    <link>https://jeffops.com/posts/2026/internal-developer-platform-expiry-date/</link>
    <guid isPermaLink="true">https://jeffops.com/posts/2026/internal-developer-platform-expiry-date/</guid>
    <pubDate>Thu, 25 Jun 2026 22:00:00 +0000</pubDate>
    <description>The best documented internal platform in the public sector hit every operational target and was shut down anyway. A platform is not an asset you build, it is a bet that a gap is worth closing, and the gap closes without telling you.</description>
    <category>platform-engineering</category><category>idp</category><category>enterprise</category><category>adoption</category><category>ops</category>
    <content:encoded><![CDATA[<p>172 digital services. More than 60 departments, agencies and local authorities. 3,200 applications. Deployments more than 122 times a day. 99.95% uptime.</p>
<p>That is <a href="https://gds.blog.gov.uk/2022/07/12/why-weve-decided-to-decommission-gov-uk-paas-platform-as-a-service/">GOV.UK PaaS in July 2022</a>, on the day the Government Digital Service announced it was shutting it down.</p>
<p>By every operational measure in the platform engineering canon, it worked. It was decommissioned anyway, and the reasons GDS gave are the most useful thing published on this subject.</p>
<p>Here is the line the rest of this hangs on: <strong>a platform is not an asset you build, it is a bet that a particular gap is worth closing, and the gap closes without telling you.</strong></p>
<p>A word on where this comes from, because it should change how you read it. I run workplace technology for 25,000 people, which is enterprise IT rather than product engineering. I have never stood up a Backstage instance. What follows is not a war story. It is what the published evidence says when you go and read it, and the evidence turns out to be stranger than the advice built on top of it.</p>
<h2 id="what-actually-closed-govuk-paas">What actually closed GOV.UK PaaS</h2>
<p>GDS gave three reasons. Growth had stalled. The large cloud providers had, in their words, "upped their game and reduced the barriers to entry". And departments had built their own cloud engineering capability and were clustering around a Kubernetes based architecture on their own.</p>
<p>The tempting move is to pick whichever of those supports the argument you already hold. Read the first on its own and this is a straightforward adoption failure. Read the second and third on their own and it is pure market movement, nobody's fault.</p>
<p>They are the same event seen from two ends.</p>
<p>Growth stalls <em>because</em> the gap is closing. A team that can now stand up its own thing in an afternoon does not join your platform, and the number that surfaces on your dashboard is flat adoption. So adoption is the symptom you can see, and scarcity is the disease underneath it. That distinction is not academic, because it changes what you monitor. Adoption tells you how you did last quarter. Scarcity tells you how long you have got.</p>
<p>Every internal platform is an arbitrage on the difference between what your organisation can do on its own and what the market makes easy. That difference is the whole value of the thing. It narrows every year, and nobody sends you a note when it does.</p>
<p><div class="figure-scroll" tabindex="0" role="group" aria-label="Diagram, scrollable"><svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 780 440" width="780" height="440" role="img" aria-labelledby="gp-t gp-d" style="max-width:100%;height:auto;display:block;margin:1.5rem 0"><title id="gp-t">Why an internal platform has an expiry date</title><desc id="gp-d">Two schematic curves. The effort of doing it yourself falls steadily as the market improves, while the effort of doing it through your platform stays roughly flat. The shaded gap between them is the value the platform provides, and it closes on its own.</desc><defs><marker id="gp-a" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"><path d="M0,0 L10,5 L0,10 z" fill="#0aa5c4"/></marker><marker id="gp-aw" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"><path d="M0,0 L10,5 L0,10 z" fill="#d2593c"/></marker></defs><line x1="110" y1="46" x2="110" y2="340" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><line x1="110" y1="340" x2="714" y2="340" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="100" y="50" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11" fill="currentColor" fill-opacity="0.62" text-anchor="end" font-weight="400">more effort</text><text x="100" y="336" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11" fill="currentColor" fill-opacity="0.62" text-anchor="end" font-weight="400">less</text><text x="714" y="358" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11" fill="currentColor" fill-opacity="0.62" text-anchor="end" font-weight="400">time</text><polygon points="110.0,88.0 119.7,97.8 129.3,104.5 139.0,110.4 148.7,115.8 158.3,120.9 168.0,125.7 177.7,130.3 187.3,134.8 197.0,139.1 206.7,143.3 216.3,147.4 226.0,151.4 235.7,155.3 245.3,159.2 255.0,163.0 264.7,166.7 274.3,170.3 284.0,173.9 293.7,177.5 303.3,181.0 313.0,184.5 322.7,187.9 332.3,191.3 342.0,194.6 351.7,197.9 361.3,201.2 371.0,204.5 380.7,207.7 390.3,210.9 400.0,214.1 409.7,217.2 419.3,220.3 429.0,223.4 438.7,226.5 448.3,229.5 458.0,232.5 467.7,235.5 477.3,238.5 487.0,241.5 496.7,244.4 496.7,242.0 487.0,241.8 477.3,241.6 467.7,241.4 458.0,241.2 448.3,241.0 438.7,240.8 429.0,240.6 419.3,240.4 409.7,240.2 400.0,240.0 390.3,239.8 380.7,239.6 371.0,239.4 361.3,239.2 351.7,239.0 342.0,238.8 332.3,238.6 322.7,238.4 313.0,238.2 303.3,238.0 293.7,237.8 284.0,237.6 274.3,237.4 264.7,237.2 255.0,237.0 245.3,236.8 235.7,236.6 226.0,236.4 216.3,236.2 206.7,236.0 197.0,235.8 187.3,235.6 177.7,235.4 168.0,235.2 158.3,235.0 148.7,234.8 139.0,234.6 129.3,234.4 119.7,234.2 110.0,234.0" fill="#0aa5c4" fill-opacity="0.14" stroke="none"/><polyline points="110.0,234.0 119.7,234.2 129.3,234.4 139.0,234.6 148.7,234.8 158.3,235.0 168.0,235.2 177.7,235.4 187.3,235.6 197.0,235.8 206.7,236.0 216.3,236.2 226.0,236.4 235.7,236.6 245.3,236.8 255.0,237.0 264.7,237.2 274.3,237.4 284.0,237.6 293.7,237.8 303.3,238.0 313.0,238.2 322.7,238.4 332.3,238.6 342.0,238.8 351.7,239.0 361.3,239.2 371.0,239.4 380.7,239.6 390.3,239.8 400.0,240.0 409.7,240.2 419.3,240.4 429.0,240.6 438.7,240.8 448.3,241.0 458.0,241.2 467.7,241.4 477.3,241.6 487.0,241.8 496.7,242.0 506.3,242.2 516.0,242.4 525.7,242.6 535.3,242.8 545.0,243.0 554.7,243.2 564.3,243.4 574.0,243.6 583.7,243.8 593.3,244.0 603.0,244.2 612.7,244.4 622.3,244.6 632.0,244.8 641.7,245.0 651.3,245.2 661.0,245.4 670.7,245.6 680.3,245.8 690.0,246.0" fill="none" stroke="#0aa5c4" stroke-width="2"/><polyline points="110.0,88.0 119.7,97.8 129.3,104.5 139.0,110.4 148.7,115.8 158.3,120.9 168.0,125.7 177.7,130.3 187.3,134.8 197.0,139.1 206.7,143.3 216.3,147.4 226.0,151.4 235.7,155.3 245.3,159.2 255.0,163.0 264.7,166.7 274.3,170.3 284.0,173.9 293.7,177.5 303.3,181.0 313.0,184.5 322.7,187.9 332.3,191.3 342.0,194.6 351.7,197.9 361.3,201.2 371.0,204.5 380.7,207.7 390.3,210.9 400.0,214.1 409.7,217.2 419.3,220.3 429.0,223.4 438.7,226.5 448.3,229.5 458.0,232.5 467.7,235.5 477.3,238.5 487.0,241.5 496.7,244.4 506.3,247.3 516.0,250.2 525.7,253.1 535.3,256.0 545.0,258.9 554.7,261.7 564.3,264.5 574.0,267.3 583.7,270.1 593.3,272.9 603.0,275.7 612.7,278.4 622.3,281.2 632.0,283.9 641.7,286.6 651.3,289.3 661.0,292.0 670.7,294.7 680.3,297.3 690.0,300.0" fill="none" stroke="#d2593c" stroke-width="2" stroke-dasharray="6 4"/><text x="128" y="66" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="13.5" fill="#d2593c" fill-opacity="1" text-anchor="start" font-weight="600">Doing it yourself</text><text x="128" y="84" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="11.5" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="400">falls every quarter, without anyone telling you</text><text x="128" y="272" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="13.5" fill="#0aa5c4" fill-opacity="1" text-anchor="start" font-weight="600">Doing it through your platform</text><text x="128" y="290" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="11.5" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="400">roughly flat, and you pay to keep it there</text><text x="240" y="200" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="12.5" fill="currentColor" fill-opacity="1" text-anchor="middle" font-weight="600">the gap you are arbitraging</text><text x="240" y="216" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="10.5" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">when this closes, so does the platform</text><circle cx="496.7" cy="242.0" r="5.5" fill="none" stroke="#d2593c" stroke-width="2"/><line x1="496.7" y1="231.0" x2="496.7" y2="132" stroke="#d2593c" stroke-width="1.5" stroke-dasharray="3 3"/><text x="508.66666666666663" y="124" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="12.5" fill="#d2593c" fill-opacity="1" text-anchor="start" font-weight="600">the abstraction stops being scarce</text><text x="508.66666666666663" y="142" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="11" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="400">nothing failed; the market moved</text><text x="110" y="406" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="12" fill="currentColor" fill-opacity="0.75" text-anchor="start" font-weight="400">GOV.UK PaaS reached this point with 172 services, 3,200 apps and 99.95% uptime, and was retired anyway.</text><text x="110" y="426" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.6" text-anchor="start" font-weight="400">Schematic. The shape is the argument; neither curve is measured.</text></svg></div></p>
<p>One detail worth sitting with. The UK government's <a href="https://roadmap-for-modern-digital-government.campaign.gov.uk/digital-and-data-infrastructure/common-systems-platforms">roadmap for modern digital government</a> commits to a fresh generation of shared platforms, including a new one called internal.gov.uk "for people who work in the public sector to connect via instant messaging functionality, forums and a secure code repository", due June 2026. It makes no reference to GOV.UK PaaS, and gives no account of why the last one was retired.</p>
<p>The organisation that wrote the best post-mortem in this field did not cite it in its own next plan.</p>
<h2 id="the-literature-was-not-written-for-you">The literature was not written for you</h2>
<p>If you run an enterprise IT estate rather than a product engineering organisation, this is the part that matters.</p>
<p>Look at who the examples are. Team Topologies publishes <a href="https://teamtopologies.com/industry-examples/tag/Platform+Teams">thirteen platform team industry examples</a>. Exactly one is public sector. The rest are Docker, Flo Health, Trade Me, Improbable, PureGym, Footasylum, Uswitch and similar. Every prescription in the field, golden paths, thinnest viable platform, stream-aligned teams, cognitive load, assumes a population of product engineers shipping services they own.</p>
<p>That is not most enterprise IT. If your estate is Microsoft 365, a service management tool, a fleet of endpoints, a dozen line-of-business applications and a handful of integrations, your users are not stream-aligned product teams. Your cognitive load problem is not deployment pipelines. And your adoption problem is not that developers prefer their own tooling, it is that a department will buy a SaaS product on a corporate card rather than wait for you.</p>
<p>The strange part is that the canonical definition already covers you. <a href="https://tag-app-delivery.cncf.io/whitepapers/platforms/">CNCF's platforms white paper</a> says platform users "include but aren't limited to app developers and operators, data scientists, COTS software operators, and information workers". Information workers are in the founding definition. Almost nothing built on top of that definition acknowledges they exist.</p>
<p>The one author writing directly for this audience is Gregor Hohpe, whose 2024 book <a href="https://architectelevator.com/book/platformstrategy/"><em>Platform Strategy</em></a> has a substantial in-house IT section, including one titled "IT Platform and IT Services Are Antonyms". His failure mode for enterprise platforms will be familiar: outdated by the time they launch, restricting rather than enabling, and mandated in a last-ditch attempt to make the economics work. His evidence base is a decade of practitioner experience rather than measurement, and he names no organisations.</p>
<h2 id="and-the-evidence-under-the-advice-is-thin">And the evidence under the advice is thin</h2>
<p>Worth knowing before you cite any of it at your own leadership.</p>
<p>Start with the number everyone repeats, some version of "70% of platform initiatives fail". <a href="https://thenewstack.io/why-up-to-70-of-platform-engineering-teams-fail-to-deliver-impact/">The New Stack repeats it</a> and, to its credit, links its sources, which is how you can check them. Follow the first link and the "up to 70%" traces to a webinar registration page titled "Platform leadership 101: Why 67% of platform initiatives are failing", whose entire stated methodology is a single sentence: "Both our community surveys and broader market studies reveal that 67% of platform leaders are missing the mark on building successful, scalable Internal Developer Platforms." Follow the second, for the claim that half of platform teams are disbanded within eighteen months, and you reach a <a href="https://www.gartner.com/en/infrastructure-and-it-operations-leaders/topics/platform-engineering">Gartner topic page</a> that does not contain that claim at all. Its only statistic is a forecast that 80% of large software engineering organisations will have platform teams by 2026.</p>
<p>So the citation is not thin. It is broken.</p>
<p>The academic position is documented, and it is worse. A <a href="https://www.frontiersin.org/journals/computer-science/articles/10.3389/fcomp.2026.1814498/full">multivocal literature review published in Frontiers in Computer Science</a> in May 2026 searched five databases and found "fewer than a dozen peer-reviewed papers from reputable venues that address platform engineering directly". No tier-A source proposes or validates a maturity model, so every maturity model in circulation comes from grey literature. And on scorecards, the main governance mechanism inside most developer portals, it found "no peer-reviewed empirical evidence of effectiveness exists despite widespread commercial adoption".</p>
<p>Be fair to the method here: a multivocal review deliberately includes grey literature, so a low proportion of academic sources is partly the point of the technique. The absolute count is the number that bites. Fewer than a dozen papers, for a discipline with a conference circuit and a tooling market.</p>
<p>The canonical artefacts have not aged well either. The <a href="https://tag-app-delivery.cncf.io/whitepapers/platform-eng-maturity-model/">CNCF Platform Engineering Maturity Model</a> was released in November 2023 and is still v1.0. The working group that owned it no longer exists: when <a href="https://www.cncf.io/blog/2025/05/07/10-years-in-cloud-native-toc-restructures-technical-groups/">CNCF restructured its technical groups in May 2025</a>, the assessment work was recorded in CNCF's own tracker as <a href="https://github.com/cncf/toc/issues/1679">"having paused in recent months due to recent TOC changes"</a>, with the authors asking that the TOC "provide guidance as to the correct place within the community to conduct this ongoing work". A <a href="https://cloud-native-platform-engineering.github.io/pemm-assessment/">pilot assessment tool</a> has since appeared outside the CNCF repositories, still asking for feedback.</p>
<p>None of that means platform engineering is wrong. It means the confident numbers in your inbox are not measurements, and you should discount accordingly, including the ones in this piece.</p>
<h2 id="what-is-actually-measured-including-the-uncomfortable-part">What is actually measured, including the uncomfortable part</h2>
<p><a href="https://dora.dev/research/2024/dora-report/">DORA's 2024 report</a>, nearly 3,000 respondents, is the strongest measured evidence in the field. Internal developer platform users showed 8% higher individual productivity and 10% higher team performance.</p>
<p>Throughput also decreased by 8%, and change stability decreased by 14%.</p>
<p>That is the field's flagship research reporting that platforms made delivery slower and less stable. DORA's own hypotheses are that a platform can add handoffs between systems and teams, or that it lets teams ship faster and therefore break more, or that instability signals teams were never properly onboarded. The report links that instability to increased developer burnout, and cautions against deploying platform engineering as a burnout remedy.</p>
<p><a href="https://dora.dev/capabilities/platform-engineering/">DORA's 2025 report</a>, <a href="https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report">nearly 5,000 technology professionals</a>, adds the finding that matters most right now: when platform quality is high, the effect of AI adoption on organisational performance is "strong and positive", and when platform quality is low that effect is "negligible". If you are being asked to demonstrate returns on AI, the platform is not a competing priority. It is the precondition. That is the strongest argument for doing this work I have found anywhere, and note that it is conditional on quality rather than on existence.</p>
<p>One discrepancy worth carrying. DORA 2025 reports 90% of organisations have adopted at least one platform and 76% have dedicated platform teams. A <a href="https://www.cncf.io/announcements/2026/03/24/cncf-and-slashdata-report-finds-platform-engineering-tools-maturing-as-organizations-prepare-for-ai-driven-infrastructure/">CNCF and SlashData survey</a> of more than 400 developers in March 2026 found only 28% have a dedicated platform engineering team. Those two are not describing the same thing. When someone quotes an adoption figure at you, ask what they counted.</p>
<h2 id="nobody-uses-these-things">Nobody uses these things</h2>
<p>Spotify, the company that created Backstage, has <a href="https://www.techtarget.com/searchitoperations/news/366558592/Behind-the-scenes-Spotify-Backstage-a-work-in-progress">publicly estimated the average Backstage adoption rate at 10%</a> inside the organisations that install it. Nine in ten of the people it was built for carry on as before.</p>
<p>The vendor surveys agree, which is notable given who pays for them. <a href="https://www.vmware.com/docs/vmw-weave-state-of-platform-engineering-report-vol-4-2025">State of Platform Engineering Volume 4</a>, published by Weave Intelligence with 518 respondents, found that "nearly 30 percent of platform teams still report that they do not measure success at all", and that adoption "is still frequently driven by extrinsic push or mandate (36.6%) rather than the intrinsic value necessary for true user pull (28.2%)".</p>
<p>Which returns you to where the idea started. When Evan Bottcher <a href="https://martinfowler.com/articles/talk-about-platforms.html">defined the digital platform</a> in 2018, one word was doing real work: "a foundation of self-service APIs, tools, services, knowledge and support which are arranged as a <strong>compelling</strong> internal product". Compelling, not mandated. A platform that needs a mandate has already failed its own definition. Even Microsoft's guidance, which has every commercial reason to encourage you, <a href="https://learn.microsoft.com/platform-engineering/about/product-mindset">says it plainly</a>: "Customers should want to use your platform, but not be mandated to use it."</p>
<p>That advice is correct, and in a regulated enterprise it is unusable as written. Which brings me to the part the literature does not cover.</p>
<h2 id="three-things-i-would-do-differently">Three things I would do differently</h2>
<p><strong>Mandate the control, never the path.</strong> In a bank, an insurer, a hospital or a government department, "don't mandate it" is not available to you, because an unpaved road is an audit finding. So mandate the outcome instead: encryption, identity, logging, network placement, data residency. Leave the route optional. This costs you nothing you actually needed, and it buys you the only honest adoption metric there is, because now the percentage of compliant workloads that came through your paved road tells you how many teams chose you when a legal alternative existed. Compare that with measuring how many workloads are compliant, which measures your auditors rather than your platform.</p>
<p><strong>Your service catalogue is the portal you already have.</strong> Microsoft's own advice is that a developer portal is <a href="https://learn.microsoft.com/platform-engineering/developer-self-service">"a destination rather than a starting point"</a>, and that you should extend the surfaces people already use before you build a new one. In enterprise IT that surface is the service management tool, because it is where people already go when they want something. Self-service in this world usually means a request with an approval attached, not a git push, and there is no shame in that. Building a second front door and then asking the organisation to learn it is how you arrive at 10%.</p>
<p><strong>Put the expiry review in the diary, and decide now who raises it.</strong> Every six months, write down what your platform makes easy, then price what it would cost a team to do that today without you, using whatever the hyperscalers shipped last quarter. If the answer is "about the same", you are running GOV.UK PaaS in 2022. The second half of that sentence is the hard one. Nobody retires a platform without embarrassment, because the person best placed to notice the gap has closed is the person whose job closes with it. GDS could do it because GDS had plenty of other work. A platform team inside a bank has no such luxury, and that, rather than sunk cost, is the real reason platforms outlive their scarcity. Name the person who is expected to raise it while naming them is still cheap.</p>
<p>Underneath all three sits the question nobody writes about: who funds this. In a product organisation the platform team is paid for out of engineering and nobody asks again. In enterprise IT it is funded by whoever will sign for it, and an unmandated central platform re-justifies its budget every year against projects with a named business sponsor, and loses. The person who signs is also the person who will be asked at the next audit why the controls are not uniform. Which is the first recommendation again, arriving from the other direction.</p>
<h2 id="the-two-things-that-are-both-true">The two things that are both true</h2>
<p>The discipline's own flagship research says platforms reduced throughput and stability. Its most-quoted failure statistic traces to a webinar title and a Gartner page that does not contain it. The company that built the most popular portal puts adoption at one in ten.</p>
<p>And the platform is still the thing that determines whether your AI spending returns anything at all.</p>
<p>You hold those together by changing what a platform is for. Not an asset you build and then defend, but a bet on a gap, priced at the start, repriced on a schedule, and closed without embarrassment when the market closes it for you. GOV.UK PaaS was not a failure. It was a correct decision at the beginning and a correct decision at the end, and the only mistake available was keeping it because shutting it down would have looked like one.</p>
<p>Start by writing down which gap you are betting on. If you cannot name it in a sentence, you do not have a platform, you have a project.</p>
<p><em>If you are building on Azure, two of the services most commonly recommended as the foundation went into maintenance mode this year. That is a separate post: <a href="/posts/2026/azure-platform-engineering-maintenance-mode/">Azure Deployment Environments and Dev Box are in maintenance mode</a>.</em></p>]]></content:encoded>
  </item>
  <item>
    <title>AI - Understand before you apply (and buy)</title>
    <link>https://jeffops.com/newsletter/understand-before-you-apply/</link>
    <guid isPermaLink="true">https://jeffops.com/newsletter/understand-before-you-apply/</guid>
    <pubDate>Mon, 22 Jun 2026 09:00:00 +0000</pubDate>
    <description>AI doesn&#x27;t think, it predicts. Why that makes it superb at some tasks, a liability at others, and how to tell which is which before you buy.</description>
    <category>Newsletter</category>
    <content:encoded><![CDATA[<p>AI is a wonderful tool. The speed at which it can perform certain tasks is just downright amazing. But it’s a tool, nothing more than that. It’s not a way to replace your workforce. Let me explain.</p>
<h2 id="how-ai-models-work">How AI models work</h2>
<p>It essentially comes down to the very basics of how AI models work. AI doesn’t think, it predicts. Instead of thinking:</p>
<blockquote>
<p>This is the correct answer</p>
</blockquote>
<p>It does this:</p>
<blockquote>
<p>Here are the top 10 plausible next words, each with a probability.</p>
</blockquote>
<p>Example (simplified):</p>
<table>
<thead>
<tr>
<th>Word</th>
<th>Probability</th>
</tr>
</thead>
<tbody>
<tr>
<td>"increase"</td>
<td>40%</td>
</tr>
<tr>
<td>"optimize"</td>
<td>30%</td>
</tr>
<tr>
<td>"improve"</td>
<td>20%</td>
</tr>
<tr>
<td>"reduce"</td>
<td>10%</td>
</tr>
</tbody>
</table>
<p>However, instead of always picking the one with the highest probability, it often samples from the list. That means sometimes it goes left, sometimes right. And you can’t predict, and therefore depend on, when it does which. Therefore, results given by AI will never be inherently consistent. It’s simply not in its nature.</p>
<p>When you think about AI, think less about it as a ‘thinking brain’ and more about it as a very advanced pattern prediction engine training on enormous amounts of data.</p>
<h2 id="why-the-randomness-is-intentionally-built-in">Why the randomness is intentionally built in</h2>
<p>If the model would always pick the #1 most likely word the responses would be repetitive, creativity would collapse and conversations would feel robotic. That same mechanism is exactly why AI is both extremely powerful and fundamentally limited and therefore very good at certain tasks, but also very bad at others.</p>
<h2 id="using-ais-randomness">Using AI’s randomness</h2>
<h3 id="creative-work">Creative work</h3>
<p>For some brainstorming and creative work this is great!</p>
<blockquote>
<p>Give me 3 examples of a company logo which name is ‘JeffOps’.</p>
<p>Explain AI to me using 3 different analogies.</p>
</blockquote>
<p>You get variety. You get ideas. You get speed. So, AI is great for marketing? Yes, as a tool. Not as a replacement.</p>
<h3 id="brainstorming">Brainstorming</h3>
<p>Let’s take a question that I’ve used many times, or variations of it:</p>
<blockquote>
<p>How should I grow and/or improve my business?</p>
</blockquote>
<p>Valid answers to this might include:</p>
<ul>
<li>Improve pricing</li>
<li>Increase marketing spending</li>
<li>Focus on retention</li>
<li>Expand to new markets</li>
</ul>
<p>All of these exist in the models’ learned patterns and are given as answers to my question as a result. So, what you’re basically seeing is the model picking different valid paths through the same landscape (set of data), but with vastly different outcomes.</p>
<h2 id="the-scalability-trap-but-is-it-a-trap">The scalability trap (but is it a trap?)</h2>
<p>A common argument for the use of AI is that it’s right most of the time. Which is true, but that’s not the real issue. The real issue is scale.</p>
<p>If AI is wrong 10% of the time, it’s manageable. But if AI runs your business (context is always important 😉), that 10% becomes hundreds of mistakes. Let's take legal work. Writing a paragraph of 200 words would mean there will be 20 wrong words in it. This could become a very heavy, and very pricy endeavor 😉</p>
<p>What works fine in isolation turns into a liability nightmare at scale.</p>
<h2 id="the-illusion-of-competence">The illusion of competence</h2>
<p>Here's the trap: AI doesn't fail loudly. It fails convincingly. Its answers sound correct, look polished, and feel complete. The grammar is clean, the structure is logical, the tone is confident. And none of that tells you whether a single word of it is true.</p>
<p>That's the part worth sitting with. With people, we use polish as a shortcut for competence — and most of the time it works. When someone speaks fluently and confidently about a topic, it's usually because they actually understand it. Producing a clear, well-structured answer normally requires knowing the material. So we've learned to trust the signal.</p>
<p>AI breaks that link completely. Remember how these models work: they don't reason toward a correct answer, they predict plausible-sounding text. Plausibility is the objective. Fluency isn't a side effect of understanding — it's the entire product. The model produces the exact same confident, polished output whether it's right or completely wrong. There's no tell. No hesitation, no "I'm not sure about this," no awkward phrasing that gives it away.</p>
<p>So every instinct you'd normally use to judge whether to trust an answer — the very cues that serve you well with humans — is exactly the thing AI generates effortlessly, regardless of truth. Your judgment isn't just unhelpful here; it actively works against you.</p>
<p>If you've spent any time in ops or engineering, this should feel especially wrong. We're trained to trust that failures are visible. A bug throws an exception. A bad deploy crashes. A broken query returns an error, a null, a stack trace — something red, something loud. You know it failed. AI doesn't give you that. It fails silently, with a straight face, and hands you a wrong answer that's indistinguishable from a right one.</p>
<p>And that's where the most dangerous errors live: not the ones that blow up, but the ones that look perfectly fine and slip straight through. The mistakes you never catch are the expensive ones.</p>
<p>This is also why "it's right most of the time" is cold comfort. If the failures announced themselves, a 90% hit rate would be easy to manage — you'd just fix the 10% you can see. But the failures don't announce themselves. They hide inside the 90% that looks identical. So the real cost isn't the error rate. It's that you can't tell which answers are the errors without checking every one yourself.</p>
<p>Which brings it back to the point: AI is brilliant at producing things that look finished. Whether they're actually correct is a separate question — and answering it is still your job.</p>
<h2 id="1-person-ai-companies">1-person AI companies</h2>
<p>You've probably seen the claim: a single person, no employees, running an entire company because AI agents handle everything. Marketing, sales, support, content, ops — all automated. One human, a full business.</p>
<p>And here's the honest part: they can do this. The agents will produce marketing copy, answer tickets, draft posts, and generate graphics around the clock. On paper it looks like a whole team.</p>
<p>But look closer at the output.</p>
<p>You'll find the misspelled graphics. The messaging that says one thing on the landing page and the opposite in the email sequence. The "facts" that are outdated or were never true to begin with. These aren't signs of a lazy operator — they're the exact failure modes I described earlier, now running unsupervised. Remember the randomness baked into the model, and the illusion of competence? A solo operator has removed the one thing that used to catch those failures: a human reading the output before it ships.</p>
<p>And it compounds, because no single agent sees the whole business.</p>
<p>This is the real reason it breaks down. AI has no genuine contextual awareness — it doesn't understand your situation, your intent, or the consequences of being wrong. It only interprets patterns in the text you hand it. In practice that means:</p>
<ul>
<li>It misses the context a human just knows — internal history, customer relationships, where the market is heading.</li>
<li>It produces answers that sound right but don't fit your specific situation.</li>
<li>It fails on edge cases, exactly where judgment matters more than pattern-matching.</li>
<li>It treats every input as a "text problem" instead of a real-world decision.</li>
</ul>
<p>It doesn't see the situation. It only sees the words about the situation.</p>
<p>Now multiply that across agents. Your marketing agent, your sales agent, and your support agent each interpret their own slice of text with no shared understanding of the business. Individually their output might be fine. Together, they drift — and the customer is the one who notices the contradictions.</p>
<p>So yes, one person can run a company on agents. The question is whether the result holds up under scrutiny — and at any meaningful scale, it usually doesn't.</p>
<p>Note: there are tools and techniques that give models more context — RAG, memory, system instructions — but it's still nowhere near what a human brings to the table.</p>
<h2 id="what-not-to-use-ai-for">What not to use AI for</h2>
<p>Let’s take legal as a more concrete example. AI handles edge cases poorly, especially when it comes to complaints and legal issues. While some models perform better than others, law is a precise and context-heavy domain.</p>
<p>AI systems don’t understand the situation—they generate responses based on patterns—so they can produce different outputs for the same input or miss critical nuances. That makes them fundamentally unpredictable in high-stakes scenarios.</p>
<p>In legal contexts, even a slightly incorrect or poorly phrased statement can have serious consequences. And because the AI isn’t accountable, the responsibility always falls entirely on you.</p>
<p>That’s why I would never trust AI to handle legal matters independently; at best, it’s a drafting or research tool, not a decision-maker. And there you have it: A research tool.</p>
<p>When it comes to legal, what you can use AI for is reviewing contracts:</p>
<ul>
<li>The contract mentions a payment window of maximum 30 days or shorter.</li>
<li>There is no scenario mentioned where the payment window is allowed to exceed 30 days.</li>
<li>The contract does not include ambiguous phrasing that could be interpreted differently under certain conditions.</li>
<li>There are no conflicting clauses that override or weaken the 30‑day requirement elsewhere in the document.</li>
<li>Standard enforcement, dispute resolution, and penalty clauses are present and aligned with that payment term.</li>
</ul>
<p>In this context, AI works very well because the task is structured, the criteria are defined in advance and more importantly: You’re asking it to verify patterns, not make judgement calls.</p>
<p>It’s essentially acting as a high-speed checklist processor.</p>
<h2 id="what-to-use-ai-for">What to use AI for</h2>
<p>Every tool has its own use case. When you understand the basics of how AI models work (read above), you can start to plot it against your business and its processes.</p>
<p>A few cases where AI can clearly be useful:</p>
<ul>
<li>Contract scanning (not contract writing!)</li>
<li>Clause extraction</li>
<li>Consistency checks</li>
<li>First-pass reviews</li>
<li>Coding (more information about this specific case in my next newsletter!)</li>
</ul>
<p>But not for:</p>
<ul>
<li>Legal interpretation</li>
<li>Risk decisions</li>
<li>Negotiation strategies</li>
<li>Liability-bearing judgements</li>
</ul>
<p>One of there more ‘beautiful’ examples I’ve encountered thus far was a judge using AI to generate a verdict. The problem was that the AI model came up with precedents that were fake. They didn’t exist!</p>
<p>I’m guessing that’s why Microsoft’s Co-pilot client in Windows has the following message under its prompt input:</p>
<blockquote>
<p>AI-generated content may be incorrect.</p>
</blockquote>
<h2 id="the-system-problem-multiple-ai-agents">The system problem (multiple AI agents)</h2>
<p>Everything so far has been about a single model. The moment you start chaining them together, the problem doesn't just add up — it multiplies.</p>
<blockquote>
<p>The Marketing AI says one thing.</p>
<p>The Sales AI says something else.</p>
<p>The Support AI contradicts both.</p>
</blockquote>
<p>Individually, their answers might be fine. Good, even. But together? Inconsistent. And in business, consistency is everything.</p>
<p>But the visible contradictions are the easy problem — at least you can see those. The harder one is what happens underneath, where the agents feed each other.</p>
<p>In a real pipeline, the output of one agent becomes the input of the next. So when an agent gets something subtly wrong — a hallucinated detail, a misread of context — the next agent doesn't question it. It treats that wrong output as ground truth and builds on top of it. The error doesn't stay contained. It travels downstream, and it grows.</p>
<p>Now bring back the maths from earlier. Say each agent is 90% reliable on its own — pretty good, right? Watch what happens when you chain them:</p>
<ul>
<li>1 agent: 90% reliable. Manageable.</li>
<li>3 agents in a chain: 0.9 × 0.9 × 0.9 ≈ 73%. Already a 1-in-4 chance something's off by the end.</li>
<li>5 agents in a chain: 0.9⁵ ≈ 59%. Now it's basically a coin flip.</li>
</ul>
<p>Every handoff is another roll of the dice. Reliability doesn't hold steady across a system — it compounds downward. And because each agent only sees its own slice of text (remember: no real contextual awareness, no shared understanding of the business), there's nothing holding the whole thing together. No agent sees the full picture. No agent owns the outcome.</p>
<p>And here's where it ties back to the illusion of competence: each individual step still looks clean. Polished output, confident tone, no errors thrown. So when the final result is wrong, good luck tracing it back — was it the marketing agent? The third handoff? Something five steps upstream? There's no stack trace. Just a confident, broken result and no clear place to point.</p>
<p>A human team isn't immune to this — people miscommunicate too. But people share context, ask each other questions, and someone, somewhere, owns the result. Strip all of that out and replace it with a chain of confident, isolated, non-deterministic components, and you don't get a leaner business. You get a faster way to be consistently wrong.</p>
<h2 id="conclusion">Conclusion</h2>
<p>Please don’t misunderstand me: I love AI and I think it’s a transformative and game-changing tool for finding and structuring information, and to an extent automating certain tasks and perhaps even very specific jobs, but not for deciding what that information means in the real world.</p>
<p>Using AI for the right tasks can enhance the productivity of people, enhance the quality of their work, broaden their experiences and knowledge, and much more. But it should be used as such: A tool to enhance, not to replace.</p>]]></content:encoded>
  </item>
  <item>
    <title>GPU costs on Kubernetes: sharing is a reliability decision</title>
    <link>https://jeffops.com/posts/2026/gpu-costs-kubernetes-sharing/</link>
    <guid isPermaLink="true">https://jeffops.com/posts/2026/gpu-costs-kubernetes-sharing/</guid>
    <pubDate>Thu, 28 May 2026 22:00:00 +0000</pubDate>
    <description>GPU spend is idle time, not utilisation. Every way of sharing a card trades isolation for density, the metrics meant to guide that choice are misleading, and for LLM serving the defaults quietly work against you.</description>
    <category>kubernetes</category><category>gpu</category><category>ai</category><category>aks</category><category>cost</category><category>finops</category><category>llm</category><category>ops</category>
    <content:encoded><![CDATA[<p>Put a number on it before anything else. An eight-GPU H100 node (<a href="https://learn.microsoft.com/azure/virtual-machines/sizes/gpu-accelerated/ndh100v5-series"><code>Standard_ND96isr_H100_v5</code></a>) lists at $98.32 an hour in East US, which is around $72,000 a month if you leave it running. The eight-GPU A100 equivalent is $27.20 an hour, about $20,000 a month. Even a single-GPU H100 node is roughly $5,000 a month. Those are Azure retail Linux rates as of August 2026 and they move, but the order of magnitude is the point: one forgotten pool outweighs most teams' entire non-production estate.</p>
<p>Nobody decides to spend that. It happens because a pool was created for a project, the project finished, and no alert fires for expensive idleness.</p>
<p>That is the shape of accelerator spend. Not expensive work, expensive waiting. And the reflex when someone notices the number is to pack more workloads onto each card, which is a reasonable instinct with a consequence people skip past: <strong>every mechanism for sharing a GPU trades isolation for density</strong>. The useful question is not how much more you can fit. It is what failure you are willing to have shared.</p>
<p>I wrote separately about <a href="/posts/2026/cutting-kubernetes-costs-without-cutting-reliability/">Kubernetes cost in general</a>, where the argument is that overspend is the gap between requests and usage. Accelerators are that argument with the volume turned up and most of the usual tooling unavailable.</p>
<p>The same note as the companion piece. I have not run GPU workloads in anger, so read this as a researched guide and not as a war story. That caveat carries more weight here than it does for ordinary compute, because published GPU utilisation figures are unusually unreliable, and a good part of this article is about why.</p>
<h2 id="how-bad-is-idle-actually">How bad is idle, actually</h2>
<p>Microsoft states the mechanism plainly: <a href="https://learn.microsoft.com/azure/architecture/reference-architectures/containers/aks-gpu/gpu-aks">you incur cost on a GPU node pool even when no GPU workload is running</a>. The harder question is what fraction of paid GPU time does useful work, and the published evidence is thinner than the confident numbers in circulation suggest.</p>
<p>The best-sourced measurement I found is the Alibaba PAI trace published at <a href="https://www.usenix.org/system/files/nsdi22-paper-weng.pdf">NSDI '22</a>. 6,742 GPUs, 1.2 million tasks. The headline pair is the one worth carrying: the <strong>median instance requested 0.5 of a GPU and used 0.042 of one</strong>. A twelve-fold gap, in production, at scale. Heavy utilisation of 95% or above accounted for only 7% of cases, and their scheduler simulation suggested half the GPUs would have sufficed with sharing enabled.</p>
<p>Two caveats I want to state rather than bury, because the article later criticises other people's methodology. <strong>The trace was collected in July and August 2020.</strong> That is a multi-tenant ML platform dominated by training and classical ML, before LLM serving existed as a workload class and before continuous batching. The request-versus-usage gap is a durable finding about human behaviour under uncertainty. The specific numbers describe notebook-scale jobs on pre-Ampere hardware, not a fleet running inference on H100s.</p>
<p>There is also a <a href="https://www.microsoft.com/en-us/research/wp-content/uploads/2024/01/gpu-util-icse2024.pdf">Microsoft Research study from ICSE 2024</a> which examined 400 deep learning jobs, and found 85% of low-utilisation issues were fixable with small code or script changes, mostly around data loading and batch size. Note that study deliberately sampled jobs already below 50% utilisation, so it characterises causes rather than fleet averages.</p>
<p>Vendor reports quoting single-digit average GPU utilisation may well be right. The ones I could check do not publish their methodology, and the companies publishing them sell optimisation software. Treat them as a prompt to measure, not as a measurement.</p>
<h2 id="your-utilisation-dashboard-misleads-in-both-directions">Your utilisation dashboard misleads in both directions</h2>
<p>Before optimising anything, know that the obvious metric does not mean what its name implies.</p>
<p><strong>SM Activity is a duty cycle.</strong> <a href="https://docs.nvidia.com/datacenter/dcgm/latest/user-guide/feature-overview.html">NVIDIA defines it</a> as the fraction of time at least one warp was active on a multiprocessor, averaged across multiprocessors, and gives two examples that both produce 0.2. On a GPU with N streaming multiprocessors, a kernel launching N/5 blocks that runs for the whole interval reports 0.2. So does a kernel launching N blocks that runs for a fifth of the interval with the SMs idle the rest of the time.</p>
<p>That is the problem in one metric: 0.2 could mean a fifth of the card all of the time, or all of the card a fifth of the time, and the number cannot tell you which. A warp waiting on memory also counts as active.</p>
<p><strong>Memory utilisation is worse, because inference servers pre-allocate.</strong> vLLM reserves GPU memory for its KV cache at startup, with <code>gpu_memory_utilization</code> defaulting to 0.92. The memory graph therefore reads 92% permanently, whether the server is saturated or idle.</p>
<p>Put those together and an idle inference pod can look busy while a loaded one looks unremarkable. What to use instead:</p>
<ul>
<li><strong>GPU-seconds allocated against GPU-seconds used</strong>, per namespace and workload. <a href="https://learn.microsoft.com/azure/aks/best-practices-gpu-observability">Microsoft's GPU observability guidance</a> recommends exactly this as a shared metric between platform and finance, because aggregate averages hide over-allocation.</li>
<li><strong>Serving metrics for inference.</strong> Tokens per second, requests waiting, KV-cache block utilisation, time to first token. These tell you whether the server is saturated; GPU percentage does not.</li>
<li><strong>DCGM enriched with pod labels</strong>, so telemetry attributes to a team rather than to a card.</li>
</ul>
<p>Same failure I described in <a href="/posts/2026/ai-observability-four-problems/">instrumenting AI systems</a>: the instrument you already trust is answering a question that stopped being the important one.</p>
<h2 id="the-three-ways-to-share-a-card">The three ways to share a card</h2>
<p><a href="https://learn.microsoft.com/azure/aks/concepts-gpu-partitioning">Microsoft's comparison</a> is a good frame, and the isolation column is the one to read first.</p>
<table>
<thead>
<tr>
<th></th>
<th>Isolation</th>
<th>Density</th>
<th>Status</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>Time-slicing</strong></td>
<td>None. Shared memory, shared fault domain</td>
<td>Highest, arbitrary N</td>
<td>Stable, any CUDA GPU</td>
</tr>
<tr>
<td><strong>MPS</strong></td>
<td>Memory capped per client, limited error containment</td>
<td>High</td>
<td>Experimental in the device plugin, incompatible with MIG</td>
</tr>
<tr>
<td><strong>MIG</strong></td>
<td>Hardware partitioning, dedicated memory, error isolation</td>
<td>Up to 7</td>
<td>Production-recommended, Ampere and later only</td>
</tr>
</tbody>
</table>
<p><a href="https://github.com/NVIDIA/k8s-device-plugin">NVIDIA's device plugin documentation</a> is blunt about time-slicing: replicas of the same GPU run in the same fault domain, so if one workload crashes they all do. It also states that requesting more than one time-sliced GPU does not guarantee a proportional share of compute. MPS is better but not safe: NVIDIA's guidance says a fatal GPU fault from one client is <a href="https://docs.nvidia.com/deploy/mps/when-to-use-mps.html">reported to all clients on the affected GPUs</a>, without indicating which one caused it, and the MPS server waits for all of them to exit. Contained to the shared GPU rather than isolated from it.</p>
<p>Two MIG constraints decide architecture rather than configuration.</p>
<p><strong>Hardware.</strong> <a href="https://docs.nvidia.com/datacenter/tesla/mig-user-guide/supported-mig-profiles.html">NVIDIA's supported profile list</a> covers A100, H100, H200 and B200 at seven instances, A30 and RTX PRO 6000 Blackwell at four, and a couple of Blackwell workstation parts at two. <strong>L4, L40S, T4 and V100 do not appear at all</strong>, which is an expensive planning surprise, because L4 and L40S are exactly the cards a cost-conscious team reaches for.</p>
<p><strong>Geometry is fixed at node pool creation.</strong> Microsoft states that partitioning is static at pool level and changes require reprovisioning; through the GPU Operator, reconfiguration stops all GPU pods on the node and may need a reboot. Since nobody can name the model they will be serving in a year, the practical advice is to pick a geometry that fits the largest model you serve today with headroom, and accept that a genuinely new model class means a new node pool.</p>
<h2 id="the-llm-traps">The LLM traps</h2>
<p>This is where general Kubernetes cost advice goes wrong, because it was written for training and batch.</p>
<p><strong>Time-slicing plus vLLM fails on the defaults, and the fix is not obvious.</strong> Set <code>replicas: 4</code> on the device plugin and deploy four vLLM servers, and the first one reserves 92% of the card. The rest die. The instinct is to blame time-slicing for having no memory cap.</p>
<p>That is not quite right, and it matters. vLLM's own <a href="https://docs.vllm.ai/en/stable/configuration/engine_args/">field documentation</a> says <code>gpu_memory_utilization</code> "is a per-instance limit, and only applies to the current vLLM instance. It does not matter if you have another vLLM instance running on the same GPU. For example, if you have two vLLM instances running on the same GPU, you can set the GPU memory utilization to 0.5 for each instance." The cap exists. It just has to be set by you, on every instance, in agreement with a replica count configured somewhere completely different, with nothing validating that the two add up.</p>
<p>So this is a footgun rather than an incompatibility. Two independent settings that must be kept consistent by hand, in different files, owned by different teams, where getting it wrong produces a crash loop at deploy time and getting it <em>nearly</em> right produces a KV cache too small to serve your batch size. You still get no fault isolation and no compute guarantee. Set both deliberately or use MIG.</p>
<p><strong>Requesting more than one time-sliced GPU silently does nothing.</strong> <code>failRequestsGreaterThanOne</code> defaults to <code>false</code> for backwards compatibility, so a pod requesting two replicas is admitted, runs, and receives no proportional share. Setting it to <code>true</code> makes the request fail honestly, but understand what "fail" means here before you flip it. The extended resource still exists on the node, so the scheduler binds the pod happily and the <strong>kubelet</strong> rejects it at admission with <code>UnexpectedAdmissionError</code>. The pod lands in <code>Failed</code>, not <code>Pending</code>, and it does not self-heal: NVIDIA's documentation says you must manually delete the pod, change the resource request and redeploy. A controller behind it will keep producing pods that keep failing. Fix every manifest requesting more than one first, then set the flag.</p>
<p><strong>Sharing the server is a real option with its own cost, not a free escape.</strong> The alternative to putting four servers on a card is one server with continuous batching handling all the traffic, with multiple adapters if you need multiple behaviours. That is what these servers are built for and it usually gets better throughput than four fragmented ones.</p>
<p>But be honest about what it does to the failure mode: one process serving every tenant is not less fault coupling, it is more. A crash, a bad rollout or an OOM takes every in-flight request and the whole KV cache with it. Time-slicing at least gives each tenant its own process to lose.</p>
<p>So the actual choice is between three failure shapes: many processes sharing a fault domain (time-slicing), one process that is a single fault domain (shared server), or genuine isolation at fixed granularity on hardware that supports it (MIG).</p>
<p><div class="figure-scroll" tabindex="0" role="group" aria-label="Diagram, scrollable"><svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 880 350" width="880" height="350" role="img" aria-labelledby="bl-t bl-d" style="max-width:100%;height:auto;display:block;margin:1.5rem 0"><title id="bl-t">Three ways to share a GPU, and the blast radius of each</title><desc id="bl-d">Time-slicing places four pods in one shared fault domain. A shared inference server places every tenant inside a single process. MIG gives each pod an isolated hardware partition.</desc><defs><marker id="bl-a" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"><path d="M0,0 L10,5 L0,10 z" fill="#0aa5c4"/></marker><marker id="bl-aw" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"><path d="M0,0 L10,5 L0,10 z" fill="#d2593c"/></marker></defs><text x="20" y="24" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="15" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="600">Time-slicing</text><text x="20" y="42" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11.5" fill="currentColor" fill-opacity="0.62" text-anchor="start" font-weight="400">one shared fault domain</text><rect x="20" y="62" width="250" height="200" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="32" y="84" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11" fill="currentColor" fill-opacity="0.62" text-anchor="start" font-weight="400">GPU</text><rect x="32" y="96" width="226" height="138" rx="4" fill="none" stroke="#d2593c" stroke-opacity="1" stroke-width="1.5" stroke-dasharray="5 4"/><rect x="44" y="110" width="92" height="40" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="90" y="135" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">pod 1</text><rect x="150" y="110" width="92" height="40" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="196" y="135" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">pod 2</text><rect x="44" y="162" width="92" height="40" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="90" y="187" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">pod 3</text><rect x="150" y="162" width="92" height="40" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="196" y="187" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">pod 4</text><text x="32" y="248" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="12" fill="#d2593c" fill-opacity="1" text-anchor="start" font-weight="600">a crash takes all four</text><text x="310" y="24" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="15" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="600">Shared server</text><text x="310" y="42" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11.5" fill="currentColor" fill-opacity="0.62" text-anchor="start" font-weight="400">one process, every tenant</text><rect x="310" y="62" width="250" height="200" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="322" y="84" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11" fill="currentColor" fill-opacity="0.62" text-anchor="start" font-weight="400">GPU</text><rect x="322" y="96" width="226" height="138" rx="4" fill="none" stroke="#d2593c" stroke-opacity="1" stroke-width="1.5" stroke-dasharray="5 4"/><rect x="336" y="110" width="198" height="108" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="435.0" y="132" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">one server process</text><text x="435.0" y="154" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">tenant A  B  C  D</text><text x="322" y="248" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="12" fill="#d2593c" fill-opacity="1" text-anchor="start" font-weight="600">the largest single domain</text><text x="600" y="24" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="15" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="600">MIG</text><text x="600" y="42" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11.5" fill="currentColor" fill-opacity="0.62" text-anchor="start" font-weight="400">hardware partitions</text><rect x="600" y="62" width="250" height="200" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="612" y="84" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11" fill="currentColor" fill-opacity="0.62" text-anchor="start" font-weight="400">GPU</text><rect x="624" y="106" width="92" height="40" rx="4" fill="none" stroke="#0aa5c4" stroke-opacity="1" stroke-width="1.5" stroke-dasharray="5 4"/><text x="670" y="131" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">pod 1</text><rect x="730" y="106" width="92" height="40" rx="4" fill="none" stroke="#0aa5c4" stroke-opacity="1" stroke-width="1.5" stroke-dasharray="5 4"/><text x="776" y="131" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">pod 2</text><rect x="624" y="158" width="92" height="40" rx="4" fill="none" stroke="#0aa5c4" stroke-opacity="1" stroke-width="1.5" stroke-dasharray="5 4"/><text x="670" y="183" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">pod 3</text><rect x="730" y="158" width="92" height="40" rx="4" fill="none" stroke="#0aa5c4" stroke-opacity="1" stroke-width="1.5" stroke-dasharray="5 4"/><text x="776" y="183" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">pod 4</text><text x="612" y="248" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="12" fill="#0aa5c4" fill-opacity="1" text-anchor="start" font-weight="600">a fault stays in its slice</text><text x="20" y="336" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="12" fill="currentColor" fill-opacity="0.7" text-anchor="start" font-weight="400">Dashed boundary = fault domain. Density rises left to right in cost, not in safety.</text></svg></div></p>
<p>The shared server wins on throughput. MIG wins on blast radius. Neither wins on both.</p>
<h2 id="dynamic-resource-allocation-and-where-it-actually-is">Dynamic Resource Allocation, and where it actually is</h2>
<p>Core Dynamic Resource Allocation <a href="https://kubernetes.io/blog/2025/09/01/kubernetes-v1-34-dra-updates/">graduated to GA in Kubernetes 1.34</a>, so the structured device allocation model has been stable for the better part of a year. Be precise about what that covers, though: the piece that actually shares one device across pods, Consumable Capacity, is still alpha, and Partitionable Devices is beta. The GA milestone is the allocation framework, not fractional sharing.</p>
<p>The vendor side is the gap. NVIDIA's DRA driver documents its multi-node NVLink ComputeDomains support as officially supported while GPU allocation features, including dynamic MIG, are not yet officially supported, and the GPU kubelet plugin ships disabled by default in the Helm chart. Check the current release notes before building a strategy on it, because this is the part most likely to have moved since August 2026. Today the production path remains the device plugin with MIG or time-slicing.</p>
<h2 id="scale-to-zero-costs-minutes-not-seconds">Scale-to-zero costs minutes, not seconds</h2>
<p>Scaling an idle inference endpoint to zero is the obvious answer, and it is a latency decision as much as a cost one.</p>
<p>Model serving images are enormous. Azure's troubleshooting guidance for the AI toolchain operator states that <a href="https://learn.microsoft.com/troubleshoot/azure/azure-kubernetes/extensions/troubleshoot-ai-toolchain-operator-addon-issues">inference images are typically 30 GB to 100 GB and the pull can take up to tens of minutes</a> depending on cluster networking. Add node provisioning if the pool is also at zero, then weight loading, then CUDA graph capture.</p>
<p>Mitigations, with their trade-offs stated rather than ranked:</p>
<ul>
<li><strong>Pre-pull and cache images on nodes.</strong> Costs disk, saves the largest single component of cold start.</li>
<li><strong>Keep a warm node pool even when pods are zero.</strong> Effective, and it is not really scaling to zero. You are paying for the node to avoid paying for the wait.</li>
<li><strong>Model weights on a shared volume rather than baked into the image.</strong> Sometimes faster, sometimes slower: RWX file storage pulling 70 GB can underperform a cached image pull, and it converts many independent cold starts into one shared bottleneck. Benchmark it rather than assuming.</li>
<li><strong>One or two warm replicas during business hours, zero overnight.</strong> The pragmatic answer for anything a human waits on.</li>
</ul>
<p>Scale-to-zero is right for batch, evaluation and internal tooling. For an interactive surface, a warm replica is cheaper than the abandoned session.</p>
<h2 id="separate-training-from-inference">Separate training from inference</h2>
<p>Standard practice, for reasons more specific than "different workloads".</p>
<p><strong>Different interruption tolerance.</strong> Training checkpoints and resumes, so it can live on spot. Inference is restartable in principle but latency-sensitive in practice: a replacement replica that takes minutes to load weights is an outage from the user's point of view, even though the process itself restarts fine.</p>
<p><strong>Different scaling shapes.</strong> Training scales gradually and wants large, tightly coupled, topology-aware multi-GPU nodes. Inference follows traffic and often prefers one or two GPUs per node, spread for availability.</p>
<p><strong>Different SKUs.</strong> Training wants A100 or H100 class with NVLink. Inference frequently runs happily on L4 or A10 at a fraction of the price.</p>
<p><strong>Different sharing modes.</strong> MIG or dedicated for latency-sensitive serving; time-slicing is defensible for experimentation where a shared fault domain costs somebody an afternoon.</p>
<p>Separate node pools with taints and tolerations is the mechanism.</p>
<h2 id="queueing-and-the-spot-caveat-nobody-states">Queueing, and the spot caveat nobody states</h2>
<p>Two large jobs each holding half the GPUs, neither able to start. Everything allocated, nothing progressing.</p>
<p><a href="https://kueue.sigs.k8s.io/docs/overview/">Kueue</a> is the Kubernetes-native answer: quota with fair sharing, fungibility across resource flavours, preemption, and <strong>all-or-nothing gang admission</strong>, which is what fixes the deadlock. Its provisioning-request integration also stops you spinning up expensive nodes for a job that then cannot be admitted. The API is at v1beta2, so treat it as maturing.</p>
<p>Which leads to a warning that follows directly and which most spot advice omits. <strong>Gang-scheduled multi-node training on spot is close to the worst case.</strong> Losing one node kills the whole job, you forfeit every GPU-hour since the last checkpoint across the entire allocation, and re-acquiring a large homogeneous allocation of scarce SKUs at the moment you need it is exactly when capacity is least available. Spot suits single-node and embarrassingly parallel work with frequent checkpoints. For large gang-scheduled runs, do the arithmetic rather than trusting either instinct: spot on that same eight-GPU H100 SKU is around $18 an hour against $98 on-demand, so you can afford to waste roughly five times the compute in restarts before spot loses. Whether you do depends on your eviction rate and your checkpoint interval, which is why both are worth measuring before the argument.</p>
<p>Where spot does fit, Microsoft's boundary is clear: work that can be checkpointed or restarted cleanly, never user-facing inference. That requires shutdown inside thirty seconds, checkpoints written outside the spot VM, idempotent jobs and pre-baked images.</p>
<p>And before the argument about whether spot is too risky, look up the actual number. Azure publishes <strong>per-SKU, per-region eviction rates in bands</strong>, queryable through the <code>SpotResources</code> table in <a href="https://learn.microsoft.com/azure/virtual-machines/spot-vms">Azure Resource Graph</a> over the trailing 28 days. The portal shows the same bands over a 7-day window, so the two will not always agree. A SKU in the 0 to 5% band and one in the 20%-plus band are not the same proposition.</p>
<h2 id="check-you-are-optimising-the-right-bill">Check you are optimising the right bill</h2>
<p>This section is last in most treatments of the subject and probably should not be.</p>
<p>Microsoft's <a href="https://learn.microsoft.com/startups/build/ai/ai-cost-optimization">AI cost optimisation guidance</a> breaks the typical bill down with <strong>tokens at 30 to 60% and GPUs at 20 to 50%</strong>. Those are indicative ranges on a guidance page rather than a measured benchmark, and they overlap, so do not treat them as precise. But the direction is worth taking seriously: if any part of your workload calls hosted models, the request path may be the larger lever, and prompt caching, routing simple queries to smaller models, batch APIs and quantisation all act on it without touching a node pool.</p>
<p>That same page is where I took the framing for this article's opening, and it is worth reading in full.</p>
<p>There is also a failure mode hiding in agent architectures: <strong>retrieval fan-out</strong>. A single chat turn can issue several hidden queries through re-rankers, query rewriters and tool-calling steps, each costing tokens and latency. Keeping retrievals per turn low and alerting on the median is only possible if the trace carries them, which is the argument from the observability piece arriving from another direction.</p>
<p>So: work out the split between token spend and infrastructure spend before you spend a quarter on cluster efficiency. If tokens dominate, most of this article is the second priority.</p>
<h2 id="what-i-would-do-first">What I would do first</h2>
<p>Measure allocated GPU-seconds against used GPU-seconds, per namespace. Everything else is guesswork without it, and the number is usually enough on its own to make the case for the rest of the work.</p>
<p>Then stop trusting GPU percentage, and move inference decisions onto serving metrics. Then go looking for pools that exist because a project needed them months ago, with the caveat that in a capacity-constrained region a pool you release may not be a pool you can get back, so check availability before deleting rather than after.</p>
<p>After that: split training from inference onto separate pools, move checkpointable single-node work to spot with the eviction rates checked, decide deliberately between a shared server and MIG for serving rather than defaulting into time-slicing, and queue the batch work so partial allocation stops holding cards hostage.</p>
<h2 id="the-trade-nobody-writes-down">The trade nobody writes down</h2>
<p>Every sharing decision here has the same shape. More density, less isolation. The saving is continuous, visible and easy to attribute. The cost is a probability of an incident that will be attributed to something else entirely when it arrives.</p>
<p>That is not an argument against sharing, and it would be a poor one. Running a single workload per A100 because it feels safer is its own large waste, and on hardware that supports MIG the dichotomy largely dissolves: you get isolation and density together, and the honest version of this article's title on that hardware is narrower, roughly "do not run time-slicing in production."</p>
<p>What remains is the part worth writing down. Which workloads may share a fault domain, which may not, and who decided. That document is worth more than any of the individual levers above, because it is the thing that stops the next cost exercise quietly making the choice for you.</p>]]></content:encoded>
  </item>
  <item>
    <title>Cutting Kubernetes costs without cutting reliability</title>
    <link>https://jeffops.com/posts/2026/cutting-kubernetes-costs-without-cutting-reliability/</link>
    <guid isPermaLink="true">https://jeffops.com/posts/2026/cutting-kubernetes-costs-without-cutting-reliability/</guid>
    <pubDate>Sun, 03 May 2026 22:00:00 +0000</pubDate>
    <description>Most Kubernetes overspend is the gap between what you requested and what you used. Closing it safely is a reliability exercise, and the saving only reaches the invoice if it removes nodes.</description>
    <category>kubernetes</category><category>aks</category><category>cost</category><category>finops</category><category>reliability</category><category>ops</category>
    <content:encoded><![CDATA[<p>Someone in finance has looked at the cluster bill and asked the obvious question, and the obvious answers are all bad. Fewer replicas. Smaller nodes. Turn off the redundancy nobody has needed yet. Each of those saves money by spending reliability, and you pay it back with interest during the next incident.</p>
<p>Here is the line worth starting from: <strong>most Kubernetes overspend is not a pricing problem. It is the gap between what you asked for and what you used.</strong></p>
<p>That changes who owns the fix. A pricing problem belongs to procurement. A requests problem belongs to whoever wrote the manifest.</p>
<p><strong>But be clear about the mechanism, because this is where cost articles cheat.</strong> You are not billed for requests. You are billed for nodes. Reducing a workload's requests from 4 GiB to 1 GiB saves you nothing at all until that freed capacity lets the autoscaler remove a node. A cluster that drops from 60% requested to 40% requested on an unchanged node count has improved a dashboard and not an invoice.</p>
<p>So the work is two-part and the second half is the one people skip: close the gap, then make sure the freed space actually consolidates. Most of this article is about the second half.</p>
<p><div class="figure-scroll" tabindex="0" role="group" aria-label="Diagram, scrollable"><svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 760 520" width="760" height="520" role="img" aria-labelledby="ch-t ch-d" style="max-width:100%;height:auto;display:block;margin:1.5rem 0"><title id="ch-t">How a smaller resource request becomes a smaller bill</title><desc id="ch-d">A five step chain from reduced requests to a lower invoice, with four labelled points at which the chain commonly breaks.</desc><defs><marker id="ch-a" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"><path d="M0,0 L10,5 L0,10 z" fill="#0aa5c4"/></marker><marker id="ch-aw" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"><path d="M0,0 L10,5 L0,10 z" fill="#d2593c"/></marker></defs><rect x="24" y="30" width="330" height="62" rx="4" fill="none" stroke="#0aa5c4" stroke-opacity="1" stroke-width="1.5"/><text x="40" y="56" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="15" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="600">1. Requests reduced</text><text x="40" y="76" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="start" font-weight="400">VPA, in-place resize</text><line x1="189.0" y1="96" x2="189.0" y2="120" stroke="#0aa5c4" stroke-width="1.5" marker-end="url(#ch-a)"/><rect x="24" y="124" width="330" height="62" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="40" y="150" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="15" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="600">2. Capacity freed</text><text x="40" y="170" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="start" font-weight="400">on the nodes you already run</text><line x1="189.0" y1="190" x2="189.0" y2="214" stroke="#0aa5c4" stroke-width="1.5" marker-end="url(#ch-a)"/><rect x="24" y="218" width="330" height="62" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="40" y="244" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="15" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="600">3. Pods consolidate</text><text x="40" y="264" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="start" font-weight="400">packing, Karpenter consolidation</text><line x1="189.0" y1="284" x2="189.0" y2="308" stroke="#0aa5c4" stroke-width="1.5" marker-end="url(#ch-a)"/><rect x="24" y="312" width="330" height="62" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="40" y="338" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="15" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="600">4. Node removed</text><text x="40" y="358" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="start" font-weight="400">autoscaler drains and deletes</text><line x1="189.0" y1="378" x2="189.0" y2="402" stroke="#0aa5c4" stroke-width="1.5" marker-end="url(#ch-a)"/><rect x="24" y="406" width="330" height="62" rx="4" fill="none" stroke="#0aa5c4" stroke-opacity="1" stroke-width="1.5"/><text x="40" y="432" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="15" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="600">5. Invoice falls</text><text x="40" y="452" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12" fill="currentColor" fill-opacity="0.62" text-anchor="start" font-weight="400">the only step finance sees</text><line x1="432" y1="202.0" x2="368" y2="202.0" stroke="#d2593c" stroke-width="1.5" marker-end="url(#ch-aw)" stroke-dasharray="4 3"/><text x="440" y="189.0" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11" fill="#d2593c" fill-opacity="1" text-anchor="start" font-weight="600">breaks here</text><text x="440" y="206.0" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="11.5" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="400">LeastAllocated spreads pods,</text><text x="440" y="220.0" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="11.5" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="400">so no node ever empties</text><line x1="432" y1="296.0" x2="368" y2="296.0" stroke="#d2593c" stroke-width="1.5" marker-end="url(#ch-aw)" stroke-dasharray="4 3"/><text x="440" y="262.0" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11" fill="#d2593c" fill-opacity="1" text-anchor="start" font-weight="600">breaks here</text><text x="440" y="279.0" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="11.5" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="400">· Injected anti-affinity makes</text><text x="440" y="293.0" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="11.5" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="400">  consolidation skip the node</text><text x="440" y="307.0" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="11.5" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="400">· consolidateAfter resets on churn</text><text x="440" y="321.0" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="11.5" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="400">· A PodDisruptionBudget blocks</text><text x="440" y="335.0" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="11.5" fill="currentColor" fill-opacity="1" text-anchor="start" font-weight="400">  the last eviction</text><text x="24" y="508" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="12" fill="currentColor" fill-opacity="0.7" text-anchor="start" font-weight="400">Steps 2 to 4 are where cost work quietly fails. A dashboard improves; the bill does not.</text></svg></div></p>
<p>A note on what this is, because it should change how you weigh it. I have not run a cluster cost programme end to end, so nothing here is a war story. It is what the documentation and the published measurements say when you go and read them, together with the failure modes that keep turning up in other people's write-ups. Every number below is linked to whoever produced it, and where the evidence is thin I have said so rather than rounded it into confidence.</p>
<h2 id="where-the-money-actually-goes">Where the money actually goes</h2>
<p>Split the bill before changing anything. Kubernetes spend lands in five buckets that behave differently:</p>
<ul>
<li><strong>Compute.</strong> Nodes. Usually the largest line and usually the one holding the waste.</li>
<li><strong>Accelerators.</strong> If you run GPUs this can dominate everything else, and it follows different rules. It has <a href="/posts/2026/gpu-costs-kubernetes-sharing/">its own article</a>.</li>
<li><strong>Storage.</strong> Persistent volumes, OS disks, snapshots, and the disks left behind when a workload moved.</li>
<li><strong>Network.</strong> Egress, cross-zone traffic, load balancers that outlived their service.</li>
<li><strong>Observability.</strong> Your monitoring bill is part of your cluster bill, and it is the one nobody attributes.</li>
</ul>
<p>Within compute, the waste is rarely the node price. The scheduler reserves what you requested whether you use it or not, and the request was a guess made once by whoever wrote the deployment. A node at 20% actual utilisation showing 90% requested capacity is not an efficient node, and no amount of SKU shopping fixes it.</p>
<p>Azure's idle-cost documentation puts it plainly: a node that is ready with no pods running, waiting for the autoscaler, is <a href="https://learn.microsoft.com/azure/aks/cost-analysis-idle-costs">mostly idle cost</a>.</p>
<h3 id="measuring-the-gap">Measuring the gap</h3>
<p>Before any of this, get the number. If you run Prometheus, this is the whole diagnosis in one query:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code class="language-promql"><span class="c1"># Memory used as a fraction of memory requested, per pod, averaged over a week.</span>
<span class="c1"># max() on both sides matters: cAdvisor can export stale duplicate series for a</span>
<span class="c1"># restarted container, and kube-state-metrics emits one series per container.</span>
<span class="kr">avg_over_time</span><span class="o">(</span>
<span class="w">  </span><span class="o">(</span>
<span class="w">    </span><span class="k">sum</span><span class="w"> </span><span class="k">by</span><span class="w"> </span><span class="o">(</span><span class="nv">namespace</span><span class="p">,</span><span class="w"> </span><span class="nv">pod</span><span class="o">)</span><span class="w"> </span><span class="o">(</span>
<span class="w">      </span><span class="k">max</span><span class="w"> </span><span class="k">by</span><span class="w"> </span><span class="o">(</span><span class="nv">namespace</span><span class="p">,</span><span class="w"> </span><span class="nv">pod</span><span class="p">,</span><span class="w"> </span><span class="nv">container</span><span class="o">)</span><span class="w"> </span><span class="o">(</span>
<span class="w">        </span><span class="nv">container_memory_working_set_bytes</span><span class="p">{</span><span class="nl">container</span><span class="o">!=</span><span class="p">&quot;&quot;,</span><span class="w"> </span><span class="nl">container</span><span class="o">!=</span><span class="p">&quot;</span><span class="s">POD</span><span class="p">&quot;}</span>
<span class="w">      </span><span class="o">)</span>
<span class="w">    </span><span class="o">)</span>
<span class="w">    </span><span class="o">/</span>
<span class="w">    </span><span class="k">sum</span><span class="w"> </span><span class="k">by</span><span class="w"> </span><span class="o">(</span><span class="nv">namespace</span><span class="p">,</span><span class="w"> </span><span class="nv">pod</span><span class="o">)</span><span class="w"> </span><span class="o">(</span>
<span class="w">      </span><span class="k">max</span><span class="w"> </span><span class="k">by</span><span class="w"> </span><span class="o">(</span><span class="nv">namespace</span><span class="p">,</span><span class="w"> </span><span class="nv">pod</span><span class="p">,</span><span class="w"> </span><span class="nv">container</span><span class="o">)</span><span class="w"> </span><span class="o">(</span>
<span class="w">        </span><span class="nv">kube_pod_container_resource_requests</span><span class="p">{</span><span class="nl">resource</span><span class="o">=</span><span class="p">&quot;</span><span class="s">memory</span><span class="p">&quot;}</span>
<span class="w">      </span><span class="o">)</span>
<span class="w">      </span><span class="o">*</span><span class="w"> </span><span class="k">on</span><span class="w"> </span><span class="o">(</span><span class="nv">namespace</span><span class="p">,</span><span class="w"> </span><span class="nv">pod</span><span class="o">)</span><span class="w"> </span><span class="k">group_left</span>
<span class="w">        </span><span class="o">(</span><span class="nv">kube_pod_status_phase</span><span class="p">{</span><span class="nl">phase</span><span class="o">=~</span><span class="p">&quot;</span><span class="s">Pending|Running</span><span class="p">&quot;}</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="mi">1</span><span class="o">)</span>
<span class="w">    </span><span class="o">)</span>
<span class="w">  </span><span class="o">)</span><span class="p">[</span><span class="s">7d</span><span class="err">:</span><span class="s">1h</span><span class="p">]</span>
<span class="o">)</span></code></pre>
<p>Swap in <code>rate(container_cpu_usage_seconds_total[5m])</code> and <code>resource="cpu"</code> for the CPU version. Anything sitting below 0.3 is your list.</p>
<p>Two things this does not do, and both matter. It gives you a ratio, so run <code>sum by (namespace, pod) (kube_pod_container_resource_requests{resource="memory"})</code> alongside it and sort by that, because a 10% ratio on a 128 MiB pod is noise and a 60% ratio on a 32 GiB one is real money. And pods with no memory request at all produce no denominator and vanish silently, which is exactly the population the governance section below is about. Find those separately with <code>kube_pod_container_info unless on (namespace, pod, container) kube_pod_container_resource_requests{resource="memory"}</code>.</p>
<h2 id="the-requests-gap-is-a-reliability-problem-in-a-finance-costume">The requests gap is a reliability problem in a finance costume</h2>
<p>Here is why "just lower the requests" is not the whole answer.</p>
<p>CPU and memory fail differently, and the asymmetry is the game.</p>
<p><strong>CPU is elastic.</strong> Exceed your CPU limit and the kernel throttles you. The pod gets slower, latency rises, a queue builds. Unpleasant and survivable, and visible on a graph if you are watching.</p>
<p><strong>Memory is not.</strong> Exceed your memory limit and the container is OOMKilled mid-request. There are leading indicators if you look for them, working set climbing and garbage collection getting busier, but the failure itself has no gradual mode. It is a cliff.</p>
<p>So the two deserve opposite instincts. Under-requesting CPU costs latency. Under-requesting memory costs availability. Treat them as one slider and you will eventually trade an outage for a saving.</p>
<p>A practical position: aggressive on CPU requests, conservative on memory, and memory limits equal to memory requests <strong>for workloads whose memory profile you have actually measured</strong>. That last clause matters. Setting limits equal to requests on an unmeasured JVM heap or a Go service with a growing cache guarantees an OOMKill at precisely the point a looser limit would have survived. Measure first, then pin.</p>
<p>CPU limits are genuinely contested. Leaving them off avoids throttling a burst that had capacity available; setting them makes behaviour reproducible and stops one workload starving its neighbours. Decide per workload class.</p>
<p>One interaction that is easy to miss: <a href="https://kubernetes.io/docs/concepts/cluster-administration/swap-memory-management/">node swap is stable as of Kubernetes 1.34</a>, and under <code>LimitedSwap</code> only Burstable pods may swap. Setting memory request equal to limit makes a container Guaranteed, which opts it out of swap entirely. If you are experimenting with swap as an overcommit lever, that advice and the paragraph above are in tension.</p>
<h2 id="right-sizing-no-longer-means-restarting-everything">Right-sizing no longer means restarting everything</h2>
<p>Historically, changing a pod's CPU or memory meant recreating it. That made right-sizing a scheduled, disruptive event, so it happened rarely, so requests drifted between rounds.</p>
<p><strong>In-place pod resize is <a href="https://kubernetes.io/docs/tasks/configure-pod-container/resize-container-resources/">stable as of Kubernetes 1.35</a></strong> and on by default. You can change a running container's allocation while potentially avoiding disruption. Paired with the Vertical Pod Autoscaler, right-sizing becomes continuous rather than scheduled.</p>
<p>Three cautions. Resizing memory downwards is the awkward direction and depends on the per-resource resize policy you set, so read the semantics; do not assume symmetry. VPA acting automatically on anything holding data deserves the same care as any other automated change to state: recommendation mode first. And <strong>do not point VPA at CPU while an HPA is scaling on CPU</strong>, because the two will fight, VPA raising requests as HPA adds replicas to the same signal. VPA on memory alongside HPA on CPU is the combination that works.</p>
<p>Recommendation mode is worth running even if you never let it act. Azure Advisor <a href="https://learn.microsoft.com/azure/advisor/advisor-reference-cost-recommendations">recommends exactly that</a> as a first step, and comparing what VPA thinks each workload needs against what it currently requests is the fastest way to size the prize.</p>
<p>For sidecar-heavy clusters, note that pod-level resources reached beta and default-on in Kubernetes 1.34, letting a pod hold one shared envelope rather than each container carrying its own padding. In-place resizing of pod-level resources reached beta in 1.36.</p>
<h2 id="the-overhead-you-pay-per-node">The overhead you pay per node</h2>
<p>Allocatable is not capacity, and the difference is a strong argument for fewer, larger nodes.</p>
<p>Every node reserves CPU and memory before your pods get anything, and <a href="https://learn.microsoft.com/azure/aks/node-resource-reservations">AKS publishes the schedule</a>. The CPU reservation is regressive: a 2-core node gives up 100 millicores, which is 5% of it; a 64-core node gives up 740 millicores, which is 1.2%. Memory reservation is the lesser of 20 MB per max-pod plus 50 MB, or 25% of system memory, putting an 8 GB node at 30 max pods at roughly 90.6% allocatable.</p>
<p>Two things compound it. Every DaemonSet is a per-node tax, so your log shipper, node exporter, CSI and CNI agents multiply by node count rather than by workload. And there is a pod-slot ceiling: <a href="https://learn.microsoft.com/azure/aks/concepts-network-ip-address-planning">AKS allows between 10 and 250 pods per node</a>, defaulting to 250 on Azure CNI Overlay and 30 on Azure CNI with standard networking. A fleet of small pods on a default standard-networking pool hits 30 long before it fills CPU or memory, which is a configuration problem rather than a law of nature, but it is one worth checking.</p>
<p>Consolidating to fewer, larger nodes removes copies of all of that. The trade is real: a larger node is a larger blast radius and coarser autoscaling granularity.</p>
<p>While counting fixed costs, the AKS Standard tier control plane is <a href="https://learn.microsoft.com/azure/architecture/aws-professional/eks-to-aks/cost-management#aks-cost-basics">$0.10 per cluster per hour</a>, roughly $73 a month flat regardless of size. Ten clusters is $730 before a single node, and underneath each sits a system node pool floor of at least two nodes at 4 vCPU minimum.</p>
<h2 id="make-the-nodes-fit-the-pods">Make the nodes fit the pods</h2>
<p>Once requests reflect reality, the next waste is nodes that do not fit them. Three pods needing 5 GB each on an 8 GB node means two nodes and a lot of stranded memory.</p>
<p><strong>Node autoprovisioning</strong>, which in AKS is <a href="https://learn.microsoft.com/azure/aks/node-auto-provisioning">built on Karpenter</a>, looks at pending pods and provisions the VM configuration that fits them, rather than making you define pools and hope.</p>
<p>The money is mostly in <strong>consolidation</strong> rather than in initial provisioning: reclaiming capacity after scale-down, after a deployment shrinks, after churn. That is where steady-state drift accumulates. If you are reading older guidance, the policy names changed: <code>WhenUnderutilized</code> is gone, and <a href="https://karpenter.sh/docs/concepts/disruption/">current Karpenter</a> offers <code>WhenEmpty</code>, <code>Balanced</code> and <code>WhenEmptyOrUnderutilized</code>, the last being the default.</p>
<p>Two things quietly stop consolidation working, and this is the part that connects back to the thesis. Karpenter's documentation warns that <strong>preferred anti-affinity and topology spread constraints reduce consolidation effectiveness</strong>, because it keeps trying to honour your preferences and skips otherwise-valid moves. And <code>consolidateAfter</code> resets whenever a pod is added or removed, so a churning node never becomes a candidate.</p>
<p>That matters more than it first appears on AKS Automatic, because Deployment Safeguards runs there in Enforce mode and <strong>injects preferred pod anti-affinity and topology spread constraints into workloads that lack them</strong>. The platform marketed for cost efficiency mandatorily adds the two constraints that impede the mechanism doing most of the saving. Both facts are on Microsoft's own pages. Neither page mentions the other.</p>
<p>Two smaller levers: <strong>Arm64 node pools</strong>, which Microsoft describes as offering <a href="https://learn.microsoft.com/azure/aks/best-practices-cost">up to 50% better price-performance</a> for scale-out workloads, with mixed architectures supported in one cluster. And <strong>keep application images small</strong>, because every scale-out downloads them and slow starts cause the overprovisioning discussed below. (Model-serving images are a different problem, covered in the GPU article.)</p>
<p>One infrastructure detail with a real recurring cost: AKS <a href="https://learn.microsoft.com/azure/aks/concepts-storage#ephemeral-os-disks-in-aks">defaults to ephemeral OS disks</a> where the VM SKU supports the requested size, and the VM price includes them, so you avoid the managed OS disk charge. That charge scales with the node: AKS defaults to a 128 GiB P10 at 1 to 7 vCPU, rising to 1024 GiB P30 at 64 vCPU and above, so the saving is larger on exactly the big nodes recommended above. The fallback is silent: ask for 100 GiB on a VM with 75 GiB of temp storage and you get a managed disk without being told. Ephemeral also rules out <a href="https://learn.microsoft.com/azure/virtual-machines/ephemeral-os-disks">disk snapshots, Azure Disk Encryption, Azure Backup and Site Recovery</a> on the node, so check your compliance posture before assuming it is free money.</p>
<h2 id="scale-to-zero-where-the-workload-allows-it">Scale to zero where the workload allows it</h2>
<p>The Horizontal Pod Autoscaler handles demand that varies. <strong>KEDA</strong> handles demand that stops, scaling on queue depth, event backlog and schedules, and going to zero. Anything queue-driven or batch should not be sitting idle overnight waiting for work that arrives at nine.</p>
<p>Two HPA defaults worth knowing. Scale-down stabilisation is 300 seconds while scale-up can double every 15 seconds, an asymmetry most people never look at. And <strong>configurable tolerance reached beta and default-on in 1.35</strong>, so you can now set different tolerances per direction: tight on scale-up for fast reaction, loose on scale-down so a 2% dip does not cause churn.</p>
<p>Thrash matters less for the replicas than for what follows them. Replica churn drives node churn, node churn drives image pulls and warm-up CPU, and it keeps resetting the consolidation timers so nodes never become candidates for removal.</p>
<p><strong>The slow-start trap raises your floor permanently.</strong> The HPA ignores a pod's CPU during the <a href="https://kubernetes.io/docs/concepts/workloads/autoscaling/horizontal-pod-autoscale/">CPU initialisation period</a>, five minutes by default, and treats not-yet-ready pods as consuming nothing. A workload taking minutes to become Ready is invisible to its own autoscaler for that window. The team concludes the HPA is too slow and raises <code>minReplicas</code>. That floor is now permanent cost. Fixing start time, with a startup probe so Ready means ready, is what lets it come back down.</p>
<h2 id="bin-packing-and-the-tool-that-undoes-it">Bin packing, and the tool that undoes it</h2>
<p>Two levers that only work as a pair.</p>
<p>The default scheduler scoring strategy is <code>LeastAllocated</code>, which spreads pods across nodes. Every node ends up partially full, which is exactly the state in which the autoscaler cannot remove any of them. Switching <code>NodeResourcesFit</code> to a packing strategy concentrates idle capacity onto whole nodes that can be deleted.</p>
<p>This is now possible on AKS: <a href="https://learn.microsoft.com/azure/aks/configure-node-binpack-scheduler">configurable scheduler profiles</a> arrived in preview for Kubernetes 1.33 and later. Read Microsoft's guidance rather than reaching for the obvious setting, because two details matter. Raw <code>MostAllocated</code> "risks saturating nodes beyond desirable limits, causing throttling or additional bottlenecks", and Microsoft recommends <code>RequestedToCapacityRatio</code> for production instead, which lets you target a utilisation band and deprioritise nodes above it. And Microsoft says you <strong>must disable the <code>PodTopologySpread</code> plugin</strong>, because it can override the <code>NodeResourcesFit</code> weighted score.</p>
<p>Which raises a combination this article has now recommended in three places and which you should not deploy blind: aggressive CPU requests, plus deliberate packing, plus a descheduler evicting from under-used nodes. CPU requests set the kernel's CPU weight (<code>cpu.weight</code> on cgroup v2), not just placement. Aggressively low requests on a deliberately saturated node means starvation under contention, missed liveness probes, restarts, and rescheduling onto other saturated nodes. Introduce these one at a time.</p>
<p>On the descheduler itself: its <code>HighNodeUtilization</code> plugin evicts from under-utilised nodes so they can be emptied, and its <a href="https://github.com/kubernetes-sigs/descheduler">documentation states</a> the plugin <strong>must</strong> be used with <code>MostAllocated</code> scoring. Without it, evicted pods spread straight back out and you have built an eviction loop that costs money. Run it as a CronJob rather than a hot loop, with <code>nodeFit: true</code> so it checks a pod can be placed before evicting it, and with eviction caps. It evicts and hopes; it does not schedule replacements.</p>
<h2 id="spot-done-in-a-way-you-will-not-regret">Spot, done in a way you will not regret</h2>
<p>Azure is explicit: spot node pools have <strong>no SLA</strong>, sit in a single fault domain, and provide <a href="https://learn.microsoft.com/azure/architecture/aws-professional/eks-to-aks/cost-management">no high-availability guarantees</a>, with eviction notice of <a href="https://learn.microsoft.com/azure/architecture/guide/spot/spot-eviction">at least 30 seconds</a> delivered best-effort.</p>
<p>The rules follow:</p>
<ul>
<li><strong>Never put stateful workloads, single-replica services, or the only capacity behind a user-facing path on spot.</strong></li>
<li><strong>Mix, for workloads that tolerate it.</strong> An on-demand baseline that can carry the service, with spot absorbing peaks. This works for stateless request handlers with fast startup. It does not work for latency-sensitive inference, where a cold replacement replica takes minutes.</li>
<li><strong>Set PodDisruptionBudgets and mean them</strong>, so a reclamation cannot take your last healthy replica.</li>
<li><strong>Make shutdown fit in thirty seconds.</strong></li>
<li><strong>Spread across zones and instance types</strong>, so one capacity squeeze does not take a whole tier.</li>
</ul>
<p>Azure publishes real eviction rates per SKU and region, so this does not have to be an argument about vibes. The <a href="/posts/2026/gpu-costs-kubernetes-sharing/">GPU article</a> covers how to query them, where the stakes make it matter more.</p>
<h2 id="commitments-and-one-thing-that-is-not-a-discount">Commitments, and one thing that is not a discount</h2>
<p>Reservations and savings plans produce a discount <a href="https://learn.microsoft.com/azure/aks/best-practices-cost">up to 72% against pay-as-you-go</a> with no runtime change at all.</p>
<p>There is a real reason to right-size first, and it is not that discounts "multiply your waste", which is arithmetically confused since the two levers commute. It is <strong>shape lock-in</strong>. A one or three year commitment is a bet on a fleet profile. Make it before you have closed the requests gap and you have committed to a shape you are about to change, and the unused portion is not refundable because you got more efficient.</p>
<p><strong>Capacity reservations are not a discount.</strong> They guarantee capacity is available to you, which is genuine reliability value in a constrained region, but Azure's billing documentation is blunt that a capacity reservation bills at full rate for the reserved quantity <a href="https://learn.microsoft.com/azure/virtual-machines/capacity-reservation-overview#pricing-and-billing">whether or not you use it</a>. They can be covered by a reservation or savings plan, which is the only way they get cheaper.</p>
<h2 id="the-bucket-nobody-attributes-your-monitoring-bill">The bucket nobody attributes: your monitoring bill</h2>
<p>Usually the fastest saving in the exercise, and the one people forget is part of the cluster bill at all.</p>
<p>Azure Advisor recommends switching from log-based Container Insights metrics to <a href="https://learn.microsoft.com/azure/advisor/advisor-reference-cost-recommendations">Managed Prometheus, which it describes as up to 80% cheaper</a> for the same metric data. If you run both, you are paying twice for substantially the same numbers. That one is close to free money.</p>
<p>The rest of the list is not free money, and I want to be careful here, because this section is where a cost article most easily commits the sin it opened by condemning. Every lever below trades some future ability to answer a question:</p>
<ul>
<li><strong>Collect logs and events only</strong> in Container Insights if you have Managed Prometheus. Low risk, mostly removes duplication.</li>
<li><strong>Move container logs to ContainerLogV2 and Basic Logs</strong> if you do not query them routinely. You lose alerting and most query capability on that table. Fine for chatty application stdout, wrong for anything you investigate.</li>
<li><strong>Turn off control plane log categories you genuinely do not use.</strong> Not <code>kube-audit</code>. Audit logs are what answer "who deleted that secret" during a security investigation, and they frequently carry a retention obligation. Trim the noisy categories, keep the forensic one, and if audit volume is the problem, use ingestion-time transformations to filter it rather than switching it off.</li>
<li><strong>Alert on metrics rather than logs</strong> where the signal exists in both.</li>
<li><strong>Use ingestion-time transformations</strong> to drop or reshape data before it lands, so you never pay for what you discard.</li>
</ul>
<p>One trap when combining these. <a href="https://learn.microsoft.com/azure/azure-monitor/logs/cost-logs#commitment-tiers">Commitment tiers</a> start at 100 GB per day and save as much as 30%, but apply <strong>only to Analytics Logs</strong>. Basic and Auxiliary Logs bill at flat per-GB rates and are excluded. The "move to Basic Logs" lever and the "buy a commitment tier" lever do not compose.</p>
<p>The same discipline appears in <a href="/posts/2026/ai-observability-four-problems/">instrumenting AI systems</a>: telemetry volume is a design decision with a price attached, and the default is rarely right.</p>
<h2 id="storage-and-network">Storage and network</h2>
<p>Persistent volumes outlive the workloads that created them, and a <code>Retain</code> reclaim policy leaves the disk behind when the claim goes. List unattached disks and orphaned snapshots against live claims periodically, and delete them with an owner attached, not with hope.</p>
<p>On disk choice, classic Premium SSD ties IOPS to capacity, so a workload needing 4,000 IOPS forces a 1 TiB disk even if it stores 50 GiB. Premium SSD v2 decouples them and includes 3,000 IOPS and 125 MB/s at no extra charge, at the cost of being LRS only with no zone-redundant option, which is a reliability trade rather than a free upgrade.</p>
<p>Cross-zone traffic is billable, so a chatty service spread across three zones pays for that redundancy twice. <code>spec.trafficDistribution</code> with <code>PreferSameZone</code> <a href="https://kubernetes.io/docs/reference/networking/virtual-ips/#traffic-distribution">went stable in Kubernetes 1.35</a> and is the current mechanism. Be honest about what it does: keeping traffic in-zone reduces egress cost and consumes the cross-zone failover headroom you were paying for.</p>
<p>Do not confuse it with the older <code>service.kubernetes.io/topology-mode: Auto</code> annotation, which has been beta since 1.23 and distributes proportionally with safeguards, including falling back to cluster-wide routing when there are too few endpoints. <code>PreferSameZone</code> deliberately dropped those heuristics for predictability: if a zone has endpoints they take all of that zone's traffic, and it falls back only when the zone has none. Simpler, and it makes overloading a zone's endpoints your problem rather than the control plane's.</p>
<p>And every load balancer and public IP left behind by a deleted service is a standing charge nobody is watching.</p>
<h2 id="when-the-governance-control-raises-the-bill">When the governance control raises the bill</h2>
<p>AKS <a href="https://learn.microsoft.com/azure/aks/deployment-safeguards">Deployment Safeguards</a> in Enforce mode assigns <strong>500 millicores and 2 GiB</strong>, as both request and limit, to any pod arriving with no resources set, and raises anything below 100m or 100Mi to those floors. That is sound governance, and on a cluster of small utility pods that previously ran unspecified, 2 GiB of <em>requested</em> memory each is a scheduling floor that the node count follows upwards.</p>
<p>Worth knowing before you rely on it as a control: Microsoft documents Gatekeeper as operating fail-open, so if the admission webhook does not respond the validation is skipped. It is a strong default, not a guarantee.</p>
<p><strong>LimitRange has a sharper edge.</strong> A container specifying <strong>only a limit</strong> gets its request set equal to that limit, and this happens <em>whether or not</em> you have set a <code>defaultRequest</code> for the namespace. Your default is simply ignored for that container. A team asks for burst headroom and silently reserves all of it for the pod's lifetime. <a href="https://kubernetes.io/docs/tasks/administer-cluster/manage-resources/cpu-default-namespace/">Upstream documents this</a>, and <a href="https://kubernetes.io/docs/concepts/policy/limit-range/">notes separately</a> that with two LimitRanges in a namespace, which default applies is not deterministic. Keep one.</p>
<p>Used deliberately the pairing works: a ResourceQuota makes requests mandatory by rejecting pods without them, and a single LimitRange supplies sane defaults.</p>
<h2 id="non-production-is-where-the-easy-money-is">Non-production is where the easy money is</h2>
<p>Development, test and staging are frequently a large share of the bill and the savings are uncontroversial.</p>
<p>Scale non-production node pools down outside working hours, with the obvious caveats: check nightly CI, scheduled integration suites and colleagues in other timezones before picking the window, and prefer a small floor to a hard zero if anything needs to run unattended. Put an expiry on ephemeral environments so a preview namespace from a merged pull request does not run until the heat death of the universe. Use spot aggressively here, because an interrupted CI runner is an inconvenience rather than an incident.</p>
<h2 id="finance-wants-a-number-this-quarter">"Finance wants a number this quarter"</h2>
<p>The fair objection to everything above: you have been told to spend a month measuring, and someone wants a saving on this quarter's report.</p>
<p>Two answers. The monitoring bill and idle non-production capacity are both actionable in days, carry little risk, and are large enough to be worth reporting. Start there, and they buy you the time for the rest.</p>
<p>And be straight about the alternative. The fast levers, cutting replicas and shrinking nodes, are available to anyone in an afternoon. They work. They also spend reliability, and that spend does not appear on the same report as the saving. If the decision is made anyway, make sure it is recorded as a decision with a trade attached, so the incident review has somewhere to point.</p>
<h2 id="the-order-i-would-do-it-in">The order I would do it in</h2>
<ol>
<li><strong>Get visibility.</strong> Cost analysis on, VPA in recommendation mode, and the requests-versus-usage query above. <a href="https://learn.microsoft.com/azure/aks/cost-analysis">AKS Cost Analysis</a> breaks spend down by Kubernetes construct, separating idle, system and unallocated charges; <a href="https://www.opencost.io/docs/configuration/azure">OpenCost</a> is the vendor-neutral option.</li>
<li><strong>Fix the monitoring bill</strong>, minus the audit logs.</li>
<li><strong>Close the requests gap.</strong> Aggressive on CPU, careful on measured memory, using in-place resize so it is not a disruptive event. Keep PodDisruptionBudgets on everything that matters, because this generates a lot of voluntary disruption.</li>
<li><strong>Make the freed space consolidate</strong>, which is the step that turns a dashboard change into an invoice change. Node autoprovisioning, packing, fewer larger nodes, and a check that nothing is injecting anti-affinity you did not ask for.</li>
<li><strong>Scale to zero</strong> what can be, non-production first.</li>
<li><strong>Introduce spot</strong> behind PDBs with an on-demand baseline.</li>
<li><strong>Commit</strong> to the steady-state remainder, once the shape has stopped moving.</li>
<li><strong>Show teams their own numbers.</strong> Showback works because the person who set <code>memory: 4Gi</code> because it seemed safe is usually happy to fix it once they can see the cost. They were being cautious with a number nobody had ever shown them.</li>
</ol>
<p>Change one thing at a time and watch it. An exercise that changes six variables in a week cannot tell you which one caused the incident. And keep an eye on the error budget while you work: if reliability degrades as costs fall, you have not optimised anything, you have sold something that was not yours to sell.</p>
<h2 id="what-this-does-not-do">What this does not do</h2>
<p>None of it will halve your bill in a fortnight, and any guide promising a specific percentage is describing someone else's cluster. What is available depends entirely on how far your requests have drifted from your usage, which is knowable only by looking.</p>
<p>What it does do is separate the savings you can take safely from the ones that quietly borrow against reliability. The bill and the reliability are not opposing forces. They are both downstream of whether your resource requests tell the truth, and of whether the truth reaches the node count.</p>]]></content:encoded>
  </item>
  <item>
    <title>AI observability is four problems wearing one name</title>
    <link>https://jeffops.com/posts/2026/ai-observability-four-problems/</link>
    <guid isPermaLink="true">https://jeffops.com/posts/2026/ai-observability-four-problems/</guid>
    <pubDate>Thu, 23 Apr 2026 22:00:00 +0000</pubDate>
    <description>Monitoring tells you the AI ran. It cannot tell you it was right. Four layers, four different denominators, and a privacy flag that protects your users by blinding your quality monitoring.</description>
    <category>observability</category><category>ai</category><category>opentelemetry</category><category>azure</category><category>ops</category><category>agents</category>
    <content:encoded><![CDATA[<p>An AI system returns a confident, well-formatted, completely wrong answer. Status 200. Latency inside the p95. No exception, nothing in the error log, dashboard green.</p>
<p>Your alerting is built on exceptions, status codes, timeouts, saturation and error rates. Every one of those is a signal that something announced itself. <strong>Monitoring tells you the system ran. It cannot tell you the system was right</strong>, and until now those were close enough to the same question that nobody had to separate them.</p>
<p>That is the illusion of competence with a dashboard in front of it.</p>
<p>The reason your telemetry feels thin the first time you point it at an AI system is not that you instrumented it badly. It is that "AI observability" is four different problems and most instrumentation treats them as one.</p>
<ol>
<li><strong>The call.</strong> One inference. Request in, tokens out.</li>
<li><strong>What you fed it.</strong> The prompt, the system instructions, and whatever you retrieved and put in front of the question.</li>
<li><strong>The tools.</strong> The functions the model is allowed to invoke. Real side effects.</li>
<li><strong>The loop.</strong> The agent that runs the other three until it decides it is finished.</li>
</ol>
<p>The fourth one is not a layer. It is the axis the other three run along, and it is the reason per-call metrics stop meaning anything. More on that below, because it is the part most instrumentation gets wrong.</p>
<p>If you only do a handful of things: get the structural attributes flowing, put one trace around the whole task, version your prompts on the span, promote the quiet failures to loud ones, and add evaluation as a separate sampled track. The rest of this is why.</p>
<h2 id="the-call-tells-you-almost-nothing-you-care-about">The call tells you almost nothing you care about</h2>
<p>The model call is the easiest thing to instrument and the most misleading thing to look at. You get latency, token counts, cost and a status. All of it is about the call, not the answer. A call can be fast, cheap, successful and wrong. There is no field for wrong.</p>
<p>The attributes worth carrying, from the <a href="https://github.com/open-telemetry/semantic-conventions-genai">OpenTelemetry GenAI semantic conventions</a>:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code>gen_ai.operation.name           chat
gen_ai.provider.name            azure.ai.openai
gen_ai.request.model            &lt;your deployment name&gt;
gen_ai.response.model           &lt;the dated version actually served&gt;
gen_ai.usage.input_tokens       3184
gen_ai.usage.output_tokens      412
gen_ai.response.finish_reasons  [&quot;length&quot;]
gen_ai.prompt.name              triage-classifier
gen_ai.prompt.version           v7
gen_ai.conversation.id          c-8f21...
</code></pre>
<p>Two of those repay attention.</p>
<p><strong><code>gen_ai.response.finish_reasons</code> is the closest thing you have to an exception.</strong> A finish reason of <code>length</code> means the model stopped because it ran out of output budget, not because it finished the thought. The answer is truncated, the call succeeded, and your user got half a sentence and a conclusion that never arrived. <code>content_filter</code> is the other one worth catching, and on a filtered endpoint it can account for more truncation than <code>length</code> does. Measure the split rather than assuming it.</p>
<p><strong><code>gen_ai.request.model</code> and <code>gen_ai.response.model</code> are separate fields for a reason.</strong> On Azure you call a deployment name rather than a model alias, and if that deployment is set to auto-update to the default version, the version underneath it rolls forward and your outputs change without you deploying anything. If the response model is not on every span, you cannot answer "what changed on Tuesday", because on your side nothing did.</p>
<h3 id="querying-finish-reasons-and-the-trap-in-it">Querying finish reasons, and the trap in it</h3>
<p>Here is where I have to be careful, because the obvious query does not work everywhere.</p>
<p><code>gen_ai.response.finish_reasons</code> is typed <code>string[]</code>. The Azure Monitor exporters flatten span attributes into <code>customDimensions</code>, which is a string-to-string map, and they do it differently per language. The .NET exporter calls <code>Convert.ToString</code> on the tag value, and for a string array that returns the literal text <code>System.String[]</code>. Python stringifies the tuple, so you get <code>"('length',)"</code> and a <code>has</code> operator matches by accident.</p>
<p>So a query written against the array works in Python, silently returns nothing in .NET, and looks identical in both cases. Agent Framework and Semantic Kernel both ship in Python and .NET, so which behaviour you get depends on the language your agents are written in, not on the framework. Check before you trust the number.</p>
<p>Set your own scalar alongside it and query that:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code class="language-python"><span class="n">finish</span> <span class="o">=</span> <span class="n">response</span><span class="o">.</span><span class="n">choices</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span><span class="o">.</span><span class="n">finish_reason</span>          <span class="c1"># &quot;stop&quot; | &quot;length&quot; | &quot;content_filter&quot;</span>
<span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;gen_ai.response.finish_reasons&quot;</span><span class="p">,</span> <span class="p">[</span><span class="n">finish</span><span class="p">])</span>   <span class="c1"># spec-conformant</span>
<span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;app.finish_reason&quot;</span><span class="p">,</span> <span class="n">finish</span><span class="p">)</span>                  <span class="c1"># queryable everywhere</span></code></pre>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code class="language-kusto"><span class="c">// Truncation and filtering as a *rate*. A raw count rises with traffic</span>
<span class="c">// and tells you nothing about quality.</span>
<span class="n">dependencies</span>
<span class="p">|</span><span class="w"> </span><span class="k">where</span><span class="w"> </span><span class="n">timestamp</span><span class="w"> </span><span class="p">&gt;</span><span class="w"> </span><span class="n">ago</span><span class="p">(</span><span class="mi">1</span><span class="n">d</span><span class="p">)</span>
<span class="p">|</span><span class="w"> </span><span class="k">where</span><span class="w"> </span><span class="n">tostring</span><span class="p">(</span><span class="n">customDimensions</span><span class="p">[</span><span class="s">&quot;gen_ai.operation.name&quot;</span><span class="p">])</span><span class="w"> </span><span class="p">==</span><span class="w"> </span><span class="s">&quot;chat&quot;</span>
<span class="p">|</span><span class="w"> </span><span class="k">extend</span><span class="w"> </span><span class="n">finish</span><span class="w"> </span><span class="p">=</span><span class="w"> </span><span class="n">tostring</span><span class="p">(</span><span class="n">customDimensions</span><span class="p">[</span><span class="s">&quot;app.finish_reason&quot;</span><span class="p">])</span>
<span class="p">|</span><span class="w"> </span><span class="k">summarize</span>
<span class="w">    </span><span class="n">cut_short</span><span class="w"> </span><span class="p">=</span><span class="w"> </span><span class="n">countif</span><span class="p">(</span><span class="n">finish</span><span class="w"> </span><span class="n">in</span><span class="w"> </span><span class="p">(</span><span class="s">&quot;length&quot;</span><span class="p">,</span><span class="w"> </span><span class="s">&quot;content_filter&quot;</span><span class="p">)),</span>
<span class="w">    </span><span class="n">total</span><span class="w">     </span><span class="p">=</span><span class="w"> </span><span class="k">count</span><span class="p">()</span>
<span class="w">  </span><span class="k">by</span><span class="w"> </span><span class="n">bin</span><span class="p">(</span><span class="n">timestamp</span><span class="p">,</span><span class="w"> </span><span class="mi">1</span><span class="n">h</span><span class="p">),</span><span class="w"> </span><span class="n">model</span><span class="w"> </span><span class="p">=</span><span class="w"> </span><span class="n">tostring</span><span class="p">(</span><span class="n">customDimensions</span><span class="p">[</span><span class="s">&quot;gen_ai.request.model&quot;</span><span class="p">])</span>
<span class="p">|</span><span class="w"> </span><span class="k">extend</span><span class="w"> </span><span class="n">pct</span><span class="w"> </span><span class="p">=</span><span class="w"> </span><span class="n">round</span><span class="p">(</span><span class="mf">100.0</span><span class="w"> </span><span class="p">*</span><span class="w"> </span><span class="n">cut_short</span><span class="w"> </span><span class="p">/</span><span class="w"> </span><span class="n">total</span><span class="p">,</span><span class="w"> </span><span class="mi">2</span><span class="p">)</span></code></pre>
<p>Watch it for a week before you set a threshold. Truncation rate is workload-specific: a summarisation service with a tight output budget may sit at several percent quite happily, while a classifier that ever truncates is broken. Alert on a deviation from your own baseline rather than an absolute number.</p>
<p>Two conventions worth matching while you are in there. The span name should be <code>{gen_ai.operation.name} {gen_ai.request.model}</code>, so <code>chat gpt-4.1</code> rather than <code>chat</code>. And the inference span <a href="https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md">should be <code>SpanKind.CLIENT</code></a>, though the spec allows <code>INTERNAL</code> for a model running in the same process. Hand-rolled spans default to <code>INTERNAL</code>, which breaks nothing but looks wrong to anyone reading your traces against the spec.</p>
<h2 id="most-confident-wrong-answers-start-upstream-of-the-model">Most confident wrong answers start upstream of the model</h2>
<p>When a model states something untrue, the instinct is to blame the model. Often enough the model behaved perfectly and faithfully summarised the wrong document.</p>
<p>If you are doing retrieval, the interesting telemetry is not in the model call at all. It is in what you handed the model: which documents came back and their identifiers, how many and how much of the context window they consumed, the score distribution and whether anything crossed the relevance threshold at all, and whether the final answer cited anything that was actually in the retrieved set.</p>
<p>That last one is the whole game. An answer citing nothing you retrieved is an answer the model produced from its own weights. It may be right, it may be stale, and it is definitely not what your retrieval system was for.</p>
<p>The conventions give you a <code>retrieval</code> operation with <code>gen_ai.data_source.id</code>, plus <code>gen_ai.retrieval.documents</code> and <code>gen_ai.retrieval.query.text</code> as opt-in attributes.</p>
<p>Prompt versioning is in the spec too, which surprised me: <code>gen_ai.prompt.name</code> and <code>gen_ai.prompt.version</code> are conditionally required on the inference span when a named template is used. Use them rather than inventing your own namespace, because conformant tooling will read them and your own attributes will be ignored. Keep custom attributes for things the spec genuinely does not cover, such as how many documents came back as opposed to how many you asked for:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code class="language-python"><span class="k">with</span> <span class="n">tracer</span><span class="o">.</span><span class="n">start_as_current_span</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;chat </span><span class="si">{</span><span class="n">model</span><span class="si">}</span><span class="s2">&quot;</span><span class="p">,</span> <span class="n">kind</span><span class="o">=</span><span class="n">SpanKind</span><span class="o">.</span><span class="n">CLIENT</span><span class="p">)</span> <span class="k">as</span> <span class="n">span</span><span class="p">:</span>
    <span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;gen_ai.operation.name&quot;</span><span class="p">,</span> <span class="s2">&quot;chat&quot;</span><span class="p">)</span>
    <span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;gen_ai.provider.name&quot;</span><span class="p">,</span> <span class="s2">&quot;azure.ai.openai&quot;</span><span class="p">)</span>
    <span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;gen_ai.prompt.name&quot;</span><span class="p">,</span> <span class="s2">&quot;triage-classifier&quot;</span><span class="p">)</span>
    <span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;gen_ai.prompt.version&quot;</span><span class="p">,</span> <span class="s2">&quot;v7&quot;</span><span class="p">)</span>
    <span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;gen_ai.conversation.id&quot;</span><span class="p">,</span> <span class="n">conversation_id</span><span class="p">)</span>
    <span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;app.retrieval.doc_count&quot;</span><span class="p">,</span> <span class="nb">len</span><span class="p">(</span><span class="n">docs</span><span class="p">))</span></code></pre>
<p>A prompt behaves like code. It is deployed, it changes behaviour, and it can be rolled back. Plenty of teams already keep prompts in version control with review; plenty do not, and in those the prompt changes without appearing in any release note. Either way, if the version is not on the span you cannot correlate a quality drop with the change that caused it.</p>
<h2 id="tool-invocation-is-honest-tool-internals-may-not-be">Tool invocation is honest. Tool internals may not be</h2>
<p>Here is the good news. A tool call either ran or it threw, took arguments you can inspect, and returned a result you can log. Spans are named <code>execute_tool {gen_ai.tool.name}</code>, with <code>gen_ai.tool.call.arguments</code> and <code>gen_ai.tool.call.result</code> available opt-in.</p>
<p>The caveat is that the spec's own <code>gen_ai.tool.type</code> includes <code>datastore</code>, and Microsoft's LangChain integration maps retrievers onto <code>execute_tool</code>. Sub-agents-as-tools and remote MCP servers land in the same span type. So tool <em>invocation</em> is deterministic and observable; a good share of what sits behind it is not.</p>
<p>The failure that matters here is not the tool erroring. That is loud, and loud you can already handle. It is the model calling the wrong tool, or the right tool with plausible but incorrect arguments, and that tool succeeding perfectly. Every metric says success. A record got updated. It was the wrong record.</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code class="language-python"><span class="k">with</span> <span class="n">tracer</span><span class="o">.</span><span class="n">start_as_current_span</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;execute_tool </span><span class="si">{</span><span class="n">tool_name</span><span class="si">}</span><span class="s2">&quot;</span><span class="p">)</span> <span class="k">as</span> <span class="n">span</span><span class="p">:</span>
    <span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;gen_ai.operation.name&quot;</span><span class="p">,</span> <span class="s2">&quot;execute_tool&quot;</span><span class="p">)</span>
    <span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;gen_ai.tool.name&quot;</span><span class="p">,</span> <span class="n">tool_name</span><span class="p">)</span>
    <span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;gen_ai.tool.call.id&quot;</span><span class="p">,</span> <span class="n">call_id</span><span class="p">)</span>
    <span class="c1"># The model&#39;s choice is the thing under test, not the function.</span>
    <span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;app.tool.args_valid&quot;</span><span class="p">,</span> <span class="n">schema_ok</span><span class="p">)</span>
    <span class="n">span</span><span class="o">.</span><span class="n">set_attribute</span><span class="p">(</span><span class="s2">&quot;app.tool.write&quot;</span><span class="p">,</span> <span class="n">is_mutating</span><span class="p">)</span></code></pre>
<p>Flagging mutating tools separately is worth the few minutes it costs. "How many write operations did agents perform against production this week, and how many were later reversed by a human" is a governance question that arrives eventually, and it is easier to have the answer than to build it retrospectively.</p>
<p>There is a design point underneath this. Every piece of work you move out of the model's prose and into a tool call becomes deterministic, testable and auditable. Tools are where an AI system stops being a text generator and starts being a system you can operate.</p>
<h2 id="the-loop-is-not-a-fourth-layer-it-is-a-change-of-denominator">The loop is not a fourth layer, it is a change of denominator</h2>
<p>A single agent run is not one model call. It is a plan, some retrieval, several tool calls, more model calls to decide what to do with the results, and a final answer.</p>
<p>The conventions handle the shape: <code>invoke_agent {gen_ai.agent.name}</code>, <code>invoke_workflow {gen_ai.workflow.name}</code>, <code>plan {gen_ai.agent.name}</code>, with <code>gen_ai.agent.id</code> and <code>gen_ai.agent.version</code> identifying who did what.</p>
<p>The trap is the unit of measurement. Every metric you instinctively reach for is per call, and per call is now the wrong denominator. What you want is cost per completed task rather than per call, steps per task as a distribution rather than an average (the tail is where the loops live), the termination reason, and the human intervention rate. A run that terminates on its step ceiling is a failure that reports as a completion.</p>
<p>Then there is the arithmetic. If each step were independently 90% reliable, five steps in a chain would be 0.9⁵, roughly 59%. That is a thought experiment rather than a measurement, and real systems are not that clean in either direction: errors are correlated, but many are also recoverable, because a validator rejects or a tool throws. I have <a href="https://www.linkedin.com/pulse/ai-understand-before-you-apply-buy-jeff-wouters-d9y1e">written about the compounding version of this before</a>. The point that survives the caveats is that reliability does not hold steady across a chain, and the output of one step becomes the input of the next, so a subtle error does not stay put. It gets treated as ground truth downstream.</p>
<p>So when the final answer is wrong, where do you look? Each step looks clean on its own. There is no stack trace, just a confident broken result.</p>
<h2 id="what-that-actually-looks-like-in-a-trace">What that actually looks like in a trace</h2>
<p><strong>A worked example, and worth being explicit that it is one.</strong> The trace below is constructed to show the shape of the failure, not transcribed from an incident I sat through. The attribute names, the span structure and the failure mode are real and come from the specification. The run is not.</p>
<p>Take a document triage agent that returns the wrong routing decision on a small fraction of cases. Nothing errors. Users notice because the wrong team receives the work.</p>
<p>The trace for one bad run:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code>invoke_agent triage-router            2.9s   ok
├─ plan triage-router                 0.4s   ok
├─ retrieval policy-index             0.2s   ok   app.retrieval.doc_count=0
├─ chat                               1.1s   ok   gen_ai.prompt.version=v7
│                                                 finish_reasons=[&quot;stop&quot;]
├─ execute_tool lookup_owner          0.1s   ok   app.tool.args_valid=true
└─ execute_tool assign_queue          0.3s   ok   app.tool.write=true
</code></pre>
<p>Every span is green. The answer is wrong. The only anomaly is <code>app.retrieval.doc_count=0</code>: the retrieval returned nothing above threshold, the model answered from its own weights anyway, and the tool faithfully assigned the queue it was told to.</p>
<p>Without that one custom attribute, the trace says a healthy agent did five healthy things. With it, the diagnosis is a retrieval threshold, not a model problem, and it takes minutes rather than a day of arguing about the prompt.</p>
<p>That is the argument for logging inputs as well as outputs at every hop. Not because you want to read them, but because when you need to walk backwards, output-only telemetry tells you the run happened and nothing about why it went wrong.</p>
<p>Which leads to the problem underneath all of this.</p>
<h2 id="you-cannot-evaluate-what-you-are-not-allowed-to-store">You cannot evaluate what you are not allowed to store</h2>
<p>To measure quality automatically you need the content. To protect people you must not keep the content. Both are true at once.</p>
<p>The specification is unambiguous: model instructions, user messages and model outputs are classed as sensitive, and instrumentations should not capture them without explicit opt-in. <code>gen_ai.input.messages</code>, <code>gen_ai.output.messages</code> and <code>gen_ai.system_instructions</code> are all opt-in for that reason.</p>
<p>Microsoft's Foundry SDK follows suit. Content recording is off unless you switch it on, via <a href="https://learn.microsoft.com/azure/foundry/observability/how-to/trace-agent-client-side#control-tracing-behavior-with-environment-variables"><code>OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT</code></a> (or the <code>Azure.Experimental.TraceGenAIMessageContent</code> switch in .NET), with a plain warning not to enable it in production unless your compliance requirements allow it. On that SDK you also need <code>AZURE_EXPERIMENTAL_ENABLE_GENAI_TRACING=true</code> before you get any GenAI spans at all, which is a confusing hour if nobody told you.</p>
<p><strong>Do not generalise that default across your stack, though.</strong> Microsoft's own <a href="https://learn.microsoft.com/azure/foundry/how-to/develop/langchain-traces">LangChain integration</a> documents content recording as <strong>enabled by default</strong>, and Agent Framework and Semantic Kernel emit traces automatically once tracing is on for the project, with a different set of variables again. "Off by default" is a property of a particular SDK, not of the ecosystem. Check the framework you actually run, in the environment you actually run it, rather than trusting a blog post about a different one.</p>
<p>Now the collision. Foundry's trace-based evaluation <a href="https://learn.microsoft.com/azure/foundry/observability/how-to/troubleshooting#trace-evaluation-issues">only reads spans where <code>gen_ai.operation.name</code> is <code>invoke_agent</code></a>, and Microsoft's wording is that if those spans have neither <code>gen_ai.input.messages</code> nor <code>gen_ai.output.messages</code>, the evaluators have no conversation content to score. No content, no automated quality score.</p>
<p>So the flag that protects your users is the same flag that blinds your quality monitoring. That is not a bug in anyone's product. It is the shape of the problem.</p>
<p><div class="figure-scroll" tabindex="0" role="group" aria-label="Diagram, scrollable"><svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 720 380" width="720" height="380" role="img" aria-labelledby="dl-t dl-d" style="max-width:100%;height:auto;display:block;margin:1.5rem 0"><title id="dl-t">The privacy and evaluation deadlock</title><desc id="dl-d">A four step cycle: scoring quality requires message content, message content is sensitive and off by default, so spans carry no content, so the evaluator returns no score.</desc><defs><marker id="dl-a" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"><path d="M0,0 L10,5 L0,10 z" fill="#0aa5c4"/></marker><marker id="dl-aw" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto-start-reverse"><path d="M0,0 L10,5 L0,10 z" fill="#d2593c"/></marker></defs><rect x="235.0" y="13.0" width="250" height="62" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="360" y="38" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="14" fill="currentColor" fill-opacity="1" text-anchor="middle" font-weight="600">You want a quality score</text><text x="360" y="59" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">groundedness, tool-call accuracy</text><rect x="435.0" y="159.0" width="250" height="62" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="560" y="184" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="14" fill="currentColor" fill-opacity="1" text-anchor="middle" font-weight="600">The evaluator needs content</text><text x="560" y="205" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">gen_ai.*.messages must be present</text><rect x="235.0" y="305.0" width="250" height="62" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="360" y="330" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="14" fill="currentColor" fill-opacity="1" text-anchor="middle" font-weight="600">Spans carry no content</text><text x="360" y="351" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">nothing to evaluate against</text><rect x="35.0" y="159.0" width="250" height="62" rx="4" fill="none" stroke="currentColor" stroke-opacity="0.28" stroke-width="1.5"/><text x="160" y="184" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="14" fill="currentColor" fill-opacity="1" text-anchor="middle" font-weight="600">Content is sensitive</text><text x="160" y="205" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="11" fill="currentColor" fill-opacity="0.62" text-anchor="middle" font-weight="400">opt-in, and off by default</text><line x1="470" y1="62" x2="545" y2="158" stroke="#d2593c" stroke-width="1.5" marker-end="url(#dl-aw)"/><line x1="545" y1="222" x2="470" y2="318" stroke="#d2593c" stroke-width="1.5" marker-end="url(#dl-aw)"/><line x1="250" y1="318" x2="175" y2="222" stroke="#d2593c" stroke-width="1.5" marker-end="url(#dl-aw)"/><line x1="175" y1="158" x2="250" y2="62" stroke="#d2593c" stroke-width="1.5" marker-end="url(#dl-aw)"/><text x="360" y="186" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="13" fill="#d2593c" fill-opacity="1" text-anchor="middle" font-weight="600">no way out</text><text x="360" y="206" font-family="-apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Helvetica, Arial, sans-serif" font-size="11" fill="currentColor" fill-opacity="1" text-anchor="middle" font-weight="400">without a deliberate trade</text></svg></div></p>
<p>What I would do in an environment with real data protection obligations, and I would want a DPIA rather than a blog post to settle it:</p>
<ul>
<li><strong>Enable content recording for a sampled slice, not the whole stream.</strong> You do not need every conversation to detect a regression.</li>
<li><strong>Send content-bearing telemetry to its own resource</strong>, with its own access control and its own retention period. Not the general workspace half of IT can query.</li>
<li><strong>Redact before export.</strong> Once it is in the pipeline it is in the backups, and unpicking that later is worse than gating it now.</li>
<li><strong>Keep the structural telemetry always on.</strong> Token counts, finish reasons, tool names, step counts, termination reasons and costs are far lower risk and answer most operational questions on their own.</li>
</ul>
<p>One correction to a claim you will see made often, including by me until someone checked it: structural telemetry is lower risk, not risk-free. <code>gen_ai.conversation.id</code> is a pseudonymous identifier, and pseudonymised data counts as personal data under GDPR wherever it can still be attributed to a person using other information you hold. In most deployments a conversation id can be joined back to a user, which is the case you should assume until someone demonstrates otherwise. Retrieved document identifiers can point at a case file about a named individual. Application Insights typically stamps rows with client city and country derived from IP, without you instrumenting anything, <a href="https://learn.microsoft.com/azure/azure-monitor/app/ip-collection">unless IP collection is disabled</a>. "No content" narrows the conversation with your DPO. It does not end it.</p>
<h2 id="what-breaks-and-what-does-not">What breaks, and what does not</h2>
<p>Your existing stack still works perfectly for availability and latency, for tool and API failures, for cost and token spend, for throughput and saturation, and for trace topology, which was designed for exactly this shape.</p>
<p>It goes quiet on correctness, where there is no signal and no threshold to alert on. On groundedness, meaning whether the answer came from your data or the model's memory. On tool selection, where the right function is called for the wrong reason. On task completion, where every step is green and the outcome is useless. And on drift, where this month's answers are worse than last month's and nothing errored in between.</p>
<p>Look back at the four layers and the same issue turns up each time. The truncated answer, the stale model version, the confidently summarised wrong document, the right tool with the wrong argument, the agent that hit its step ceiling and reported success. None of them is an availability failure. Every one is a judgement failure, and judgement does not throw exceptions.</p>
<p>The first list is monitoring. The second needs evaluation, which is a different discipline you add rather than a dial you turn.</p>
<h2 id="what-this-costs-and-what-you-would-actually-be-buying">What this costs, and what you would actually be buying</h2>
<p>The instrumentation is genuinely cheap. If you already run OpenTelemetry, the GenAI conventions are a set of attributes and a few new span types, not a new stack. Azure Monitor ingests them as ordinary spans and Application Insights has <a href="https://learn.microsoft.com/azure/azure-monitor/app/agents-view">an agent view</a> built on them.</p>
<p>The honest caveat, as of August 2026: <strong>no <code>gen_ai.*</code> attribute is marked Stable.</strong> The whole namespace is Development, it has <a href="https://github.com/open-telemetry/semantic-conventions-genai">moved into its own repository</a>, and it is iterating quickly. Adopt it, because a moving standard still beats everyone inventing their own names, but put a thin translation layer between the conventions and any dashboard you would be annoyed to rewrite.</p>
<p>Two things I would not pretend away.</p>
<p><strong>Ingestion is not free and this design is ingestion-heavy.</strong> Twenty spans for one user question, with inputs and outputs at every hop, into a backend billed per gigabyte. Microsoft says so directly in its own tracing guidance: trace data is subject to your retention settings and Azure Monitor pricing, and you should consider adjusting sampling rates or retention. Cost your ingestion before you enable content recording broadly, not after the first invoice.</p>
<p><strong>Evaluation is where the products earn their money.</strong> The vocabulary is standardised and free. Golden dataset curation, annotation queues, human review interfaces tied to specific spans, judge calibration, inter-annotator agreement, dataset versioning and online eval scheduling are not, and none of them is an attribute you can set. If you buy something in this category, buy it for that, and know that the tracing layer underneath it is a commodity you already own. Buying a product to emit spans is the part that does not make sense.</p>
<h2 id="making-it-survive-scale">Making it survive scale</h2>
<p><strong>Sample on outcome, and know what that costs you.</strong> Head sampling decides before the run has an outcome, so it discards the interesting runs at the same rate as the boring ones. What you want is tail sampling: keep everything that scored badly, errored, looped or made the user retry, and discard most of the clean successes. Be clear that this is a real architectural addition rather than a setting. Application Insights' own sampler is trace-ID-hash based and outcome-blind, and Microsoft states that full tail-based sampling <a href="https://learn.microsoft.com/azure/azure-monitor/app/application-insights-faq">is not currently supported</a>; doing it properly means an OpenTelemetry Collector with the <code>tail_sampling</code> processor in front of the exporter. This is the one place where "you already own the pipeline" stops being true.</p>
<p><strong>Keep prompt text out of metric dimensions.</strong> It belongs on a sampled span. As a metric label it is unbounded cardinality, which hurts ingest and query performance and shows up on the bill.</p>
<p><strong>Watch the cost of watching.</strong> Judging every response with a model adds a second non-deterministic system to the one you were trying to understand, and a bill that varies enormously with the judge model and rubric size. Judge a sample. Use cheap deterministic proxies for the rest: regeneration rate, abandonment, escalation to a human, and the edit distance between what the AI drafted and what actually got sent.</p>
<p>That last signal is a useful one and an imperfect one, so treat it carefully. If people heavily rewrite what the assistant produces, something is wrong regardless of what your groundedness score says. But it is confounded, because people edit for tone and house style as much as for correctness, and a low edit rate can mean the output is good or that users have stopped reading it. Make it a KPI and you will get paste-accept behaviour with the fixes made downstream where you cannot see them.</p>
<h2 id="where-to-start">Where to start</h2>
<p><strong>Get the structural attributes flowing.</strong> Operation name, provider, request and response model, token usage, finish reasons, conversation id. Low risk, immediate value, and no privacy discussion needed beyond confirming that conversation ids fall under your existing telemetry policy.</p>
<p><strong>Put one trace around the whole task.</strong> One conversation id and one root span from the user's question through every call, retrieval, tool and sub-agent. Without it you have a pile of unrelated spans and no way to ask whether the task worked.</p>
<p><strong>Version everything on the span.</strong> Prompt name and version, tool schema version, agent version. Five attributes, and the difference between diagnosing a regression quickly and arguing about it for a day.</p>
<p><strong>Promote the quiet failures.</strong> Truncation and content filtering, retrieval returning nothing above threshold, agents terminating on a step ceiling. These are already in your telemetry and nothing is looking at them.</p>
<p><strong>Add evaluation as a separate sampled track.</strong> Groundedness and tool-call accuracy on a slice of production traffic, plus a golden set you run on every prompt change. Decide up front who owns it, because this is where it usually stalls: the app team has the context, the platform team has the pipeline, and neither has the budget line.</p>
<p><strong>Instrument the humans.</strong> Override rate, escalation rate, regeneration rate.</p>
<h2 id="why-trust-an-ai-to-grade-an-ai">Why trust an AI to grade an AI</h2>
<p>Fair question to finish on. If AI fails convincingly, why would an AI evaluator be any better at spotting it?</p>
<p>Partly it is not, and anyone selling LLM-as-a-judge as solved is overselling. Judge models have the same non-determinism, and there is a real literature on their position bias, verbosity bias and preference for their own outputs.</p>
<p>The reason it is still worth doing is narrower than it first looks. Grading a specific answer against specific retrieved documents and a defined rubric is a bounded, structured task with the criteria fixed in advance. Deciding what the user should be told is an open judgement call. That is the same line I drew around <a href="https://www.linkedin.com/pulse/ai-understand-before-you-apply-buy-jeff-wouters-d9y1e">contract review</a>: excellent at checking whether a document contains a clause matching defined criteria, poor at telling you what your exposure is. Verification, not interpretation.</p>
<p>So use the judge for what it is good at, sample it, calibrate against human review often enough to notice drift, and never let it be the only thing between a broken system and your users.</p>
<p>Please don't misunderstand me. I am not arguing that AI systems are unobservable, or that you should wait for the conventions to settle before building anything. I am arguing that the instruments you already trust were built to answer a question that has quietly stopped being the important one, and nothing will tell you that, because the dashboard stays green either way.</p>
<p>Monitoring tells you the system ran. Evaluation tells you it was right. You need both, and you currently have one.</p>
<hr />
<p><em>Everything here reflects my own personal views and experience, not those of my employer or any organisation I'm affiliated with.</em></p>]]></content:encoded>
  </item>
  <item>
    <title>Branding your reMarkable Paper Pro</title>
    <link>https://jeffops.com/posts/2025/branding-your-remarkable-paper-pro/</link>
    <guid isPermaLink="true">https://jeffops.com/posts/2025/branding-your-remarkable-paper-pro/</guid>
    <pubDate>Tue, 21 Jan 2025 23:00:00 +0000</pubDate>
    <description>Enabling developer mode on a reMarkable Paper Pro, getting SSH working over WiFi, remounting the root filesystem so you can actually change anything, and putting it all back afterwards.</description>
    <category>remarkable</category><category>hardware</category><category>linux</category><category>ssh</category><category>toys</category>
    <content:encoded><![CDATA[<p><img alt="A reMarkable Paper Pro" src="https://jeffops.com/posts/2025/branding-your-remarkable-paper-pro/featured-image.jpg" /></p>
<p>Earlier this month I could not wait any longer, so I ordered a reMarkable Paper Pro. Two weeks later, at the short end of the quoted two to four, it arrived, and I had an afternoon free to start playing with it.</p>
<p>I will spare you the happy-happy-joy-joy reaction to the user experience. This post is about the branding, which means developer mode, SSH, and a writable root filesystem.</p>
<h2 id="step-one-basic-setup">Step one: basic setup</h2>
<p>Boot the device and go through the setup. I am not going to walk through it, because the experience is genuinely good and you do not need me for it.</p>
<h2 id="step-two-developer-mode">Step two: developer mode</h2>
<p>This is the important part, and the fiddly one.</p>
<ol>
<li>Open the left sidebar with the three horizontal lines at the top left of the screen.</li>
<li>Go to <strong>Settings</strong> at the bottom left. The <strong>General settings</strong> menu opens.</li>
<li>Under <strong>Paper tablet</strong>, tap <strong>Software</strong>.</li>
<li>Turn on the <strong>Advanced</strong> section with the toggle to its right.</li>
<li>Tap <strong>Developer mode</strong> and follow the instructions.</li>
<li>Tap <strong>back</strong> at the top left to return to <strong>General settings</strong>.</li>
<li>Under <strong>Personal</strong>, tap <strong>Account</strong>.</li>
<li>Tap <strong>Reset</strong> next to <strong>Factory reset</strong>. The device now performs the factory reset.</li>
</ol>
<p>One warning that cost me time. The instructions tell you a factory reset is required. In my case the device rebooted and did <strong>not</strong> perform one, which is misleading, because everything afterwards behaves as though it did until it suddenly does not. Steps 6 to 8 are how you make it actually happen.</p>
<h2 id="getting-the-username-password-and-ip-address">Getting the username, password and IP address</h2>
<p>You need these before you can connect to anything.</p>
<p>Three horizontal lines at the top left of the home screen, then <strong>Settings</strong>, then <strong>About</strong> under Help, then <strong>Copyright &amp; Licenses</strong>. The username, the password and the IP addresses the device is using are all on that screen.</p>
<h2 id="enabling-ssh-over-wifi">Enabling SSH over WiFi</h2>
<p>SSH over WiFi is off by default, which is the right default. The device ships with a utility to turn it on.</p>
<ol>
<li>Connect the device to your laptop with the USB-C cable.</li>
<li>SSH to <code>10.11.99.1</code>, which is the address it presents over USB.</li>
<li>Turn on SSH over WiFi:</li>
</ol>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code class="language-bash">rm-ssh-over-wlan<span class="w"> </span>on</code></pre>
<p>From then on you can do maintenance whenever the device is on the WiFi rather than tethered. When you are finished, turn it off again:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code class="language-bash">rm-ssh-over-wlan<span class="w"> </span>off</code></pre>
<h2 id="mounting-the-drive">Mounting the drive</h2>
<p>This is the step that will otherwise have you staring at permission errors that make no sense.</p>
<p>By default only <code>/home</code> is mounted writable. Everywhere else you will be told you cannot modify, delete or write, while the file permissions cheerfully tell you that you can. To anyone not steeped in Linux that reads as a bug rather than a mount option.</p>
<p>Connect over SSH and remount the root filesystem read-write:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code class="language-bash">mount<span class="w"> </span>-o<span class="w"> </span>remount,rw<span class="w"> </span>/</code></pre>
<p>Two things to hold in your head while it is mounted this way. Making <code>/</code> writable means you can put the device into a state where it no longer boots, so be deliberate about what you touch. And a reboot restores the normal mount, so if you lose your nerve, restarting it puts the safety back on.</p>
<h2 id="last-step-disabling-developer-mode">Last step: disabling developer mode</h2>
<p>Considerably more work than enabling it was.</p>
<ol>
<li>Connect the reMarkable Paper Pro to your computer with the USB cable, and leave it connected for the whole recovery process.</li>
<li>Activate recovery: long-press the power button for 30 seconds, then shortly after press it again for 3 seconds.</li>
<li>On your computer, open the desktop app.</li>
<li>In the side menu, go to <strong>Settings › Help › Recovery</strong>. Note that Recovery is only under Help <em>inside Settings</em>; it is not under Help in the toolbar, which is where I looked first.</li>
<li>Activate recovery on the device again with the same long-press then short-press.</li>
<li>In the desktop app, click <strong>Start recovery › Restore data › Continue</strong>. Leave the app running. It takes several minutes and the app shows progress.</li>
<li>When it finishes, click <strong>Close</strong>, close the desktop app, and turn the tablet on.</li>
</ol>
<p>And do not forget to actually do this. Leaving developer mode enabled on a device you carry around is the sort of thing you only regret once.</p>]]></content:encoded>
  </item>
  <item>
    <title>HowTo test Architecture in .NET with NetArchTest</title>
    <link>https://jeffops.com/posts/2023/howto-test-architecture-in-dotnet-with-netarchtest/</link>
    <guid isPermaLink="true">https://jeffops.com/posts/2023/howto-test-architecture-in-dotnet-with-netarchtest/</guid>
    <pubDate>Sun, 12 Nov 2023 23:00:00 +0000</pubDate>
    <description>Have you ever wondered how to ensure that your .NET code follows the architectural design and conventions that you have chosen? Do you want to avoid the common pitfalls of violating the principles of separation of concer</description>
    <category>dotnet</category><category>netarchtest</category><category>architecture</category><category>testing</category><category>dev</category>
    <content:encoded><![CDATA[<p>Have you ever wondered how to ensure that your .NET code follows the architectural design and conventions that you have chosen? Do you want to avoid the common pitfalls of violating the principles of separation of concerns, dependency inversion, or layering? If so, then you might be interested in NetArchTest, a fluent API for .NET Standard that can enforce architectural rules in unit tests.</p>
<p>Recently I stumbled upon the NetArchTest package by Ben Morris.</p>
<h2 id="what-is-netarchtest">What is NetArchTest?</h2>
<p>NetArchTest is a .NET library that allows you to create tests that enforce conventions for class design, naming, and dependency in .NET code bases. It is inspired by the ArchUnit library for Java. It uses a fluid API that allows you to string together readable rules that can be used in test assertions. You can use it with any unit test framework and incorporate it into your build pipeline.</p>
<p>It is available as a NuGet package.</p>
<h2 id="how-to-use-netarchtest">How to use NetArchTest?</h2>
<p>The basic steps to use NetArchTest are:</p>
<ol>
<li>Select a set of types from a path, assembly, or namespace using the static <code>Types</code> class.</li>
<li>Filter the types using one or more predicates, such as <code>ResideInNamespace</code>, <code>HaveDependencyOn</code>, <code>ImplementInterface</code>, etc. You can chain the predicates using <code>And</code> or <code>Or</code> conjunctions.</li>
<li>Apply one or more conditions using the <code>Should</code> or <code>ShouldNot</code> methods, such as <code>BeSealed</code>, <code>BeAbstract</code>, <code>HaveNameStartingWith</code>, etc.</li>
<li>Obtain a result from the rule by using an executor, such as <code>GetTypes</code> to return the types that match the rule or <code>GetResult</code> to determine whether the rule has been met. The result will also return a list of types that failed to meet the conditions.</li>
</ol>
<p>Here are some examples of rules that you can create with NetArchTest:</p>
<p>Classes in the presentation layer should not directly reference repositories:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code class="language-csharp"><span class="kt">var</span><span class="w"> </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">Types</span><span class="p">.</span><span class="n">InCurrentDomain</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">That</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">ResideInNamespace</span><span class="p">(</span><span class="s">&quot;MyProject.Presentation&quot;</span><span class="p">)</span>
<span class="w">    </span><span class="p">.</span><span class="n">ShouldNot</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">HaveDependencyOn</span><span class="p">(</span><span class="s">&quot;MyProject.Data&quot;</span><span class="p">)</span>
<span class="w">    </span><span class="p">.</span><span class="n">GetResult</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">IsSuccessful</span><span class="p">;</span></code></pre>
<p>Classes in the data layer should implement <code>IRepository</code>:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code class="language-csharp"><span class="kt">var</span><span class="w"> </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">Types</span><span class="p">.</span><span class="n">InCurrentDomain</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">That</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">ResideInNamespace</span><span class="p">(</span><span class="s">&quot;MyProject.Data&quot;</span><span class="p">)</span>
<span class="w">    </span><span class="p">.</span><span class="n">Should</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">ImplementInterface</span><span class="p">(</span><span class="k">typeof</span><span class="p">(</span><span class="n">IRepository</span><span class="p">))</span>
<span class="w">    </span><span class="p">.</span><span class="n">GetResult</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">IsSuccessful</span><span class="p">;</span></code></pre>
<p>All the service classes should be sealed:</p>
<pre tabindex="0" role="group" aria-label="Code, scrollable"><code class="language-csharp"><span class="kt">var</span><span class="w"> </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">Types</span><span class="p">.</span><span class="n">InCurrentDomain</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">That</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">ImplementInterface</span><span class="p">(</span><span class="k">typeof</span><span class="p">(</span><span class="n">IService</span><span class="p">))</span>
<span class="w">    </span><span class="p">.</span><span class="n">Should</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">BeSealed</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">GetResult</span><span class="p">()</span>
<span class="w">    </span><span class="p">.</span><span class="n">IsSuccessful</span><span class="p">;</span></code></pre>
<h2 id="want-to-read-more-on-its-origin">Want to read more on its origin?</h2>
<p>Ben's written a nice blog post about how he came to write this package. Especially if you're interested in what motivates people, and the path they've walked, take a look at his blog post.</p>
<h2 id="why-use-netarchtest">Why use NetArchTest?</h2>
<p>NetArchTest can help you to:</p>
<ul>
<li>Maintain the consistency and quality of your code base over time.</li>
<li>Avoid the need for manual code reviews or static analysis tools that may not capture your specific architectural requirements.</li>
<li>Create a self-testing architecture that can be verified by automated tests.</li>
<li>Communicate and document your architectural design and conventions through code.</li>
</ul>
<h2 id="conclusion">Conclusion</h2>
<p>NetArchTest is a powerful and easy-to-use library that can help you to test your architecture in .NET. It can help you to enforce the rules and conventions that you have chosen for your code base and avoid the common pitfalls of architectural decay. You can use it with any unit test framework and integrate it into your build pipeline. If you are interested in learning more about NetArchTest, you can check out its GitHub repository or its NuGet page. Happy testing!</p>]]></content:encoded>
  </item>
</channel>
</rss>
