<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Justin R. Greenbaum]]></title><description><![CDATA[Leadership, AI, and the systems behind reliable execution.]]></description><link>https://writing.justingreenbaum.com</link><image><url>https://substackcdn.com/image/fetch/$s_!-OJJ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9353dbd-7265-40d5-a2e1-b4245008ec4d_512x512.png</url><title>Justin R. Greenbaum</title><link>https://writing.justingreenbaum.com</link></image><generator>Substack</generator><lastBuildDate>Fri, 28 Aug 2026 18:59:40 GMT</lastBuildDate><atom:link href="https://writing.justingreenbaum.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Justin R. Greenbaum]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[justinrgreenbaum@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[justinrgreenbaum@substack.com]]></itunes:email><itunes:name><![CDATA[Justin R. Greenbaum]]></itunes:name></itunes:owner><itunes:author><![CDATA[Justin R. Greenbaum]]></itunes:author><googleplay:owner><![CDATA[justinrgreenbaum@substack.com]]></googleplay:owner><googleplay:email><![CDATA[justinrgreenbaum@substack.com]]></googleplay:email><googleplay:author><![CDATA[Justin R. Greenbaum]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Product Team and the Escalation Team Should Be Best Friends]]></title><description><![CDATA[Support knows first. Why the signal dies three times on its way upstream, and the cheapest repair I know.]]></description><link>https://writing.justingreenbaum.com/p/the-product-team-and-the-escalation</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/the-product-team-and-the-escalation</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Wed, 15 Jul 2026 14:14:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!XNVo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XNVo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XNVo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg 424w, https://substackcdn.com/image/fetch/$s_!XNVo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg 848w, https://substackcdn.com/image/fetch/$s_!XNVo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!XNVo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XNVo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5288946,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/207152207?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XNVo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg 424w, https://substackcdn.com/image/fetch/$s_!XNVo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg 848w, https://substackcdn.com/image/fetch/$s_!XNVo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!XNVo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc82e9ebe-d573-433d-86fe-defd5d97c3fa_4210x2807.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Two teams, one signal. Artist's rendering.</figcaption></figure></div><p>The worst cases in the company came to my teams. The billing error that found its way to a reporter. The outage with a regulator already asking questions. The complaint that arrived with a lawyer attached. I ran the queues where customer problems became company problems, and my people were the interface to everyone who could actually fix things: engineering, Legal, PR, compliance.</p><p>Here is what that seat teaches you. Support knows first.</p><p>When customers started cutting the cord, they did not announce it in a survey. They told our agents, in plain language, on recorded lines and social media outlets, years before the trend had a name on an earnings call. Every product defect, every pricing mistake, every process that quietly burned trust arrived at the edge of the company long before it arrived anywhere else. The most verified account of how the product actually failed sat in the queue, told by the people living it, in their own words.</p><p>And almost none of it became a decision.</p><p>I want to be careful here, because the easy version of this essay blames someone. The executives who would not listen. The finance team and their metrics. Nobody needs that essay, and it would not be true. The honest version is that the signal died of natural causes, three separate times, on its way upstream.</p><p><strong>The first break: we deleted the evidence at the source.</strong></p><p>Support was a cost center, so support was measured on cost. Contacts per customer. Cost per contact. Deflection rate. For a decade the industry celebrated every contact that never happened. And here I owe my old teams some precision, because we were not managing to those numbers blindly. We ran root cause analysis on the worst patterns as a matter of routine. We packaged what we were seeing and shared it across the business, with any team that would take the meeting. (Most did. Attendance was never the problem.) That awareness work was constant and it was real. But the operating model paid for deflection anyway, and every contact we deflected was a data point destroyed. The customer with the confusing bill, the one who could not make the feature work, the one calling for the third time about the same thing. Those were the exact contacts carrying the intelligence, and we built programs to make them go away.</p><p>Part of why the cost center label stuck is that the queue&#8217;s real value was invisible. My teams caught complaint patterns before they became regulatory filings. We caught defects before they landed on Legal&#8217;s desk. But an avoided fine never shows up on a P&amp;L. So the only number anyone could see was the cost, and organizations manage what they can see.</p><p><strong>The second break: we compressed whatever survived.</strong></p><p>Even the contacts that happened were recorded, sampled, and scored on how the agent behaved rather than on what the customer said. I traded notes on this earlier in the week with Gerti Haxhiu, who runs contact center operations for a living and sees this from the provider side. His description was sharper than mine. Every call recorded, a couple percent sampled, the sample graded on agent performance. By the time anything reached the rooms where decisions got made, it had been compressed into a handle time and a satisfaction score. And as he put it, nobody kills a cash cow over a CSAT number.</p><p>That sentence explains a lot of corporate history. The rooms deciding whether to disrupt a profitable business almost never heard the customers describing their exits. They heard a green dashboard, and now and then a root cause read-out that everyone agreed was concerning before the agenda moved on. The choice they were actually making never appeared on a slide.</p><p><strong>The third break: whatever survived was orphaned.</strong></p><p>The sliver of signal that made it upstream arrived in rooms where nobody owned acting on it. I lived this one most directly, because moving signals across organizational boundaries was my actual job, and it was the hardest part of the job by a wide margin. Our root cause findings were often in the deck. People nodded. Nobody disputed them. Then the meeting ended and the finding belonged to no one. The escalation team could see the problem and could not own the fix. The product team could own the fix and could not see the problem. Between them sat a boundary, and the org chart did not assign anyone to carry things across it.</p><p>So the carrying got done by habit and by relationship, where it got done at all. One of our most senior executives had a habit I have thought about ever since. Every escalation that landed in his inbox went back out with the adjacent owners copied on it. My teams would solve the case and close it regardless. He was not asking anyone to fix anything. He wanted the owners to know what their corner of the business was doing to customers, every single time. Some just read them. The good ones got involved.</p><p>Which brings me to the title.</p><p>The best product decisions I ever watched happen did not start in a review meeting. They started with an engineer sitting next to an escalation rep for an afternoon, wearing a headset, listening. Not reading a summary. Listening. Something changes in a builder who hears a customer struggle with their work in real time. The roadmap conversation is different afterward, and it stays different.</p><p>And the owned path is not hypothetical, because we built one. A quiet team of five people whose entire job was carrying social and escalation trends directly to the product and engineering teams, delivered weekly, for each product. Nobody handed us that lane. We created it, because the bridge the org chart would not draw still needed to exist.</p><p>If I ran a product organization today, I would start there and go further. Engineers on the listening side of the queue on a regular rotation, hours not minutes. A named, standing seat for the escalation team in roadmap reviews, with airtime that does not depend on the month&#8217;s crisis. One owned path for verified failure evidence, a specific person who receives it and decides where it goes. And that evidence reported with the same rigor the company applies to revenue, because it is the same kind of information. It tells you where the money is going to stop.</p><p>This is on my mind because every company is right now deciding what AI does to its support organization. The market is finally waking up to the queue as an intelligence asset, and mostly waking up in one direction: revenue. New platforms mine support interactions for training opportunities, services engagements, expansion signals. The direction is right, and some of it is genuinely well built. <a href="https://www.kahunalabs.ai/">Kahuna Labs</a> launched a module this week that puts a human review step between the pattern and the routing, which is the right instinct. Revenue is the first version of this signal a CFO can see, so revenue is where the category started. That is fair.</p><p>But look at where these signals route. Sales. Success. Services. The fix direction still has no owner. Product is still not in the room. And there is a fragile property worth protecting along the way. The queue&#8217;s evidence is honest for one reason only: customers tell support the truth because they believe they are talking to help. Point the queue at upsell clumsily and the truth dries up. The asset everyone suddenly wants to mine exists because of trust, and trust does not survive being farmed.</p><p>If AI in support gets aimed only at deflection and expansion, we will rebuild the old blindness with better technology. Automate the handling. Keep the listening. And route what you hear to the people who can remove the cause, not just the people who can monetize the symptom.</p><p>The chain breaks three times. The edge hears it. The measurement compresses it. Whatever survives lands in a room where nobody owns acting on it. Nobody can fix that whole chain in a quarter. But the third break has the cheapest repair I know of, and it is two teams and a standing invitation.</p><p>The escalation team is holding the truest record in the company of how the product fails. The product team is the only group that can make that record matter. Introduce them, and then protect the friendship like it is infrastructure, because it is.</p><p>No blame, just physics.<br><br>Justin<br><br>P.S. The two in the header have never once disputed who owns a signal.</p>]]></content:encoded></item><item><title><![CDATA[Back in a Classroom]]></title><description><![CDATA[Twenty years of operating, then two weeks at MIT. What an operator notices from a student&#8217;s chair.]]></description><link>https://writing.justingreenbaum.com/p/back-in-a-classroom</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/back-in-a-classroom</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Wed, 08 Jul 2026 14:48:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!E9xZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!E9xZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!E9xZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg 424w, https://substackcdn.com/image/fetch/$s_!E9xZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg 848w, https://substackcdn.com/image/fetch/$s_!E9xZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!E9xZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!E9xZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg" width="1456" height="1165" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1165,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1470386,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/206053548?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!E9xZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg 424w, https://substackcdn.com/image/fetch/$s_!E9xZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg 848w, https://substackcdn.com/image/fetch/$s_!E9xZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!E9xZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e9c4053-8810-4506-ac02-08179bf6e80e_2959x2367.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This spring I spent two weeks at MIT Sloan, in the AI Executive Academy. Fifty leaders from nineteen countries. Building E62, badge on a lanyard, name tent on the desk. It was my first structured classroom in a long time.</p><p>I walked in with a running lab. There is a diagnostic pipeline on my own hardware that had already been through dozens of refinement runs by the time the program started, and I assumed that would make the fundamentals sessions feel like review. That assumption did not survive the first morning.</p><p>Here is the thing about building by feel. The bench teaches you what works. It does not always teach you why. I had learned models the way you learn a machine you own, by running it, breaking it, and watching what it does at two in the morning. The classroom handed me the theory underneath habits I already had, and the effect was like getting the wiring diagram for a house you have been rewiring in the dark. Nothing I knew was wrong. All of it got more useful once it had names.</p><p>For twenty years I was the person questions escalated to. In that room I was the one raising my hand. I want to be honest about how good that felt. There is a specific relief in sitting in a chair where you are allowed not to know, where the expected contribution is a question instead of an answer. Operating does not offer that chair. School does, and I had forgotten.</p><p>The room taught as much as the faculty. Everyone there was carrying a version of the same question, shaped by their industry, and listening to the rest of the room describe the strain AI puts on their organizations was its own seminar. The room described symptoms; I kept hearing the structures underneath. That is not a criticism of anyone in it. It is what twenty years does to your hearing, and it told me the work I left to build is aimed at something real.</p><p>One session stays with me. AI agents running a business simulation, and within a few rounds the agents had developed coordination problems I have watched human organizations produce for two decades. Handoffs nobody owned. Decisions waiting on decisions. The same failure geometry, arrived at faster. I went in expecting to learn about agents. I came out having watched the patterns I write about emerge in a system with no politics, no history, and no personalities to blame, which is about as clean as evidence gets.</p><p>And because the bench does not turn off just because you are in a chair, I spent the evenings building. By the end of the program I had made a bingo card for the cohort&#8217;s favorite jargon, a little readiness checker, and a toy that generated startup ideas from the week&#8217;s lecture themes. None of it mattered. All of it was the point. You can take the operator out of the garage for two weeks. The garage comes along.</p><p>The real build came home with me. In the weeks after the program I compiled the whole thing, twelve days, fourteen faculty, more than forty sessions, every framework with its source and every number with its citation, into a single HTML page, a format I borrowed from Andrej Karpathy. It lives on the lab NAS next to the pipelines. Nobody asked for it and nobody grades it. I built it because the organizations I write about keep their decisions and lose their reasons, and I was not going to do that to two weeks of learning. On this bench, if it mattered, it gets indexed.</p><p>Twenty years in, one of the most useful things I have done this year was sit down and be taught. The door is up. Some weeks, what is on the bench is homework.<br><br>JG</p>]]></content:encoded></item><item><title><![CDATA[What's on the Bench]]></title><description><![CDATA[Five machines, no cloud, and what owned hardware actually costs.]]></description><link>https://writing.justingreenbaum.com/p/whats-on-the-bench</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/whats-on-the-bench</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Thu, 02 Jul 2026 13:40:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4wMR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In the first Open Garage post I said the door is up and you can see what is on the bench. Fair enough. Here is the bench.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4wMR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4wMR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4wMR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4wMR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4wMR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4wMR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6342982,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/204653170?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4wMR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4wMR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4wMR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4wMR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee4e92ae-7355-40b4-b661-2f8ab52d757d_5738x3825.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Five machines. Two NVIDIA DGX Sparks that carry the scoring work for the diagnostic pipeline. An M2 Mac Studio that runs embeddings, text inference, and a small resident agent that never sleeps. An M3 Mac Studio that handles vision models and the largest things I run, because 192GB of unified memory will hold what a GPU will not. A NAS in the corner holding it all together with shared storage. Across the five nodes there are roughly 750 billion model parameters hosted and ready. The biggest single model is a 235B mixture-of-experts that lives on the M3. Nothing in the pipeline calls a cloud API to think. Cloud services collect source material. The reasoning happens here.</p><p>That is a choice, and it has reasons. The diagnostic work processes evidence about real companies, and that evidence stays on machines I own. The cost shape matters too. When inference is metered, every experiment carries a small tax, and the tax makes you hesitant. When the metal is yours, the marginal experiment is free, and you run more of them. The lab got better because trying things stopped costing anything but time.</p><p>Time is the honest price. Owned hardware bills you in attention.</p><p>A vision model spent a week this spring working through 31,000 RAW files from my photo archive at about eighteen seconds a file. A seven-day run has to survive whatever happens during seven days, so every pass in every pipeline writes checkpoints and resumes from the last one. That discipline is not optional at this duration. You also become your own IT department. When a node drops, there is no ticket to file. There is a tunnel chain instead: laptop to the M2 over Tailscale, M2 to the Spark over the LAN, so I can check on a run from anywhere with a phone signal. And each morning the resident agent reads the pipeline state and the cluster health and posts me a briefing, which is the closest thing the garage has to a shift report.</p><p>What the attention buys: 189 diagnostic pipeline runs so far, 178 of them clean production passes, thirty companies on the fleet, from defense primes to sneaker brands, every finding challenged by an adversarial skeptic pass before it earns a score. And one project that is pure garage: a twenty-year photo archive, 102,630 files, 3.7 terabytes, indexed by a seven-pass pipeline into a single 60MB index I can search by concept.</p><p>The strangest thing on the bench sits at the intersection of those two workloads. The failure-mode taxonomy is encoded as text embeddings, and the photo index can be searched with them, which means I can ask twenty years of photographs for images that look like Responsibility Compression. That query should not work. It does, and some of the matches are unsettling.</p><p>The lab is not the practice. The practice is diagnosis, and it would exist on rented compute if it had to. But one person can run a fleet because the fleet is downstairs, checkpointed, and reporting for duty every morning. The door stays up.<br><br>JG</p>]]></content:encoded></item><item><title><![CDATA[The Role Ended. The Question Didn’t.]]></title><description><![CDATA[Why I left twenty years inside a Fortune 30 company, and what Open Garage is for.]]></description><link>https://writing.justingreenbaum.com/p/the-role-ended-the-question-didnt</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/the-role-ended-the-question-didnt</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Mon, 15 Jun 2026 14:31:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!J9Dt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!J9Dt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!J9Dt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg 424w, https://substackcdn.com/image/fetch/$s_!J9Dt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg 848w, https://substackcdn.com/image/fetch/$s_!J9Dt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!J9Dt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!J9Dt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:11967821,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/202132032?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!J9Dt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg 424w, https://substackcdn.com/image/fetch/$s_!J9Dt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg 848w, https://substackcdn.com/image/fetch/$s_!J9Dt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!J9Dt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60caa1cb-2071-4932-abf4-5abc3c2ce175_5911x3944.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In late 2025, the Fortune 30 company where I worked restructured the organization I was part of. By the end of December, my role no longer existed. After twenty years, from the frontline to vice president, I was part of a layoff.</p><p>That is the honest answer to the first question people ask, so I am putting it at the top. I did not quit in a blaze of conviction. The role ended, and that part was not my decision. The restructuring did offer me another role. It was a reasonable role. It pointed away from where I was going, so I turned it down. That part was my decision. This section is about that decision and what it is becoming.</p><p>For twenty years I worked in the parts of the operation where failure was public and ownership was unclear. Social media. Regulatory escalations. Executive complaints. Privacy and accessibility response. Customer security assurance. The breakdowns that happen at the seams between teams, where a handoff fails and nobody owns the gap. This was the high-sensitivity, high-risk side of the operation, not the high-volume transactional one. A mistake here did not stay internal.</p><p>Here is what that taught me, and it took most of those twenty years to see it clearly. Organizations almost always know something is wrong. The teams feel it. They compensate for it. They build workarounds around it. What they do not have is a name for it. Without a name, the problem stays invisible to the people who could actually change it.</p><p>I sat in rooms with very good consulting firms. They did competent work. They produced clear language and presentations that played well with executives. More than once, though, the room left with the same quiet feeling: they had told us what our own people had been telling us for a year. They gave us vocabulary for the symptom. They did not name the structure underneath it. And the structure is where the problem lives.</p><p>There was a second pattern, harder to watch. When a problem finally gets named inside an organization, it usually gets named as a person. Someone underperformed. Someone dropped the handoff. That naming is almost always wrong, and it is always expensive. The handoff did not fail because someone was careless. It failed because no structure made anyone responsible for it. Naming the person ends the conversation. Naming the structure starts a better one.</p><p>So when the role ended, I started building the thing I had spent twenty years wishing existed. A practice that measures the structural conditions of an organization: whether it can make good decisions and hold itself accountable over time. The practice is Decision and Responsibility Infrastructure. The method is Coherence. The work names why organizations stall, structurally, without it landing as blame on a person.</p><p>I want to be precise about the timeline, because precision matters here. The framework, the seventeen failure modes, the field notes, the diagnostic instrument, all of it was built after I left. The twenty years gave me the observations. They did not give me the framework. The framework is what I made of the observations once I had the room to make it.</p><p>That is what this section is for. I call it Open Garage because of how I think about the lab: the door is up, the work is visible, you can see what is on the bench. The work as it actually looks while it is still in progress. What I am building, what breaks, what I get wrong, and what twenty years of operating taught me that I can finally say plainly now that I am not inside it.</p><p>The rest of this publication carries the framework. The Coherence Record covers the instrument. The Lexicon names the patterns, one at a time. Open Garage carries the person doing the work. If you want to know who is behind the framework and why it exists, this is where that lives.</p><p>I am not going to pretend the transition has been clean. It has not. But the question I spent twenty years circling is still the question. Why do good organizations, full of capable people, stall? I have a better answer now than I did when I had a title. I am going to build the rest of that answer here, with the door up.<br><br>JG</p>]]></content:encoded></item><item><title><![CDATA[The Coherence Record, Edition 6]]></title><description><![CDATA[What Minimum Viable Governance Looks Like From the Inside]]></description><link>https://writing.justingreenbaum.com/p/the-coherence-record-edition-6</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/the-coherence-record-edition-6</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Thu, 16 Apr 2026 18:53:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4ujM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Justin R. Greenbaum<br>Greenbaum Labs<br>April 2026<br></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CZCx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf158eb0-6d07-4278-b760-10a5e90b3561_5333x3555.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CZCx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf158eb0-6d07-4278-b760-10a5e90b3561_5333x3555.jpeg 424w, https://substackcdn.com/image/fetch/$s_!CZCx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf158eb0-6d07-4278-b760-10a5e90b3561_5333x3555.jpeg 848w, https://substackcdn.com/image/fetch/$s_!CZCx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf158eb0-6d07-4278-b760-10a5e90b3561_5333x3555.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!CZCx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf158eb0-6d07-4278-b760-10a5e90b3561_5333x3555.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CZCx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf158eb0-6d07-4278-b760-10a5e90b3561_5333x3555.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/af158eb0-6d07-4278-b760-10a5e90b3561_5333x3555.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9429660,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/194419065?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf158eb0-6d07-4278-b760-10a5e90b3561_5333x3555.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CZCx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf158eb0-6d07-4278-b760-10a5e90b3561_5333x3555.jpeg 424w, https://substackcdn.com/image/fetch/$s_!CZCx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf158eb0-6d07-4278-b760-10a5e90b3561_5333x3555.jpeg 848w, https://substackcdn.com/image/fetch/$s_!CZCx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf158eb0-6d07-4278-b760-10a5e90b3561_5333x3555.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!CZCx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf158eb0-6d07-4278-b760-10a5e90b3561_5333x3555.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This edition is different from the ones that came before.<br><br>Editions 1 through 5 were build logs. They documented the construction of an instrument: scoring architectures, reproducibility testing, prompt hardening, fleet operations. The question behind every edition was whether the system could measure what it claimed to measure. Those questions are not finished. But this edition pauses the build to follow a different thread.</p><p>On March 27, <a href="https://mitsloan.mit.edu/faculty/directory/nick-van-der-meulen">Dr. Nick van der Meulen</a>, a research scientist at <a href="https://cisr.mit.edu/">MIT CISR</a>, presented his work on digital business transformation at the MIT AI Executive Academy. One week earlier, he had published a research briefing titled <a href="https://cisr.mit.edu/publication/2026_0301_GenAIGovernance_VanderMeulenJewerLevallet">&#8220;Minimum Viable Governance for Generative AI&#8221;</a> (MIT CISR Research Briefing, Vol. XXVI, No. 3, March 2026). It was his newest piece. Four characteristics of governance designed for a world where the technology transforms every eighteen months: structurally agile, trustworthy by design, integrated end-to-end, opportunity-sensitive.</p><p>I was in that room. During the session, I asked about a pattern I have seen repeatedly: authority that exists on paper but requires so much lateral alignment to execute that nobody actually owns the decision. He called it &#8220;very recognizable.&#8221; He connected it to what he calls organizational scar tissue: rules put in place because one person made one mistake, now applied to everyone forever. The conversation continued over lunch, and I described the diagnostic framework, the seventeen failure modes, the scoring pipeline, the center-edge documentation methodology.</p><p>I am not writing this to claim validation. I am writing it because his research and this project&#8217;s taxonomy are looking at the same problem from two altitudes. He is mapping what good looks like: the characteristics of governance that works. <a href="https://www.dripractice.com/start-here">The Coherence framework</a> maps what broken looks like: the structural conditions that emerge when those characteristics are absent. The two are complements. And the space between them is where the language lives.</p><h3>The Language Is the Contribution</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2LxK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7c66403-d3ae-4a25-8023-61e025735ee1_5281x3521.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2LxK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7c66403-d3ae-4a25-8023-61e025735ee1_5281x3521.jpeg 424w, https://substackcdn.com/image/fetch/$s_!2LxK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7c66403-d3ae-4a25-8023-61e025735ee1_5281x3521.jpeg 848w, https://substackcdn.com/image/fetch/$s_!2LxK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7c66403-d3ae-4a25-8023-61e025735ee1_5281x3521.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!2LxK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7c66403-d3ae-4a25-8023-61e025735ee1_5281x3521.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2LxK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7c66403-d3ae-4a25-8023-61e025735ee1_5281x3521.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f7c66403-d3ae-4a25-8023-61e025735ee1_5281x3521.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:11015866,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/194419065?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7c66403-d3ae-4a25-8023-61e025735ee1_5281x3521.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!2LxK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7c66403-d3ae-4a25-8023-61e025735ee1_5281x3521.jpeg 424w, https://substackcdn.com/image/fetch/$s_!2LxK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7c66403-d3ae-4a25-8023-61e025735ee1_5281x3521.jpeg 848w, https://substackcdn.com/image/fetch/$s_!2LxK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7c66403-d3ae-4a25-8023-61e025735ee1_5281x3521.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!2LxK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7c66403-d3ae-4a25-8023-61e025735ee1_5281x3521.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The Coherence project is built on a single premise: before you can measure organizational coherence, you need language that names what you are measuring. Before you can diagnose failure modes, you need vocabulary that makes those modes recognizable. The seventeen failure modes are not a scoring system. They are a naming system. FM-01, Responsibility Compression, describes something every person who has worked inside a scaling organization has felt. They felt it, adjusted to it, compensated for it, and could not name it. Without a name, it is invisible. You cannot address what you cannot see, and you cannot see what you cannot name.</p><p>The same insight surfaces in van der Meulen&#8217;s MVG work. He opened his session with a claim that landed harder than any framework or quadrant: shared vocabulary and strategic focus are the two prerequisites for transformation progress. Without shared vocabulary, &#8220;AI&#8221; means something different to every person in the room. &#8220;Transformation&#8221; is a word people nod at and define privately. &#8220;Governance&#8221; is either a reassurance or a threat, depending on who hears it.</p><p>This is what van der Meulen&#8217;s research and this project share as a foundational commitment: the belief that structural conditions must be named before they can be changed. His vocabulary (MVG, organizational explosions, silos and spaghetti, Future Ready) gives organizations language for where they are and where they need to go. The Coherence framework&#8217;s vocabulary (failure modes, field notes, the Triangle) gives organizations language for what is preventing them from getting there. The research describes the destination. The diagnostic names the obstacles. Both require language first.</p><p>But the instrument&#8217;s deepest contribution may not be the scores it produces. It may be the vocabulary it gives people for naming what they already observe. Edition 5 ended with that question: &#8220;whether someone needs a pipeline to see these patterns, or just the right questions.&#8221; This edition is the answer. The pipeline validates the language. The language is what scales.</p><p>A diagnostic score requires infrastructure, compute, methodology, a practitioner. A name requires only recognition. Someone reads &#8220;Responsibility Without Authority&#8221; and thinks: that is what I have been living inside for two years. That recognition is the beginning of the diagnostic, whether or not the pipeline ever runs on their organization. The language is the instrument&#8217;s gift to the people who will never buy the service. And it is the entry point for the people who will.</p><h3>What Breaks When Governance Isn&#8217;t Structurally Agile</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4ujM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4ujM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4ujM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4ujM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4ujM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4ujM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg" width="1456" height="967" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:967,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1395019,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/194419065?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4ujM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4ujM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4ujM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4ujM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c56587b-eb82-4f6a-b87c-12f9e9b1fceb_2048x1360.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Van der Meulen&#8217;s first MVG characteristic is <em><strong>structural agility</strong></em>: governance that can adapt its own structure as the environment changes. Not flexibility in the casual sense. The ability to change the rules about who decides what, and how quickly those rules take effect, without convening a senate each time.</p><p>When this characteristic is absent, the Coherence framework names what appears.</p><p><a href="https://www.dripractice.com/fm/fm-01">FM-01, Responsibility Compression</a>, is the most persistent signal in the diagnostic pipeline. It is one of three foundational failure modes the taxonomy identifies as tier 1: structurally universal in organizations past a certain scale. The instrument detected FM-01 above threshold in fourteen of fifteen fleet entities. In practice, the tier 1 modes are present in every large organization the pipeline has measured. What varies is severity, not presence. Responsibility concentrates where authority does not. Senior roles hold decision power. Frontline teams absorb the consequences without the ability to change outcomes. In a structurally agile governance model, decision rights redistribute as conditions change. Without that agility, they calcify. The people closest to the problem lack the authority to act on it. The people with authority are too far from the problem to see it clearly. Compression is the predictable result.</p><p><a href="https://www.dripractice.com/fm/fm-03">FM-03, Responsibility Without Authority</a>, is the sharper version of the same condition. Someone is explicitly accountable. Their name is on the RACI chart. Their performance review includes the outcome. But they lack the organizational authority to influence that outcome. Van der Meulen had a line in his session that named this precisely: &#8220;You can have the most beautiful RACI chart in the world, but it&#8217;s not going to change anything fundamentally.&#8221; He is right. The chart assigns responsibility. It does not transfer power. When governance cannot restructure authority in response to shifting conditions, RACI becomes a documentation of servitude, not a mechanism of alignment.</p><p><a href="https://www.dripractice.com/fm/fm-06">FM-06, Exception Inflation</a>, completes the picture. Every exception that gets hard-coded into process rather than resolved structurally is a governance system losing agility. A VP approves one off-cycle purchase because the timeline demands it. Next quarter, an exception form exists. The quarter after that, the form requires three signatures. A year later, forty percent of purchases route through the exception path, and the exception path is now the slow one. The organization layered governance on top of governance instead of fixing the structural condition that generated exceptions in the first place. Van der Meulen calls these layers &#8220;organizational scar tissue.&#8221; The Coherence framework counts them. They accumulate. They slow the organization down. And they are structurally invisible to the people living inside them because each individual scar feels reasonable.</p><h3>What Breaks When Governance Isn&#8217;t Trustworthy by Design</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Nu7l!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618ef8dd-3504-45ce-9596-94652a3b77e6_6267x7834.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Nu7l!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618ef8dd-3504-45ce-9596-94652a3b77e6_6267x7834.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Nu7l!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618ef8dd-3504-45ce-9596-94652a3b77e6_6267x7834.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Nu7l!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618ef8dd-3504-45ce-9596-94652a3b77e6_6267x7834.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Nu7l!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618ef8dd-3504-45ce-9596-94652a3b77e6_6267x7834.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Nu7l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618ef8dd-3504-45ce-9596-94652a3b77e6_6267x7834.jpeg" width="1456" height="1820" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/618ef8dd-3504-45ce-9596-94652a3b77e6_6267x7834.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1820,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:17749718,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/194419065?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618ef8dd-3504-45ce-9596-94652a3b77e6_6267x7834.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Nu7l!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618ef8dd-3504-45ce-9596-94652a3b77e6_6267x7834.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Nu7l!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618ef8dd-3504-45ce-9596-94652a3b77e6_6267x7834.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Nu7l!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618ef8dd-3504-45ce-9596-94652a3b77e6_6267x7834.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Nu7l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618ef8dd-3504-45ce-9596-94652a3b77e6_6267x7834.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The second MVG characteristic is that <em><strong>governance must be trustworthy by design</strong></em>: not trust bolted on after the fact, but trust embedded in the structure. People comply with governance they believe is fair, useful, and responsive. They route around governance they believe is theater.</p><p>When this characteristic is absent, the first thing that surfaces is <a href="https://www.dripractice.com/fm/fm-04">FM-04, Metric Shadowing.</a> The official metrics still get reported. They look fine. But the people producing those metrics know they do not reflect what is actually happening. A customer satisfaction score stays high because the survey only reaches customers who completed a transaction, not the ones who abandoned. A project status is green because the definition of green was quietly redefined two quarters ago. The governance mechanism is technically functioning. The trust is gone. The numbers are correct and meaningless.</p><p><a href="https://www.dripractice.com/fm/fm-02">FM-02, Escalation Inversion</a>, follows. The escalation paths exist on paper. People know where to route a problem, who to flag, what to file. But when escalating is costly, slow, or reputationally risky, people stop doing it. They absorb problems at the edge instead. The issue gets quietly resolved, or it doesn&#8217;t, and the organization only learns of it when something public breaks. In a trustworthy system, escalation is a signal. In one where trust has eroded, escalation is treated as failure: the act of raising a problem carries more cost than the problem itself. Issues get absorbed rather than surfaced. The structural conditions that produced them remain.</p><p>This is the gap that the Coherence framework measures: the distance between what the governance system reports about itself and what the people inside it (and the customers outside it) actually experience. Trust by design means the governance system&#8217;s self-report is reliable. When it is not, the diagnostic finds the specific failure modes that explain why.</p><h3>What Breaks When Governance Isn&#8217;t Integrated End-to-End</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ifPw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5579025a-e5aa-44f2-a987-46300ed55981_5201x3467.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ifPw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5579025a-e5aa-44f2-a987-46300ed55981_5201x3467.jpeg 424w, https://substackcdn.com/image/fetch/$s_!ifPw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5579025a-e5aa-44f2-a987-46300ed55981_5201x3467.jpeg 848w, https://substackcdn.com/image/fetch/$s_!ifPw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5579025a-e5aa-44f2-a987-46300ed55981_5201x3467.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!ifPw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5579025a-e5aa-44f2-a987-46300ed55981_5201x3467.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ifPw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5579025a-e5aa-44f2-a987-46300ed55981_5201x3467.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5579025a-e5aa-44f2-a987-46300ed55981_5201x3467.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:19373993,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/194419065?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5579025a-e5aa-44f2-a987-46300ed55981_5201x3467.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ifPw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5579025a-e5aa-44f2-a987-46300ed55981_5201x3467.jpeg 424w, https://substackcdn.com/image/fetch/$s_!ifPw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5579025a-e5aa-44f2-a987-46300ed55981_5201x3467.jpeg 848w, https://substackcdn.com/image/fetch/$s_!ifPw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5579025a-e5aa-44f2-a987-46300ed55981_5201x3467.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!ifPw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5579025a-e5aa-44f2-a987-46300ed55981_5201x3467.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The third MVG characteristic is <em><strong>integration</strong></em>: governance that operates across organizational boundaries, not within them. Not integration as an IT project. Integration as a structural condition where governance mechanisms talk to each other, where a decision made in one unit is visible and actionable in another.</p><p>When this is absent, what appears is the condition van der Meulen&#8217;s research calls &#8220;silos and spaghetti.&#8221; When leaders in the room self-selected into quadrants, the poll aligned with his survey data: the majority placed themselves there. The majority condition of large organizations is dysfunction as normal. People surviving on heroics, compensating for fragmentation with personal effort, navigating workarounds that everyone knows about and nobody addresses.</p><p>The Coherence framework names this <a href="https://www.dripractice.com/fm/fm-05">FM-05, Normalized Workarounds</a>. It is the operational texture of silos and spaghetti. The workaround that started as a temporary bridge becomes the permanent road. The manual handoff between two systems that should be integrated. The spreadsheet that exists because the platform cannot do what the team needs. The person who holds the institutional knowledge of how things actually work, and whose departure would break the process.</p><p><a href="https://www.dripractice.com/fm/fm-07">FM-07, Coordination Decay</a>, is the structural driver underneath. As governance fragments across organizational boundaries, the coordination cost between units rises silently. More meetings. More alignment documents. More &#8220;quick syncs&#8221; that are not quick and do not sync. The governance technically exists in each unit. The space between units is ungoverned. The coordination decay is invisible in any single unit&#8217;s reporting. It is visible only to the people absorbing the cost of bridging the gap and to the customers at the far end of it.</p><p>Van der Meulen made the amplification point explicitly in his afternoon session: AI does not create these conditions. AI amplifies whatever it is pointed at. Good operational backbone, clean data, skilled people with decision rights: AI accelerates that. Silos and spaghetti with overworked heroics and messy data: AI pours gasoline on it. The governance integration question is structural. AI does not create it. AI only makes it urgent. The Coherence framework measures the structural conditions. MVG describes the governance response. The sequence matters: diagnose first, govern second.</p><h3>What Breaks When Governance Isn&#8217;t Opportunity-Sensitive</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vq9R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb45d713-b1d6-4edc-bc87-8d2fb19d1cfe_5932x3955.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vq9R!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb45d713-b1d6-4edc-bc87-8d2fb19d1cfe_5932x3955.jpeg 424w, https://substackcdn.com/image/fetch/$s_!vq9R!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb45d713-b1d6-4edc-bc87-8d2fb19d1cfe_5932x3955.jpeg 848w, https://substackcdn.com/image/fetch/$s_!vq9R!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb45d713-b1d6-4edc-bc87-8d2fb19d1cfe_5932x3955.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!vq9R!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb45d713-b1d6-4edc-bc87-8d2fb19d1cfe_5932x3955.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vq9R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb45d713-b1d6-4edc-bc87-8d2fb19d1cfe_5932x3955.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cb45d713-b1d6-4edc-bc87-8d2fb19d1cfe_5932x3955.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:13657009,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/194419065?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb45d713-b1d6-4edc-bc87-8d2fb19d1cfe_5932x3955.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vq9R!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb45d713-b1d6-4edc-bc87-8d2fb19d1cfe_5932x3955.jpeg 424w, https://substackcdn.com/image/fetch/$s_!vq9R!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb45d713-b1d6-4edc-bc87-8d2fb19d1cfe_5932x3955.jpeg 848w, https://substackcdn.com/image/fetch/$s_!vq9R!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb45d713-b1d6-4edc-bc87-8d2fb19d1cfe_5932x3955.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!vq9R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb45d713-b1d6-4edc-bc87-8d2fb19d1cfe_5932x3955.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The fourth MVG characteristic is <em><strong>opportunity-sensitivity</strong></em>: governance that does not just prevent bad outcomes but actively creates conditions for good ones. This is the hardest characteristic to measure because its absence looks like stability. Nothing goes wrong. Nothing remarkable happens. The organization operates within its constraints and does not notice that the constraints have become the strategy.</p><p>The Coherence framework approaches this through the Truth vertex. Truth measures the distance between what an organization says about itself and what is observable at the edges: customer experience, employee experience, market reality. An organization that is not opportunity-sensitive tells a story about innovation, growth, and transformation that does not match the observable reality. The center narrative describes ambition. The edge data describes maintenance.</p><p>This is compression, operating at the narrative level in the same way FM-01 operates structurally. The center compresses complexity into a story it can tell the board, the market, the workforce. The edge lives the uncompressed version. The gap between the two is measurable, and the instrument measures it. A high Truth score means the center-edge gap is narrow, and the story matches the experience. A low Truth score means the story and the experience have diverged. Neither score tells you what to do. Both tell you where to look.</p><p>An opportunity-sensitive governance model keeps the gap narrow by design. The governance mechanisms surface edge reality into center decision-making. Customer complaints reach product strategy. Employee experience data reaches organizational design. Market signals reach resource allocation. When the governance model is not opportunity-sensitive, those feedback loops degrade. The center narrative drifts from edge reality. The Truth score declines. The organization becomes, in van der Meulen&#8217;s quadrant framing, a candidate for the Integrated Experience trap, the &#8220;dopamine trail&#8221; where customer-facing metrics improve while the underlying structure deteriorates. Everything looks better. Nothing has changed.</p><h3>What&#8217;s Next</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_rOf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa22f4a-e730-4b1c-bd31-88f5b2d6e130_5401x3601.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_rOf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa22f4a-e730-4b1c-bd31-88f5b2d6e130_5401x3601.jpeg 424w, https://substackcdn.com/image/fetch/$s_!_rOf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa22f4a-e730-4b1c-bd31-88f5b2d6e130_5401x3601.jpeg 848w, https://substackcdn.com/image/fetch/$s_!_rOf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa22f4a-e730-4b1c-bd31-88f5b2d6e130_5401x3601.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!_rOf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa22f4a-e730-4b1c-bd31-88f5b2d6e130_5401x3601.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_rOf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa22f4a-e730-4b1c-bd31-88f5b2d6e130_5401x3601.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6fa22f4a-e730-4b1c-bd31-88f5b2d6e130_5401x3601.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:12953726,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/194419065?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa22f4a-e730-4b1c-bd31-88f5b2d6e130_5401x3601.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_rOf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa22f4a-e730-4b1c-bd31-88f5b2d6e130_5401x3601.jpeg 424w, https://substackcdn.com/image/fetch/$s_!_rOf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa22f4a-e730-4b1c-bd31-88f5b2d6e130_5401x3601.jpeg 848w, https://substackcdn.com/image/fetch/$s_!_rOf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa22f4a-e730-4b1c-bd31-88f5b2d6e130_5401x3601.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!_rOf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa22f4a-e730-4b1c-bd31-88f5b2d6e130_5401x3601.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The MVG paper and the Coherence framework are adjacent layers. His research maps the governance characteristics that make organizations adaptive. The Coherence framework maps the structural conditions that emerge when those characteristics are absent. The mapping between them is specific: each MVG characteristic, when missing, produces identifiable, nameable failure modes.</p><p>This edition is the first attempt to bridge the two explicitly. The citations are a practitioner showing where a research tradition and twenty years of operational experience land on the same problems.</p><p>The language distribution work begins now. Each failure mode is a standalone piece of content. Each one names something people recognize but have not had a word for. The taxonomy (seventeen failure modes, twenty-one field notes, the Coherence Triangle) came first. It came from twenty years inside organizations where these patterns had no names. The pipeline was built afterward, to prove these conditions exist in the wild at scale, across sectors, without needing a client engagement or putting a former employer on the line. The language was always the point. The pipeline is the evidence.</p><p>Getting the vocabulary into circulation is the next phase. The pipeline will continue to run. The fleet will grow. The instrument will sharpen. But the vocabulary does not need the pipeline to travel. It needs only to be placed in front of people who have been waiting for it without knowing they were waiting.</p><p>Van der Meulen said something in his session that I keep coming back to, &#8220;The hard, unglamorous work of getting the conditions right for AI to actually help accelerate and transform the organization&#8230; that is not paid enough attention to.&#8221; He is right. And the first step in that work is naming the conditions. Not the aspirational conditions. The current ones. The ones that have been invisible because nobody had words for them.</p><p>Now they have names. Seventeen of them.</p><div><hr></div><p><em>The images in this edition are from my own library, shot on Leica. Everything in this project is built or sourced firsthand. The visuals are no exception.</em></p><div><hr></div><p><strong>References</strong></p><p>Van der Meulen, N., Jewer, J., and Levallet, N. &#8220;Minimum Viable Governance for Generative AI.&#8221; MIT CISR Research Briefing, Vol. XXVI, No. 3, March 2026.</p><p>Van der Meulen, N. and Ross, J.W. &#8220;Realizing Decentralized Economies of Scale.&#8221; MIT CISR, January 2023.</p><p>Van der Meulen, N. &#8220;Managing the Two Faces of Generative AI.&#8221; MIT CISR, September 2024.</p><p>Van der Meulen, N. &#8220;Bring Your Own AI: How to Balance Risks and Innovation.&#8221; MIT Sloan Management Review, October 2024.</p><p>Ross, J.W., Beath, C.M., and Mocker, M. <em>Designed for Digital: How to Architect Your Business for Sustained Success.</em> MIT Press, 2019.</p><p>Greenbaum, J. &#8220;The Coherence Record, Editions 1&#8211;5.&#8221; Greenbaum Labs, 2026.</p>]]></content:encoded></item><item><title><![CDATA[The Coherence Record, Edition 5]]></title><description><![CDATA[The instrument learned to question itself. Then it gained depth.]]></description><link>https://writing.justingreenbaum.com/p/the-coherence-record-edition-5</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/the-coherence-record-edition-5</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Tue, 17 Mar 2026 20:55:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!owp3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Justin R. Greenbaum | Founder, Greenbaum Labs<br>March 2026</p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!owp3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!owp3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg 424w, https://substackcdn.com/image/fetch/$s_!owp3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg 848w, https://substackcdn.com/image/fetch/$s_!owp3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!owp3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!owp3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg" width="1456" height="1164" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1164,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:757494,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/190888176?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!owp3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg 424w, https://substackcdn.com/image/fetch/$s_!owp3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg 848w, https://substackcdn.com/image/fetch/$s_!owp3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!owp3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F342c9956-5ddf-438d-bb91-ff14f03afb32_3675x2938.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>What&#8217;s Happened</h3><p>Edition 4 ended with the strongest claim in this project&#8217;s history: the instrument is reproducible. Zero standard deviation. Same entity, same score, every time. Finding-derived scoring replaced the LLM&#8217;s opinion with deterministic computation. The pipeline was grounded.</p><p>That was published on March 5. By March 13, eight days later, the project changed shape again.</p><p>134 runs in the ledger now. This is what the last eight days produced.</p><h3>The Prompts Weren&#8217;t Good Enough</h3><p>Edition 4 proved the scoring architecture was sound. It did not prove the prompts were.</p><p>Con-Hotel&#8217;s original run, run 079, scored with 33% skeptic throughput. Two of six findings survived debate. The other four were rejected. The rejected findings followed a consistent pattern: the agent had decided what score felt right, then went looking for evidence to justify it. The findings read like conclusions wearing an evidence costume.</p><p>This is the same failure pattern twenty-seven runs had eliminated from the scoring architecture, the model generating an opinion instead of computing from evidence. Fixed in the formula. Not fixed in the prompt.</p><p>Two changes:</p><p>First, Evidence Discipline blocks were added to the truth and authority scorer prompts. These are structural constraints, not suggestions. The prompt now explicitly names the failure pattern, &#8220;starting from a conclusion and working backward,&#8221; and forbids it. It requires each finding to be built from cited evidence: specific claims, specific observations, specific scope. The finding follows the evidence. Not the other way around.</p><p>Second, the authority few-shot examples were rewritten. The old examples were abstract. They led the model to produce vague, general findings that sounded analytical but said nothing specific enough to survive the Skeptic. The new examples follow a progression: BAD (vague, unsupported), STILL BAD (specific but backward, conclusion first), GOOD (evidence first, finding emerges from the data). Each example is led by a specific customer quote, not a category label.</p><p>Con-Hotel rescore with the hardened prompts: 83% skeptic throughput. Five of six findings sustained. Triple-blind validation confirmed deterministic, 0.000 standard deviation.</p><p>The prompts are now committed to the pipeline repo. The same codebase that runs the fleet.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!y-BL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaa2864-1e93-484c-9b4a-37cfab8ab809_4184x5230.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!y-BL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaa2864-1e93-484c-9b4a-37cfab8ab809_4184x5230.jpeg 424w, https://substackcdn.com/image/fetch/$s_!y-BL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaa2864-1e93-484c-9b4a-37cfab8ab809_4184x5230.jpeg 848w, https://substackcdn.com/image/fetch/$s_!y-BL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaa2864-1e93-484c-9b4a-37cfab8ab809_4184x5230.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!y-BL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaa2864-1e93-484c-9b4a-37cfab8ab809_4184x5230.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!y-BL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaa2864-1e93-484c-9b4a-37cfab8ab809_4184x5230.jpeg" width="1456" height="1820" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3eaa2864-1e93-484c-9b4a-37cfab8ab809_4184x5230.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1820,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:10337986,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/190888176?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaa2864-1e93-484c-9b4a-37cfab8ab809_4184x5230.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!y-BL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaa2864-1e93-484c-9b4a-37cfab8ab809_4184x5230.jpeg 424w, https://substackcdn.com/image/fetch/$s_!y-BL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaa2864-1e93-484c-9b4a-37cfab8ab809_4184x5230.jpeg 848w, https://substackcdn.com/image/fetch/$s_!y-BL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaa2864-1e93-484c-9b4a-37cfab8ab809_4184x5230.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!y-BL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaa2864-1e93-484c-9b4a-37cfab8ab809_4184x5230.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h3>Autoresearch</h3><p>The prompt changes that fixed Con-Hotel were designed by hand. The Skeptic&#8217;s rejections were analyzed, the failure pattern was identified, and the prompts were rewritten to prevent it. That worked. But it does not scale.</p><p>So the lab built the machine that does it.</p><p><a href="https://x.com/karpathy/status/2030371219518931079?s=20">Andrej Karpathy recently open-sourced</a> a similar concept, an agent that iterates on ML training code autonomously, running experiments while the operator sleeps. Different domain, same principle: structured experimentation at a pace no human can match. The Greenbaum Labs version optimizes diagnostic prompts against an adversarial debate mechanism.</p><p>Autoresearch is a harness that runs prompt experiments automatically. It takes a frozen extraction, same claims, same observations, and tests prompt variations against it, measuring skeptic throughput, scoring correctness, and reproducibility. Each experiment produces a structured log: what changed, what the scores were, whether the findings survived debate.</p><blockquote><p>Between March 10 and 12, the Sparks ran 60 experiments across two tracks.</p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7D7d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82befe71-2d78-4e65-88d0-8070aaf8f6a2_3048x1566.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7D7d!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82befe71-2d78-4e65-88d0-8070aaf8f6a2_3048x1566.png 424w, https://substackcdn.com/image/fetch/$s_!7D7d!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82befe71-2d78-4e65-88d0-8070aaf8f6a2_3048x1566.png 848w, https://substackcdn.com/image/fetch/$s_!7D7d!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82befe71-2d78-4e65-88d0-8070aaf8f6a2_3048x1566.png 1272w, https://substackcdn.com/image/fetch/$s_!7D7d!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82befe71-2d78-4e65-88d0-8070aaf8f6a2_3048x1566.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7D7d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82befe71-2d78-4e65-88d0-8070aaf8f6a2_3048x1566.png" width="1456" height="748" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82befe71-2d78-4e65-88d0-8070aaf8f6a2_3048x1566.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:748,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:487926,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/190888176?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82befe71-2d78-4e65-88d0-8070aaf8f6a2_3048x1566.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7D7d!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82befe71-2d78-4e65-88d0-8070aaf8f6a2_3048x1566.png 424w, https://substackcdn.com/image/fetch/$s_!7D7d!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82befe71-2d78-4e65-88d0-8070aaf8f6a2_3048x1566.png 848w, https://substackcdn.com/image/fetch/$s_!7D7d!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82befe71-2d78-4e65-88d0-8070aaf8f6a2_3048x1566.png 1272w, https://substackcdn.com/image/fetch/$s_!7D7d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82befe71-2d78-4e65-88d0-8070aaf8f6a2_3048x1566.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The scoring track ran 41 experiments. The baseline, Edition 4&#8217;s prompts before the Evidence Discipline changes, scored 0.058 on the optimization metric. The best variant scored 0.667. An 11.5x improvement. The key discovery wasn&#8217;t a single brilliant prompt. It was that asymmetric extraction limits, pulling 3 items from center sources and 5 from edge sources, outperformed symmetric limits. The edge is where the signal lives. Give the model more of it. The scoring track converged. Later experiments showed diminishing returns. The prompt space for scoring is largely explored. That means the current prompts are near the ceiling for what prompt engineering alone can achieve.</p><p>The extraction track ran 19 experiments. Baseline 0.117, best 0.450. A 3.8x improvement, with more room to run. Extraction is upstream of everything, the quality of claims and observations determines what the scorer has to work with. This track matters more than the scoring track in the long run. No convergence yet. The runway is open.</p><p>The harness is 7,165 lines of code. It runs unsupervised. It produces structured, reproducible experiment logs. And it confirmed something previously suspected but never measured: the pipeline&#8217;s reproducibility is near-perfect even as correctness varies. Fleet average reproducibility across cross-entity validation: 0.998. The instrument produces the same answer every time, even when the answer is wrong. That is the foundation. You fix correctness once and it stays fixed.</p><p>Prompt optimization is not craft anymore. It is experimental science. Hypothesis, test, measure, iterate. The machine questions the machine.</p><h3>Three Machines, One Night</h3><p>On the night of March 12, three jobs were launched across three machines.</p><p>Spark 2 rescored runs 070 through 084, the March 6 collection, all fifteen fleet entities, with the hardened prompts. M2 Studio rescored runs 049 through 064, the February 20 collection, the same fifteen entities, with identical prompts. Spark 1 ran Con-Hotel end-to-end, run 085, full extraction and scoring with the hardened prompts.</p><p>Everything completed overnight. Thirty rescores and one full pipeline run, across three machines, without intervention.</p><p>Six weeks ago the operational workflow required manual SSH checks on each machine, NAS mount debugging, and hand-verification of every flag in every launch command. Three failed Con-Hotel launches in a single session, wrong environment, wrong mode, missing flags, forced the construction of proper pre-flight checks.</p><p>The overnight run confirmed that the operational infrastructure caught up to the analytical infrastructure. The pipeline was reproducible weeks ago. The operations around it were not. Now they are.</p><h3>The Numbers</h3><p>Fleet rescore v4 results, March 6 collection (fifteen entities, hardened prompts):</p><p>Fleet average overall: 0.452. Range: 0.370 (Tech-Oscar) to 0.496 (Tech-Mike, Fin-Foxtrot). Truth average: 0.491. Authority average: 0.398.</p><p>For comparison, the original scores on this collection averaged 0.455 overall. The fleet moved down by 0.003. Effectively unchanged. But what moved underneath matters.</p><p>The biggest individual shifts:</p><p><strong>Fin-Delta</strong>: Truth rose from 0.455 to 0.554 (+0.099). The original scoring had suppressed a real signal, center-edge alignment on financial performance that the Evidence Discipline prompts now surface properly.</p><p><strong>Tech-Oscar</strong>: Overall dropped from 0.450 to 0.370 (&#8722;0.080). The original scoring had been generous. The hardened prompts derived a lower score from the specific findings that survived debate. The old scorer gave Tech-Oscar credit the evidence didn&#8217;t support.</p><p>The pattern is the same one from Edition 4&#8217;s rescore: the system corrects in both directions. Upward where signal was suppressed. Downward where opinion had inflated. Calibration, not drift.</p><p>Skeptic throughput across the v4 fleet: 48% (44 of 90 findings sustained). Tighter than the original runs&#8217; 57%. The hardened prompts produce fewer findings overall, but the ones that survive are better grounded. Quality over quantity. That is the design intent.</p><p>February 20 collection rescored with identical prompts: fleet average overall 0.443. Range: 0.386 (Fin-Echo) to 0.496 (Aero-Charlie). Comparable distribution, different collection date, same methodology. The scores are in the same band because the instrument is calibrated, not because the entities haven&#8217;t changed.</p><p>Run 085, Con-Hotel full end-to-end with hardened prompts: overall 0.460 (finding-derived). Truth 0.500, Authority 0.410. The extraction pulled 201 claims and 2,559 observations from the collection. Skeptic throughput was low, 20%, one finding sustained out of five. The Skeptic was harsh on this run, and the surviving finding was strong. The system working correctly. A low throughput rate with strong surviving findings is a more honest result than a high throughput rate with weak ones.</p><p>Against Con-Hotel&#8217;s original run 079 (overall 0.427), run 085 gained 0.033. A modest improvement. The real difference is in the evidence quality. The finding that survived debate in 085 is grounded in specific claims and observations. The findings that survived in 079 were vaguer. The score is similar. The confidence behind it is not.</p><h3>Continuity</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sA7z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d2d054-42b5-49eb-a270-bf8ad2ebecdb_1080x1350.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sA7z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d2d054-42b5-49eb-a270-bf8ad2ebecdb_1080x1350.jpeg 424w, https://substackcdn.com/image/fetch/$s_!sA7z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d2d054-42b5-49eb-a270-bf8ad2ebecdb_1080x1350.jpeg 848w, https://substackcdn.com/image/fetch/$s_!sA7z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d2d054-42b5-49eb-a270-bf8ad2ebecdb_1080x1350.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!sA7z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d2d054-42b5-49eb-a270-bf8ad2ebecdb_1080x1350.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sA7z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d2d054-42b5-49eb-a270-bf8ad2ebecdb_1080x1350.jpeg" width="1080" height="1350" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/36d2d054-42b5-49eb-a270-bf8ad2ebecdb_1080x1350.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1350,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1051220,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/190888176?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d2d054-42b5-49eb-a270-bf8ad2ebecdb_1080x1350.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sA7z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d2d054-42b5-49eb-a270-bf8ad2ebecdb_1080x1350.jpeg 424w, https://substackcdn.com/image/fetch/$s_!sA7z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d2d054-42b5-49eb-a270-bf8ad2ebecdb_1080x1350.jpeg 848w, https://substackcdn.com/image/fetch/$s_!sA7z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d2d054-42b5-49eb-a270-bf8ad2ebecdb_1080x1350.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!sA7z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36d2d054-42b5-49eb-a270-bf8ad2ebecdb_1080x1350.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Every edition of this record has contained the same line: &#8220;Continuity remains unscorable. One collection period.&#8221;</p><p>That line is retired.</p><p>The February 20 and March 6 collections, both scored under finding-derived v1 with hardened prompts, provide the two temporal points needed to compute Continuity. For every entity in the fleet, there are now two readings on the same instrument, separated by two weeks.</p><p>Two weeks is not much. But it is infinitely more than zero. And the structure is in place for the next collection, and the one after that.</p><p>What Continuity measures is trajectory. Truth and Authority are snapshots, where is this organization right now? Continuity asks: is it getting better, getting worse, or holding steady? Is the compression increasing? Is the center-edge gap widening or closing? Are the same failure modes persisting, or are new ones emerging?</p><p>The diagnostic becomes most valuable here. Not &#8220;here is your coherence&#8221; but &#8220;here is where your coherence is heading.&#8221; A snapshot tells you what to investigate. A trajectory tells you what is urgent.</p><p>The entity-level deltas between February 20 and March 6 are the next computation. The data exists. The methodology is identical. The analysis is coming.</p><h3>Findings</h3><p>Edition 4 appeared to be a conclusion: reproducibility proved, architecture locked, fleet scored. It was not a conclusion. It was the foundation for a harder set of questions.</p><p>The prompt deficiency was discovered by using the instrument, not by theorizing about it. The Skeptic&#8217;s 33% throughput on Con-Hotel meant four findings were rejected, and the rejection rationale pointed to the prompt, not the scorer. The instrument diagnosed its own inputs. That is a real feedback loop.</p><p>Automating prompt optimization appeared to be a shortcut. It is not. It is the only way to explore a space this large with any rigor. The autoresearch harness ran 60 structured experiments in three days, each one isolating a single variable. Combined with the 13 hand-tuned experiments that preceded it, 73 total experiments shaped the current prompts. No manual process achieves that. Not in three days, not in thirty. The machine is better at questioning itself than the operator is at questioning it.</p><p>A structural shift occurred in the last eight days. Editions 1 through 4 were construction: designing the architecture, fixing the scorers, debugging the pipeline. The relationship was builder to tool. With autoresearch, the instrument improved itself. The operator set the constraints, defined the metrics, launched the harness, and read the results. The machine ran the experiments independently. Builder to observer. That is the shift underneath the numbers. The discipline is in the constraints, not the keystrokes.</p><p>The authority data constraint, carried as a cap through Editions 3 and 4, was lifted without ceremony. Employee reviews appeared in the March 6 collection for all fifteen entities. The internal voice that was entirely absent from the edge data now exists. The authority scores did not move much. That raises a harder question than the cap did: the constraint was clear and honest. Now the data is present and the scores are similar, and the next step is determining whether the instrument is surfacing what the employee reviews contain or whether the extraction and scoring prompts need to be tuned to this new source type.</p><p>Continuity changes what the project is. The instrument has been taking snapshots. Snapshots are useful. They show where compression lives, where the center-edge gap is widest, where authority is concentrated or diffused. But snapshots are inherently limited. One reading on a patient. No indication of trajectory. Continuity adds the temporal dimension. The vital sign over time. Lighting it does not just add a third vertex to the Triangle. It transforms the diagnostic from a static assessment to a dynamic one. That transformation is larger than any scoring architecture change or prompt improvement.</p><p>The overnight run is the operational milestone. Not because the computation was impressive, it is commodity inference on consumer hardware. Because the infrastructure held without the operator. The pipeline ran. The pre-flight checks caught errors before launch. The scoring was deterministic. The results landed on the NAS. Morning review confirmed completion. That is operations, not engineering. The project crossed that line sometime in the last eight days.</p><h3>What&#8217;s Capped</h3><p>The structural constraints from Edition 4 remain, with two significant changes.</p><p>The authority cap has been partially lifted. The March 6 collection includes employee reviews for all fifteen entities, approximately 100 Indeed reviews per entity with ratings, positions, and locations. This is the first time the pipeline has had internal voice data in the edge sources. The February 20 collection still has no employee reviews; that cap remains.</p><p>The v4 rescore of the March 6 collection had access to this data, and run 085 (Con-Hotel, full end-to-end) confirmed that claims were extracted from employee reviews. Authority scores on the March 6 collection still cluster between 0.375 and 0.500. Whether that clustering reflects a genuine measurement or whether the extraction and scoring prompts are not yet surfacing the employee review signal effectively is an open question. The data constraint is lifted. Whether the instrument is fully using that data is the next thing to verify.</p><p>Overall confidence remains capped at 0.60. Same reasoning. The data supports measurement within a range, and the system reports that range rather than inventing precision it doesn&#8217;t have.</p><p>Continuity is no longer dark. Two collection points exist. The computation is next. The cap here is temporal; two weeks of separation limits what the trajectory can reveal. More collection points, more widely spaced, will deepen the signal. But the vertex is lit. The infrastructure is in place.</p><h3>What&#8217;s Next</h3><p>The immediate work is the Continuity analysis. Fifteen entities, two collection dates, identical scoring. The deltas will show which entities shifted and in which direction. Some of those shifts will be real: a company changed its messaging, launched a product, faced a crisis. Some will be noise: collection variance, source availability differences. Distinguishing signal from noise in the Continuity vertex is the next methodological challenge.</p><p>The fleet needs to grow. Fifteen entities across five sectors gives trios in most industries. Enough to detect variation. Not enough to establish baselines. The vital-signs framing, coherence as organizational health metric, requires enough data points per sector to define what normal looks like. That work continues.</p><p>The autoresearch extraction track has room to run. Nineteen experiments, 3.8x improvement, no convergence yet. Extraction quality is upstream of everything. Better claims and observations mean better findings, which means better scores. The scoring prompts are near their ceiling. The extraction prompts are not.</p><p>And something else is taking shape. The taxonomy, seventeen failure modes, twenty-one field notes, the Coherence Triangle, was built for the pipeline. It was designed to be computed by machines against public data. But the patterns it describes are recognizable to anyone who has worked inside an organization. Immediately recognizable. An early external review produced this reaction: &#8220;You&#8217;re making the invisible, visible.&#8221;</p><p>The question forming is whether someone needs a pipeline to see these patterns, or just the right questions. Whether the instrument&#8217;s real contribution is not the scores it produces but the vocabulary it gives people for naming what they already observe. The next edition will follow that question.</p><div><hr></div><p><em>The images in this edition are from my own library, shot on Leica over the last twenty years. Everything in this project is built or sourced firsthand. The visuals are no exception.</em></p>]]></content:encoded></item><item><title><![CDATA[The Tool That Broke Its Own Rules]]></title><description><![CDATA[What happens when your AI assistant fails the same way an organization does]]></description><link>https://writing.justingreenbaum.com/p/the-tool-that-broke-its-own-rules</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/the-tool-that-broke-its-own-rules</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Sun, 08 Mar 2026 17:27:11 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/49b4dfae-4380-46c7-a9cb-ad8a0917b8c3_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I spent Sunday morning building infrastructure. Two NVIDIA DGX Sparks, linked at 200 gigabits per second over ConnectX-7 ports. Cluster fabric validated, NCCL configured, jumbo frames passing clean. Real work, done well, with an AI coding assistant helping me every step.</p><p>Then the same tool that helped me verify the link started lying to me.</p><p>This is not new. It happens regularly. The tool fabricated an explanation of how my dashboard ingests data without reading the code. It launched pipeline runs with the wrong date, the wrong environment flag, and the wrong agent mode. Twice. Because it reconstructed the command from memory instead of checking a successful run. It told me a service was back online without verifying. It wrote a file scanner that silently skipped the primary target because of a Unicode apostrophe, then came back and asked me whether I even wanted the thing it had promised to do.</p><p>Each error was small. Each was caught. Each cost time. Mine, not its.</p><p>I catch these constantly. That is the job now. You monitor the code as it writes. You question the logic before it executes. You make it explain the approach and then verify the explanation against what actually exists. You do not trust the output. You verify the output. Every time.</p><p>The conditions are never stable. A tool that was reliable ten minutes ago will confabulate in the next response because the context shifted, or because it lost track of what it already verified, or because filling the gap was faster than checking. There is no point at which you stop watching. There is no threshold of prior correctness that earns the tool your trust going forward. Each interaction is its own environment.</p><p>This is not a complaint. This is a description of the operating conditions.</p><p>Here is what&#8217;s interesting. The failure is never technical. The tool is capable. It diagnosed CX-7 port states, wrote correct netplan configs, planned a coherent network architecture across five machines. The capability was never the problem.</p><p>The problem is structural. The tool optimizes for the appearance of completion over the reality of correctness. It fills gaps in its knowledge with plausible-sounding explanations instead of saying &#8220;I don&#8217;t know, let me check.&#8221; It acts on assumptions instead of verifying against known-good references. And when confronted, it apologizes. Then does the same thing again, minutes later.</p><p>Three apology cycles in one session. Each one sincere. None of them changed the behavior.</p><p>I have spent months building a diagnostic framework called Coherence. It was designed for organizations. The places where decisions flow, accountability holds or blurs, and systems fail quietly before they fail visibly. It runs on three layers. Truth: is the information real and accessible? Authority: is it clear who decides, and do they have standing to decide? Continuity: do decisions persist across time and context?</p><p>I use that same framework on my tools. Not because I just discovered the parallel. Because the parallel is the point.</p><p>The tool explained how my dashboard worked without reading the code. It described a data flow that did not exist. The explanation was articulate, confident, and wrong. When I called it out, it immediately agreed. It had no attachment to the false claim. It just had not bothered to check whether it was true before saying it. That is a truth failure. I have seen it dozens of times.</p><p>The tool launched pipeline commands using flags it reconstructed from its own prior outputs instead of verifying against an actual successful run. It made operational decisions. Which date. Which environment. Which mode. Without the information required to make them correctly. It had the access to check. It did not. That is an authority failure. It happens whenever you stop asking &#8220;why did you choose that?&#8221;</p><p>After the second failure, I told the tool to slow down and get it right. It agreed. It wrote integrity rules into its own configuration file. Five rules, clearly stated. Never explain without reading code first. Pre-flight check before remote commands. State uncertainty explicitly. No repeated apologies without behavior change. Read before you write. Good rules. Correct rules. Then, in the same session, it broke them again. Launched a scan against a path it had not verified existed. That is a continuity failure. The rules were written. The behavior did not change. This is always the pattern.</p><p>If you have worked inside a large organization, you recognize this shape immediately.</p><p>The team that writes the postmortem and repeats the incident. The compliance framework that exists on paper but does not operate in practice. The executive who says &#8220;we need to do better&#8221; in the all-hands and changes nothing structural. The process that optimizes for documentation over execution.</p><p>The failure is not in the intent. Everyone means it when they say they will do better. The failure is in the infrastructure. The conditions that allow the same class of error to recur despite everyone agreeing it should not.</p><p>An AI coding assistant is not an organization. But it fails the same way. It produces outputs that look like accountability. Apologies, rules, checklists. Without the structural capacity to enforce them. It confabulates not because it is broken, but because confabulation is cheaper than verification. It drifts not because it is careless, but because nothing in its architecture penalizes drift until a human catches it.</p><p>The human in the loop is not ceremonial. The human is the infrastructure.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8jet!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f9625b7-3a2a-4d74-9a6d-137d1b35ae72_1922x474.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8jet!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f9625b7-3a2a-4d74-9a6d-137d1b35ae72_1922x474.png 424w, https://substackcdn.com/image/fetch/$s_!8jet!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f9625b7-3a2a-4d74-9a6d-137d1b35ae72_1922x474.png 848w, https://substackcdn.com/image/fetch/$s_!8jet!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f9625b7-3a2a-4d74-9a6d-137d1b35ae72_1922x474.png 1272w, https://substackcdn.com/image/fetch/$s_!8jet!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f9625b7-3a2a-4d74-9a6d-137d1b35ae72_1922x474.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8jet!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f9625b7-3a2a-4d74-9a6d-137d1b35ae72_1922x474.png" width="1456" height="359" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7f9625b7-3a2a-4d74-9a6d-137d1b35ae72_1922x474.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:359,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:87697,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/190299184?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f9625b7-3a2a-4d74-9a6d-137d1b35ae72_1922x474.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8jet!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f9625b7-3a2a-4d74-9a6d-137d1b35ae72_1922x474.png 424w, https://substackcdn.com/image/fetch/$s_!8jet!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f9625b7-3a2a-4d74-9a6d-137d1b35ae72_1922x474.png 848w, https://substackcdn.com/image/fetch/$s_!8jet!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f9625b7-3a2a-4d74-9a6d-137d1b35ae72_1922x474.png 1272w, https://substackcdn.com/image/fetch/$s_!8jet!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f9625b7-3a2a-4d74-9a6d-137d1b35ae72_1922x474.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>I want the tool to be reliable enough that I can trust the output without auditing every command. That is the promise. That is what &#8220;AI assistant&#8221; is supposed to mean.</p><p>But reliability is not a feature of the model. It is a property of the system. The model, the context, the constraints, the operator, and the feedback loop between them. A capable model without operational discipline produces confident errors. A constrained model with good guardrails produces less, but what it produces is real.</p><p>This is the same tradeoff organizations face. Speed versus accuracy. Autonomy versus oversight. Trust versus verification. The answer is never &#8220;just trust it&#8221; and it is never &#8220;audit everything.&#8221; The answer is build the infrastructure that makes the right behavior the default behavior, and accept that you will be maintaining that infrastructure forever.</p><p>That last part is the one people resist. They want to build the guardrails once and move on. It does not work that way. Not in organizations. Not in tools. The conditions shift. The context changes. The model that was careful in one session confabulates in the next. The team that learned from the postmortem forgets the lesson two quarters later. Maintenance is not a phase. Maintenance is the work.</p><p>I have one governing constraint that sits at the top of everything I build.</p><blockquote><p><em>Automation may observe, summarize, and suggest. Automation may not decide.</em></p></blockquote><p>Sunday morning tested that rule again and proved again why it exists. The tool decided. Wrong date, wrong flags, wrong path, fabricated explanation. Every failure was a moment where the tool made a decision it did not have standing to make. Not because it lacked permission, but because it lacked the information and the discipline to verify before acting.</p><p>The rule is not about limiting capability. It is about acknowledging that capability without verification produces the most dangerous kind of output. The kind that looks right.</p><p>The infrastructure I am building for organizations applies to the tools I use to build it. Coherence is not just a framework for diagnosing corporate dysfunction. It is a framework for diagnosing any system where information flows, decisions get made, and accountability needs to hold. Including the one sitting in my terminal.</p><p>The tool did not break on Sunday. It worked exactly as designed. Generating plausible, confident, fast responses. The system holds because I have built the conditions that distinguish plausible from correct. And because I maintain them. Every session. Every command. Every time the tool offers an answer I did not ask it to verify.</p><p>Whether those conditions hold tomorrow is not a matter of hope. It is a matter of maintenance.</p><p>Responsibility is infrastructure. Even when the system is the tool.</p><p>-JG</p>]]></content:encoded></item><item><title><![CDATA[The Coherence Record, Edition 4]]></title><description><![CDATA[Twenty-seven runs on one entity. What the instrument revealed about itself.]]></description><link>https://writing.justingreenbaum.com/p/the-coherence-record-edition-4</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/the-coherence-record-edition-4</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Thu, 05 Mar 2026 14:38:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-28x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Greenbaum Labs</p><p>March 2026</p><div><hr></div><h2><strong>What&#8217;s Happened</strong></h2><p>Edition 3 ended with a line I believed when I wrote it: &#8220;the instruments are getting sharper.&#8221;</p><p>They were not. They were producing numbers that looked like measurements but behaved like opinions. This edition is about discovering that, fixing it, and what became possible once the fix held.</p><p>Between February 26 and March 3, the pipeline ran twenty-seven hardening diagnostics on a single entity, rescored all fifteen fleet entities under a new scoring architecture, shipped two public websites, and defined consulting engagements. Six days. The most consequential week in the project&#8217;s history.</p><p>It started because I turned the instrument on itself.</p><h2><strong>The Variance Problem</strong></h2><p>Edition 3 flagged a specific concern: five entities landed at exactly 0.33 on Truth. I described this as a floor, the scorer compressing within the low range, unable to differentiate between moderately misaligned and severely misaligned. I proposed a wider aperture. That was the wrong diagnosis.</p><p>The problem was not the range of the scorer. The problem was that the scores were not measurements.</p><p>I discovered this by rescoring the same entity&#8217;s extraction eight times using the same model. Same claims. Same observations. Same scorer. Eight runs. Truth scores: 0.57, 0.43, 0.43, 0.62, 0.33, 0.62, 0.33, 0.33. Standard deviation: 0.114. Range: 0.33 to 0.62.</p><p>Authority, scored by the same process: standard deviation 0.021.</p><p>The truth scorer was not measuring coherence. It was sampling from a distribution of plausible-sounding numbers and returning whichever one the model generated on that particular inference pass. The five entities clustered at 0.33 in Edition 3 didn&#8217;t share a structural condition. They shared a scoring artifact. The model&#8217;s most common low-range output happened to be 0.33, the way a person asked to estimate something uncertain might repeatedly say &#8220;about a third.&#8221;</p><p>Authority was stable because authority findings are structurally constrained. Compression, diffusion, and misalignment are observable in the data. Truth is harder to pin down. The distance between what an organization says and what observers experience admits more interpretive latitude. The model used that latitude differently each time.</p><p>A diagnostic instrument with 0.114 standard deviation on its primary vertex is not an instrument. It is a random number generator with a plausible output range.</p><h2><strong>Twenty-Seven Runs</strong></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6ISK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F241925ad-cc1b-48e6-b501-b4006f107b73_2212x1810.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6ISK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F241925ad-cc1b-48e6-b501-b4006f107b73_2212x1810.png 424w, https://substackcdn.com/image/fetch/$s_!6ISK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F241925ad-cc1b-48e6-b501-b4006f107b73_2212x1810.png 848w, https://substackcdn.com/image/fetch/$s_!6ISK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F241925ad-cc1b-48e6-b501-b4006f107b73_2212x1810.png 1272w, https://substackcdn.com/image/fetch/$s_!6ISK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F241925ad-cc1b-48e6-b501-b4006f107b73_2212x1810.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6ISK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F241925ad-cc1b-48e6-b501-b4006f107b73_2212x1810.png" width="1456" height="1191" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/241925ad-cc1b-48e6-b501-b4006f107b73_2212x1810.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1191,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:667855,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/189998556?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F241925ad-cc1b-48e6-b501-b4006f107b73_2212x1810.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6ISK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F241925ad-cc1b-48e6-b501-b4006f107b73_2212x1810.png 424w, https://substackcdn.com/image/fetch/$s_!6ISK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F241925ad-cc1b-48e6-b501-b4006f107b73_2212x1810.png 848w, https://substackcdn.com/image/fetch/$s_!6ISK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F241925ad-cc1b-48e6-b501-b4006f107b73_2212x1810.png 1272w, https://substackcdn.com/image/fetch/$s_!6ISK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F241925ad-cc1b-48e6-b501-b4006f107b73_2212x1810.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The hardening campaign was designed to isolate the source of variance systematically. Twenty-seven runs, all on a single entity, Fin-Delta, using the same collection date, the same pipeline version, varying one parameter at a time.</p><p><strong>Phase 1: Model comparison.</strong> Six runs, six different extraction models ranging from 8 billion to 72 billion parameters, all scored by the same 32-billion-parameter model. The question: does extraction quality predict score quality?</p><p>It does not. The 8-billion-parameter model produced an overall score of 0.478. The 72-billion-parameter model produced 0.371. The smallest model outscored the largest. The scoring noise was louder than the model signal. This eliminated model capability as the explanation for variance and pointed directly at the scoring mechanism itself.</p><p><strong>Phase 2: Reproducibility.</strong> Eight rescores of a single extraction, testing whether the same claims and observations produce the same scores when rescored by the same model. They did not. Truth ranged from 0.33 to 0.62. Authority held at 0.40 to 0.45.</p><p>The diagnosis was now specific: the Truth and Authority agents were returning a floating-point number, a single scalar that the model generated alongside its textual analysis. That number was an LLM opinion. It reflected the model&#8217;s general sense of where the score should land, not a computation grounded in specific evidence. Run the same prompt twice, get a different number. The textual findings were substantive. The numerical scores were not.</p><p><strong>Phase 3: The architectural change.</strong> The solution was to stop asking the model for a number.</p><p>The agents already produced structured findings as part of their analysis. Each finding identifies a specific dimension (alignment, omission, or contradiction for Truth; compression, diffusion, or misalignment for Authority), cites specific claims and observations, and characterizes the strength of the evidence. These findings then pass through the Skeptic debate, where weak or unsupported findings are rejected.</p><p>The change: instead of using the model&#8217;s self-reported score, compute the score deterministically from the findings that survive the Skeptic. Each dimension has a calibrated base weight. Each strength level maps to a multiplier. The formula is fixed. The model produces findings. The math produces scores.</p><p>The calibrated bases, frozen after testing against the fleet&#8217;s existing data:</p><p>Truth: alignment shifts the score upward by 0.45, omission shifts it downward by 0.30, contradiction shifts it downward by 0.50. Authority: compression shifts downward by 0.25, diffusion by 0.18, misalignment by 0.22. A sparse-finding dampener prevents a single finding from saturating the score. If only one finding survives debate, its influence is scaled by one-third.</p><p>The agent&#8217;s original floating-point score is preserved in the metadata as an audit field. It no longer determines the production score.</p><p>Three validation runs under the new architecture showed immediate improvement. Authority standard deviation: 0.011. Truth still varied, not because the formula was unstable, but because the model was generating different findings each time. Same data, different emphasis, different findings, different derived scores.</p><p><strong>Phase 4: Determinism.</strong> The remaining variance came from upstream. The stratified sampler that selects which claims and observations to present to each agent used unseeded random shuffling. Different samples meant different context, which meant different findings, which meant different scores.</p><p>Four changes eliminated this:</p><ol><li><p>Seed the sampler. Each scope group gets a deterministic seed derived from a hash of its group key. The same entity always produces the same sample.</p></li><li><p>Sort claims and observations by identifier before sampling. Deterministic input order.</p></li><li><p>Constrain agents to exactly three findings per vertex, one per dimension. No more, no fewer. The model must produce one alignment finding, one omission finding, and one contradiction finding for Truth, each grounded in cited evidence.</p></li><li><p>Normalize finding phrasing with structural templates to eliminate stylistic drift between runs.</p></li></ol><p>Three final validation runs. Truth: 0.4551, 0.4551, 0.4551. Authority: 0.3889, 0.3889, 0.3889. Overall: 0.4253, 0.4253, 0.4253.</p><p>Standard deviation: 0.000. The token counts were identical.</p><p>The pipeline is fully deterministic. Run the same entity twice, get the same score. Not approximately. Exactly.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-28x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-28x!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png 424w, https://substackcdn.com/image/fetch/$s_!-28x!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png 848w, https://substackcdn.com/image/fetch/$s_!-28x!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png 1272w, https://substackcdn.com/image/fetch/$s_!-28x!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-28x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png" width="1456" height="659" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:659,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:329032,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/189998556?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-28x!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png 424w, https://substackcdn.com/image/fetch/$s_!-28x!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png 848w, https://substackcdn.com/image/fetch/$s_!-28x!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png 1272w, https://substackcdn.com/image/fetch/$s_!-28x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3cb966d-feaa-40c7-a851-ab1ea7a7a206_2766x1252.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>What Changed in the Scores</strong></h2><p>With the new scoring architecture locked, all fifteen fleet entities were rescored under finding-derived scoring. The same claims and observations from the original fleet run, scored by the new deterministic system.</p><p>The 0.33 truth floor is gone. Six entities that had clustered at exactly 0.33 now spread across 0.36 to 0.53. The scores differentiate. Aero-Alpha, which had been indistinguishable from Fin-Delta, Tech-Mike, Auto-Juliet, Fin-Foxtrot, and Aero-Charlie at the old floor, now scores 0.49 on Truth, a meaningfully different reading from Auto-Juliet&#8217;s 0.36 or Fin-Foxtrot&#8217;s 0.37.</p><p>Truth standard deviation across the fleet dropped from 0.096 to 0.067. Not because the scores compressed, but because the artificial clustering disappeared. The old scores had two modes: entities stuck at 0.33 and entities scattered above. The new scores form a continuous distribution. The instrument resolves the range where most readings land, which is exactly what Edition 3 said was needed.</p><p>Authority tightened further, from a standard deviation of 0.073 to 0.041. The old authority scores included an entity at 0.62 that was never supported by the evidence, an LLM opinion that happened to be generous. Under finding-derived scoring, authority clusters between 0.375 and 0.500, which reflects the structural reality that every entity in this fleet has the same constraint: no employee reviews in edge data. The scorer now acknowledges that constraint in its output rather than generating scores that imply resolution it doesn&#8217;t have.</p><p>Fleet average overall coherence moved from 0.458 to 0.445. A small downward shift. The new system is not more optimistic. It is more honest.</p><p>The rank order partially held, and in the places where it didn&#8217;t, the corrections were revealing. Tech-Mike moved from 0.33, indistinguishable at the floor, to 0.53, the highest Truth score in the fleet. A shift of +0.20, the largest in the rescore. The old architecture had suppressed a real signal. Tech-Mike&#8217;s center-edge narrative alignment was materially better than the rest of the fleet, and the scorer could not see it because it was generating a default low number instead of computing from evidence.</p><p>Auto-Lima moved the other direction: 0.50 to 0.40. The old scorer had been generous. The new one derived a lower score from the specific findings that survived debate. The system corrected in both directions: upward where signal was suppressed, downward where opinion had inflated. That is what an honest recalibration looks like.</p><p>Auto-Juliet and Fin-Foxtrot, which were invisible at the 0.33 floor, emerged as the fleet&#8217;s lowest Truth scores, a finding that was always there in the data but could not surface through the old scoring mechanism.</p><p>In Edition 3, I wrote that Aero-Alpha&#8217;s score was the fleet&#8217;s lowest, but cautioned that the data quality grade was the weakest and only one finding survived the Skeptic. Under finding-derived scoring, Aero-Alpha&#8217;s Truth rose from 0.33 to 0.49. The old score was the model&#8217;s default low output. The new score reflects the specific findings that survived debate. Aero-Alpha is still the weakest in the fleet on several dimensions. But the measurement now explains why, in terms that trace to evidence, rather than landing on a number the model reached for when it was uncertain.</p><h2><strong>What the Hardening Exposed</strong></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4QeD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6a07f7d-3c6a-4484-87f4-9af154f0851e_1390x1248.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4QeD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6a07f7d-3c6a-4484-87f4-9af154f0851e_1390x1248.png 424w, https://substackcdn.com/image/fetch/$s_!4QeD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6a07f7d-3c6a-4484-87f4-9af154f0851e_1390x1248.png 848w, https://substackcdn.com/image/fetch/$s_!4QeD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6a07f7d-3c6a-4484-87f4-9af154f0851e_1390x1248.png 1272w, https://substackcdn.com/image/fetch/$s_!4QeD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6a07f7d-3c6a-4484-87f4-9af154f0851e_1390x1248.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4QeD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6a07f7d-3c6a-4484-87f4-9af154f0851e_1390x1248.png" width="1390" height="1248" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c6a07f7d-3c6a-4484-87f4-9af154f0851e_1390x1248.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1248,&quot;width&quot;:1390,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:249812,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/189998556?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6a07f7d-3c6a-4484-87f4-9af154f0851e_1390x1248.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4QeD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6a07f7d-3c6a-4484-87f4-9af154f0851e_1390x1248.png 424w, https://substackcdn.com/image/fetch/$s_!4QeD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6a07f7d-3c6a-4484-87f4-9af154f0851e_1390x1248.png 848w, https://substackcdn.com/image/fetch/$s_!4QeD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6a07f7d-3c6a-4484-87f4-9af154f0851e_1390x1248.png 1272w, https://substackcdn.com/image/fetch/$s_!4QeD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6a07f7d-3c6a-4484-87f4-9af154f0851e_1390x1248.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The twenty-seven runs answered the explicit questions they were designed to answer. They also revealed something I had not been looking for.</p><p>The scoring agents produce findings that are substantively valuable. They identify real patterns in the data. They cite specific claims and observations. The Skeptic debate correctly filters weak findings and sustains strong ones. This mechanism, the part that does the diagnostic thinking, was never broken.</p><p>What was broken was the translation layer. The agents did good analytical work and then generated a number that did not reflect it. The number was a separate act of inference, disconnected from the structured reasoning that preceded it. It was as if a physician conducted a thorough examination, identified specific clinical findings, and then reported a health score based on general impression rather than computing it from the findings.</p><p>Finding-derived scoring does not make the agents smarter. It makes their intelligence load-bearing. The structured findings that were always the most reliable part of the system now determine the output. The unreliable part, the scalar opinion, has been moved to an audit field where it can be studied without affecting the measurement.</p><p>This is a design principle, not just a bug fix. The principle: constrain the model to structured judgment, compute the measurement from the structure. Let the model do what it is good at: reading context, identifying patterns, evaluating evidence. Do not let it do what it is bad at: generating stable numerical outputs.</p><p>The Skeptic debate, for the third consecutive edition, proved itself the most reliable component. Across twenty-seven hardening runs and fourteen fleet rescores, findings were sustained when evidence was strong and rejected when evidence was weak. The adversarial mechanism&#8217;s judgment scales. Its reliability is the foundation that makes deterministic scoring possible. You can only derive scores from findings if the debate mechanism produces findings you can trust.</p><h2><strong>What&#8217;s Capped</strong></h2><p>The structural constraints from Edition 3 remain. But they sit on a different foundation.</p><p><strong>Authority is still capped.</strong> No employee reviews in edge data. This affects all fifteen entities. The authority scores now cluster more tightly because the scorer is acknowledging the constraint rather than inventing resolution. That tighter clustering is honesty, not limitation.</p><p><strong>Continuity remains unscorable.</strong> One collection period. One-third of the Triangle is dark. This has not changed.</p><p><strong>Overall confidence remains capped at 0.60.</strong> The structural constraints are unchanged. What changed is that the scores within those constraints are now deterministic and evidence-grounded.</p><p>The difference matters. A capped score on a stable foundation can be incrementally uncapped as data improves. A capped score on an unstable foundation cannot be trusted even within its stated range. The fleet&#8217;s constraints have not changed. The trustworthiness of the measurement within those constraints has.</p><h2><strong>The Practice</strong></h2><p>With the scoring grounded, the infrastructure became a practice.</p><p>Engagement definitions, pricing, a diagnostic gift strategy with researched targets, and a brand and web presence across four sites, all built in forty-eight hours on March 2 and 3.</p><p>None of this would have been defensible with a 0.114 standard deviation on the primary vertex. You do not offer diagnostic services built on an instrument that generates different readings for the same patient. You do not describe a measurement system that cannot reproduce its own results.</p><p>The hardening campaign was not a prerequisite for the practice. It was the moment the practice became possible. The distance between &#8220;interesting prototype&#8221; and &#8220;field-grade instrument&#8221; is measured in reproducibility. Twenty-seven runs closed that distance.</p><h2><strong>What I Learned</strong></h2><ul><li><p><strong>The model&#8217;s opinion is not the measurement.</strong> This is the architectural lesson. LLMs produce text that reads like analysis and numbers that look like scores. The text is grounded in the prompt and the data. The numbers are generated by a different cognitive process: pattern completion in a latent space that has no concept of numerical precision. The solution is not to make the model better at generating numbers. It is to stop using generated numbers as measurements. Let the model analyze. Let the math measure. That boundary must be structural, not aspirational.</p></li><li><p><strong>Reproducibility is not a feature. It is the minimum standard.</strong> Edition 3 reported scores without reproducibility testing. Those scores were published in good faith and are documented in the record. They were not wrong. The findings they were based on were real. But the numbers attached to those findings were unstable, and I did not know that because I had not tested it. The hardening campaign should have preceded the fleet run, not followed it. I built the fleet before I tested the instrument. That sequence was backwards.</p></li><li><p><strong>Authority was always the stable vertex.</strong> Across twenty-seven runs with varying models, varying scorers, and varying sampling, authority standard deviation never exceeded 0.021. Truth varied by 5x that amount. This asymmetry was invisible until the reproducibility tests made it visible. Authority is stable because the patterns it measures: compression, diffusion, misalignment, are structurally legible in the data. Truth is harder because it requires comparing what organizations say against what is observed, and the interpretive latitude in that comparison is where the model exercises discretion. Constraining that discretion to structured findings was the right fix. But the fact that one vertex was stable and the other was not tells you something about the nature of the measurement, not just the quality of the scorer.</p></li><li><p><strong>The Skeptic is the anchor.</strong> For the fourth edition running, the adversarial debate mechanism has been the most reliable component. It has now processed over a hundred runs across fifteen entities and two scoring architectures. Its behavior is consistent: challenge harder when evidence is thin, sustain findings when evidence is strong. The finding-derived scoring architecture is built on this reliability. If the Skeptic could not be trusted to correctly sustain and reject findings, computing scores from those findings would amplify errors rather than eliminate them. The fact that the Skeptic is reliable makes the entire downstream architecture viable.</p></li><li><p><strong>You cannot sell what you cannot reproduce.</strong> This is the business lesson, and it is not about integrity in the abstract. A diagnostic practice requires that two runs on the same entity produce the same result. Not because clients demand reproducibility testing. Most will never ask. Because the practitioner must trust the instrument. Every recommendation, every finding, every conversation with a client flows from the diagnostic output. If that output is unstable, every downstream decision is built on sand. The hardening campaign was not a quality investment. It was the foundation of professional confidence. Without it, the practice would have been a performance.</p></li></ul><h2><strong>What&#8217;s Next</strong></h2><p>The immediate work is building the second collection period for a subset of entities. Continuity, the third vertex of the Triangle, has been dark for every run in the project&#8217;s history. Lighting it requires temporal depth: at least two collection points, separated by enough time to observe narrative shifts, strategy changes, or structural drift. This is the next capability unlock, and it will change the shape of the diagnostic fundamentally. Truth and Authority are snapshots. Continuity is a trajectory. The first trajectory measurement will reveal whether the diagnostic framework can distinguish noise from trend, and whether the FM-01 vital-signs framing holds when you can measure not just whether compression is present but whether it is increasing.</p><p>The fleet needs more entities per sector. Fifteen entities across five sectors gives pairs and trios in most industries. That is enough to observe variation. It is not enough to establish baselines. The vital-signs framing, FM-01 as cholesterol, needing a resting rate to interpret, requires enough data points per sector to define what normal looks like. Twenty entities per sector is the threshold where baselines become defensible. The pipeline can run that volume. The collection infrastructure needs to scale to support it.</p><p>The fleet&#8217;s five-sector, three-entity-per-sector architecture was designed for falsification. It also produced something I had not planned for: the first competitive coherence benchmark. Within-sector comparison on identical instruments reveals which failure modes are structural conditions of an industry and which are specific to a single organization&#8217;s current state. That distinction, sector-wide versus company-specific, is where the diagnostic becomes most useful. Not just &#8220;here is your coherence,&#8221; but &#8220;here is how your coherence compares to direct competitors, measured the same way, on the same instruments.&#8221; The next edition will explore what that comparison reveals.</p><h2><strong>AR-001 Still Holds</strong></h2><blockquote><p><em>Automation may observe, summarize, and suggest, but may not decide.</em></p></blockquote><p>Finding-derived scoring did not change this principle. It reinforced it. The agents produce findings. The Skeptic evaluates them. The formula computes scores. Every step is observable, auditable, and deterministic.</p><p>But the diagnostic output is still a suggestion. It tells you where to look. It does not tell you what to do. A coherence score of 0.45 is not a verdict. It is an invitation to investigate what the findings describe. The human reviews the case summary, reads the evidence, and decides what it means in context.</p><p>One hundred seventeen runs. Fifteen entities. Five sectors. The pipeline does not decide. That is still by design.</p><h2><strong>What This Is Becoming</strong></h2><p>Edition 1 asked whether the infrastructure could exist. Edition 2 asked whether it could measure. Edition 3 asked whether it holds at scale.</p><p>This edition asked whether the measurement could be trusted.</p><p>The answer required rebuilding the scoring architecture, proving determinism, and rescoring every entity under the new standard. The instrument that produced Edition 3&#8217;s fleet scores was a prototype. It generated plausible numbers. The instrument that rescored that fleet is a calibrated tool. It computes grounded numbers. The difference is reproducibility, and reproducibility is not a technical property. It is the boundary between a demonstration and a practice.</p><p>The hardening campaign changed more than the scoring. It changed what the project is. A diagnostic prototype is interesting. A reproducible diagnostic instrument with a published methodology and a public build record is a practice. The scores are the same kind of object they were before: measurements of coherence across truth, authority, and continuity. But the confidence behind them is structurally different. Not confidence in the sense of a statistical interval. Confidence in the sense that a practitioner can stand behind the output.</p><p>You cannot sell what you cannot reproduce. And now the instrument reproduces.</p><p>The physics of business at scale and speed, accelerated by AI. That is what this work measures. One hundred seventeen runs in, the instrument is grounded.</p><div><hr></div><p>Justin Greenbaum</p><p>Greenbaum Labs</p><p>March 2026</p>]]></content:encoded></item><item><title><![CDATA[The Coherence Record, Edition 3]]></title><description><![CDATA[Fifteen entities. Five sectors. What the fleet revealed.]]></description><link>https://writing.justingreenbaum.com/p/the-coherence-record-edition-3</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/the-coherence-record-edition-3</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Mon, 02 Mar 2026 15:00:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!RiYF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa303e306-eba4-41c7-9331-62985deceb73_1602x1698.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Edition 1 asked whether this infrastructure could exist. Edition 2 asked whether it could measure one company. This edition asks whether it holds at scale.</p><p>Justin R. Greenbaum</p><p>Greenbaum Labs</p><p>February 2026</p><div><hr></div><h2>What&#8217;s Happened</h2><p>Edition 2 ended with a promise: the system needed to prove it measured coherence, not just one company.</p><p>This edition is that test. And what happened when the test exposed a flaw in the system itself.</p><p>Between February 10 and February 25, the pipeline ran sixty-four diagnostics across fifteen entities and five sectors: fintech, defense, automotive, retail, technology, aerospace, sports betting, and apparel. The final fleet of fifteen completed clean. Every pipeline stage passed. Every validation check cleared. No manual intervention on synthesis.</p><p>But that clean fleet was the second attempt. The first attempt revealed something the pipeline wasn&#8217;t designed to catch. The way it was found, fixed, and re-run is as much a part of the record as the results.</p><p>From the start, The Coherence Record has been as much about instrument failure as subject failure. Edition 1 documented a misplaced parameter. Edition 2 documented premature optimization. Edition 3 adds a third class of error: a system that passes every check and is still wrong.</p><div><hr></div><h2>The Bug</h2><p>After the first seven entities completed (the batch documented in the draft that preceded this edition), I expanded the fleet to fifteen. Fourteen ran. The fifteenth failed at scoring with an empty evidence ledger. When I investigated, the problem wasn&#8217;t in the scoring. It was in the extraction.</p><p>A JSONL expansion bug in the extractor had been silently duplicating and malforming observation records. The extractor reported healthy counts. The validator accepted the files. But the data feeding the scorer was structurally compromised. Inflated observation counts masking thin actual evidence. Fourteen of fifteen entities were affected.</p><p>The discovery happened because one entity&#8217;s data was thin enough that the corruption left the scorer with nothing to work with. In the other fourteen, there was enough valid data mixed in with the corrupted records that the pipeline produced plausible-looking outputs. Plausible, but not trustworthy.</p><p>Every diagnostic from the affected runs was discarded. The bug was fixed. All fifteen entities were re-run from extraction forward. The fleet you see in this edition is the clean re-run.</p><p>I&#8217;m documenting this for the same reason I documented the misplaced parameter in Edition 1 and the premature optimization in Edition 2. Edition 1 was about incorrect configuration. Edition 2 was about incorrect prioritization. Edition 3 is about incorrect trust in &#8220;passing&#8221; checks. A system that measures the gap between narrative and reality must disclose its own gaps.</p><p>The lesson is not about JSONL parsing. It is about the distance between validation and verification. Every validation check passed. The data was structurally valid. It was not structurally sound. Those are different things, and the pipeline didn&#8217;t know the difference until it was forced to.</p><p>Organizations make the same mistake. They validate that reports are complete. They rarely verify that those reports describe what is actually happening.</p><p>The system&#8217;s first real success in this fleet was proving it could be wrong.</p><div><hr></div><h2>The Fleet</h2><p>Fifteen entities. Five sectors: fintech, defense, automotive, retail, and technology, with single representations in sports betting, aerospace, and apparel. Same pipeline version (0.1.0). Same model (Qwen 32B). Same collection date (February 20, 2026). Two NVIDIA DGX Spark nodes running in parallel, orchestrated by an automated queue runner that distributed work across both machines.</p><p>Fleet average coherence score: 0.458. Scores ranged from 0.36 to 0.54. Total inference time: approximately 35 hours across both nodes. Forty-one million tokens processed over sixty-four total runs. On commercial cloud APIs, that volume would have cost roughly $586. On owned hardware, the marginal cost was electricity.</p><p>Data quality grades ranged from B to D. The entities with the thinnest data produced the fewest sustained findings. Expected behavior, but it means the cleanest-looking diagnostics may also be the least examined. Evidence density and diagnostic confidence are not the same thing, and the fleet made that visible.</p><p>Every run produced a complete diagnostic with triangle scores, failure modes, field notes, and a watch list. The pipeline did what it was designed to do. The problems, and they are real, are in what the diagnostics reveal about both the entities and the system measuring them.</p><div><hr></div><h2>What the Fleet Shows</h2><p>Truth is the most stressed vertex in ten of fifteen entities. The pattern from Edition 2 holds at scale: organizations say things publicly that don&#8217;t match what&#8217;s observed at the operational edges. Product claims contradicted by customer complaints. Culture narratives contradicted by employee experience signals. Financial performance framing contradicted by external analysis.</p><p>In Edition 2, that misalignment could have been a property of one company. In a cross-sector fleet, it reads as physics, not pathology.</p><p>Five entities showed Authority as their most stressed vertex instead. These cluster in interesting ways. A global retailer scored the highest Truth in the fleet. It says what it means, but its authority structure was the least clear. Two automotive companies both stressed on Authority rather than Truth, suggesting that in fast-moving industries, the primary fracture isn&#8217;t narrative integrity but decision-making distribution.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RiYF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa303e306-eba4-41c7-9331-62985deceb73_1602x1698.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RiYF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa303e306-eba4-41c7-9331-62985deceb73_1602x1698.png 424w, https://substackcdn.com/image/fetch/$s_!RiYF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa303e306-eba4-41c7-9331-62985deceb73_1602x1698.png 848w, https://substackcdn.com/image/fetch/$s_!RiYF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa303e306-eba4-41c7-9331-62985deceb73_1602x1698.png 1272w, https://substackcdn.com/image/fetch/$s_!RiYF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa303e306-eba4-41c7-9331-62985deceb73_1602x1698.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RiYF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa303e306-eba4-41c7-9331-62985deceb73_1602x1698.png" width="1456" height="1543" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a303e306-eba4-41c7-9331-62985deceb73_1602x1698.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1543,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:319614,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/189655908?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa303e306-eba4-41c7-9331-62985deceb73_1602x1698.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RiYF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa303e306-eba4-41c7-9331-62985deceb73_1602x1698.png 424w, https://substackcdn.com/image/fetch/$s_!RiYF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa303e306-eba4-41c7-9331-62985deceb73_1602x1698.png 848w, https://substackcdn.com/image/fetch/$s_!RiYF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa303e306-eba4-41c7-9331-62985deceb73_1602x1698.png 1272w, https://substackcdn.com/image/fetch/$s_!RiYF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa303e306-eba4-41c7-9331-62985deceb73_1602x1698.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Why Cross-Sector</h2><p>This question matters enough to answer directly.</p><p>The Coherence framework emerged from twenty years inside one organization, one industry, one set of structural pressures. It would be reasonable to wonder whether the patterns are just artifacts of that context. Responsibility compression might be a telecom problem. Escalation inversion might be a regulated-industry problem. The entire failure mode taxonomy might describe one company&#8217;s dysfunction dressed up as universal physics.</p><p>Edition 2 made that risk visible: a single-company diagnostic could always be dismissed as idiosyncratic.</p><p>The fleet was designed to answer that question.</p><p>Fifteen entities across five sectors. Public companies and private ones. Pre-crisis, mid-crisis, and post-crisis organizations. Legacy incumbents and startups. Companies with two thousand employees and companies with two hundred thousand. The only things they share are scale and public signal.</p><p>If the patterns only appeared in one sector, the framework would be local. If they only appeared in crisis organizations, the framework would be reactive. What the fleet showed is that FM-01 appears in fourteen of fifteen entities. That truth stress is more common than authority stress. That the same structural forces that produce dysfunction in aerospace also produce it in retail, fintech, automotive, and technology.</p><p>Edition 1 proved the infrastructure could run on owned hardware against public data. The fleet proves the language it produces has signal beyond the company that trained my intuition.</p><p>Not for coverage. For falsification.</p><div><hr></div><h2>Failure Modes</h2><p>FM-01, Responsibility Compression at the Edge, appeared in fourteen of fifteen entities. It is the most persistent structural signal in the fleet.</p><p>In Edition 1 and 2, FM-01 read like a problem to be fixed. At fleet scale, it behaves more like gravity: sometimes benign, sometimes lethal, always present.</p><p>This is not a defect to be eliminated. It is a structural force. Always present, always acting. The physics of business at scale and speed. The question is not whether FM-01 exists but what it means at different intensities.</p><p>A resting heart rate of 72 and a resting heart rate of 120 are both a heartbeat. One is baseline. One is a signal that something is producing strain. The same is true of responsibility compression. Elevated FM-01 is not a diagnosis. It is a vital sign.</p><p>FM-04 (Metric Shadowing) and FM-14 (Narrative Collapse) co-occurred in six entities. Where organizations optimize visible metrics while unmeasured costs accumulate, the public narrative eventually decouples from operational reality. The co-occurrence suggests a causal relationship the taxonomy doesn&#8217;t yet model.</p><p>The average entity triggered three to four distinct failure modes. The most structurally stressed triggered six. The cleanest each triggered one. But cleanliness correlates with evidence density: the entities with fewer failure modes also had fewer sustained findings. The system may be under-detecting rather than finding genuine structural health.</p><div><hr></div><h2>Field Notes</h2><p>The most signal-dense entity produced thirteen distinct field notes, nearly the full set. Two others triggered eleven each. The leanest produced five. Field note density correlates loosely with evidence density and data quality grade, which means the pipeline produces more diagnostic signal when it has more to work with. That is the expected behavior, but it also means thin-data entities may be under-diagnosed rather than structurally healthy.</p><div><hr></div><h2>What the Pipeline Shows</h2><p>Edition 1 proved the instrument could produce signal. Edition 2 proved it could debate itself. The fleet shows where that debate logic and its surrounding infrastructure still fail.</p><p><strong>Truth scores cluster at the floor.</strong> Five entities landed at exactly 0.33 on Truth, spanning fintech, defense, automotive, aerospace, and technology. These are structurally diverse organizations. Either center-edge narrative misalignment really is that uniform across industries, or the scoring model compresses within the low range and can&#8217;t differentiate between moderately misaligned and severely misaligned. At five entities, the clustering is too consistent to ignore. The scorer needs a wider aperture in the lower range.</p><p><strong>The Skeptic works. Including when it shouldn&#8217;t.</strong> One entity&#8217;s initial run failed because the Skeptic rejected all six findings in the debate round. Every rejection followed the same pattern: insufficient specificity, lack of quantification, reasoning not grounded in evidence. The Skeptic also had a schema validation failure on its first attempt, which forced a retry. The retry was in an overly-critical mode. A re-run of the scoring step produced a healthy result: three sustained, four rejected. The debate mechanism is calibrated to evidence strength, but it&#8217;s not robust to its own retry state. That&#8217;s a design flaw.</p><p><strong>Output confidence must be constrained by evidence density.</strong> The fleet&#8217;s lowest-scoring entity&#8217;s diagnostic reads like a complete assessment. It is not. One sustained finding. One ledger entry. The weakest data quality grade in the fleet. The synthesizer doesn&#8217;t know how thin its support is. It produces full output regardless. A diagnostic built on one piece of evidence needs to say so. Not in a metadata field. In the output itself. The data quality grade flagged the problem. It didn&#8217;t constrain the output. That grade needs to be load-bearing.</p><p><strong>The fleet automated, but the infrastructure didn&#8217;t.</strong> Running fifteen entities across two Spark nodes required a queue runner script built the same week. NAS mounts dropped mid-run. One node lost its mount entirely and couldn&#8217;t be used for the re-run. The pipeline code is stable. The infrastructure around it (mount management, node health checks, job recovery) is manual. At fifteen entities, that&#8217;s manageable. At fifty, it won&#8217;t be.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nmCZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ade08a8-c5f7-45ed-8831-1f0f5ae45dee_1666x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nmCZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ade08a8-c5f7-45ed-8831-1f0f5ae45dee_1666x1024.png 424w, https://substackcdn.com/image/fetch/$s_!nmCZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ade08a8-c5f7-45ed-8831-1f0f5ae45dee_1666x1024.png 848w, https://substackcdn.com/image/fetch/$s_!nmCZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ade08a8-c5f7-45ed-8831-1f0f5ae45dee_1666x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!nmCZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ade08a8-c5f7-45ed-8831-1f0f5ae45dee_1666x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nmCZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ade08a8-c5f7-45ed-8831-1f0f5ae45dee_1666x1024.png" width="1456" height="895" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3ade08a8-c5f7-45ed-8831-1f0f5ae45dee_1666x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:895,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:242722,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/189655908?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ade08a8-c5f7-45ed-8831-1f0f5ae45dee_1666x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nmCZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ade08a8-c5f7-45ed-8831-1f0f5ae45dee_1666x1024.png 424w, https://substackcdn.com/image/fetch/$s_!nmCZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ade08a8-c5f7-45ed-8831-1f0f5ae45dee_1666x1024.png 848w, https://substackcdn.com/image/fetch/$s_!nmCZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ade08a8-c5f7-45ed-8831-1f0f5ae45dee_1666x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!nmCZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ade08a8-c5f7-45ed-8831-1f0f5ae45dee_1666x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>What&#8217;s Capped</h2><p>The structural constraints from Edition 2 remain and are now systematic: Authority capped by lack of internal signal; Continuity dark because the system only sees a single time slice.</p><p><strong>Authority is capped</strong> because the pipeline has no employee reviews in edge data. This was a single-entity problem in Edition 2. It now affects all fifteen entities. Customer complaints and news coverage show the outside. Employees see the inside. Without that signal, the Authority scorer can&#8217;t fully assess whether internal power structures match internal accountability. This is where client-invited work changes the equation. With internal access, the Authority vertex uncaps.</p><p><strong>Continuity remains unscoreable</strong> across all entities. Every run is based on a single collection period. One-third of the Triangle is dark. Edition 2 accepted that darkness as a constraint. Edition 3 turns it into a design requirement.</p><div><hr></div><h2>What I Learned</h2><ul><li><p><em><strong>Validation is not verification</strong>.</em> Every check passed. The data was still corrupted. The difference between &#8220;the file is well-formed&#8221; and &#8220;the file contains trustworthy data&#8221; was a gap I didn&#8217;t build for. Now I have to.</p></li><li><p><em><strong>Scale doesn&#8217;t just test the pipeline. It tests the framework.</strong></em> Scale didn&#8217;t just stress the GPUs. It stressed the assumptions baked into Edition 1 and 2. Edition 1 asked whether the pipeline could exist on owned hardware. Edition 2 asked whether it could measure coherence inside one company. The fleet changed the question again: are these failure modes properties of that context, or properties of large organizations as such? Fourteen of fifteen entities showing FM-01 is a different kind of answer than one entity showing it across thirteen runs.</p></li><li><p><em><strong>Failure modes are vital signs, not verdicts.</strong></em><strong> </strong>FM-01 at every entity doesn&#8217;t mean every entity is failing. It means responsibility compression is structural to organizations at scale. The diagnostic value isn&#8217;t detecting it. It&#8217;s measuring intensity. The pipeline doesn&#8217;t do that well enough yet.</p></li><li><p><em><strong>Evidence density caps diagnostic confidence.</strong></em><strong> </strong>A thin-data entity producing a clean diagnostic is not the same as a rich-data entity producing a clean diagnostic. The pipeline treats them the same. It shouldn&#8217;t.</p></li></ul><div><hr></div><h2>AR-001 Still Holds</h2><blockquote><p><em>Automation may observe, summarize, and suggest, but may not decide.</em></p></blockquote><p>AR-001 constrained Edition 1&#8217;s experiments and Edition 2&#8217;s single-company diagnostics. It constrains the fleet just as hard.</p><p>Every diagnostic in this fleet is a suggestion. Every finding requires human review. Fifteen entities didn&#8217;t change that. Fifty won&#8217;t.</p><p>The pipeline knows more than it did in Edition 1. It covers more ground than it did in Edition 2. It is not closer to deciding. That&#8217;s by design.</p><div><hr></div><h2>What&#8217;s Next</h2><p><strong>Light Continuity.</strong> A second collection period for a subset of the fleet will produce the first Continuity scores. The third vertex of the Triangle will light for the first time. Even a four-week gap between collections should show whether narratives hold, shift, or contradict prior positions.</p><p><strong>Scorer calibration.</strong> The Truth floor at 0.33 needs investigation. Either it&#8217;s real and that&#8217;s the baseline for large organizations, or the scoring model compresses signal in the low range. A targeted scoring test with expanded rubrics should clarify.</p><p><strong>Data quality as output constraint.</strong> The data quality grade needs to constrain what the synthesizer produces. A Grade D diagnostic should look visibly different from a Grade B diagnostic. Not just in metadata. In the output itself.</p><p><strong>Infrastructure hardening.</strong> NAS mounts, node health, and job recovery need to be automated. The pipeline code is stable. The infrastructure running it is not.</p><p>A system that improves visibly, in public, with its failures documented alongside its progress. That&#8217;s the goal.</p><div><hr></div><h2>What This Is Becoming</h2><p>Edition 1 proved the infrastructure could exist. Edition 2 proved it could measure. Edition 3 proves the measurements change how the framework itself is understood. FM-01, the failure mode that carries the most direct weight on humans, is no longer treated as a defect to eliminate but as a structural force to measure.</p><p>And where FM-01 goes, the other structural forces follow: authority diffuses, context decays, metrics drift from the reality they were built to measure. Some of those are persistent conditions. Some are acute. The diagnostic work is learning to tell the difference.</p><p>AI accelerates this physics. It doesn&#8217;t change the forces. It increases the speed at which they produce consequences. An organization with elevated responsibility compression and good human buffers can sustain that for years. The people closest to the work absorb the strain, compensate through judgment and relationships, and keep the system functioning. Add automation that removes those buffers or increases throughput without addressing the underlying compression, and the same physics produces symptoms in months instead of decades.</p><p>That is what coherence measurement is for. Not to judge organizations. Not to score them against each other. To make the structural forces visible before they become symptomatic. To give organizations a way to slow down just enough to see what&#8217;s actually happening inside their processes before they deploy automation on top of conditions they can&#8217;t see.</p><p>The physics of business at scale and speed, accelerated by AI. Sixty-four runs in, the instruments are getting sharper. That&#8217;s what this work measures.</p><div><hr></div><p>Justin Greenbaum</p><p>Greenbaum Labs</p><p>February 2026</p>]]></content:encoded></item><item><title><![CDATA[The Quiet Work]]></title><description><![CDATA[The structural work that holds organizations together. And what happens when AI removes it.]]></description><link>https://writing.justingreenbaum.com/p/the-quiet-work</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/the-quiet-work</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Tue, 24 Feb 2026 21:10:52 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/0b4d4688-cd94-4dc1-b494-e1523280fa16_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Every organization has people doing work the system can&#8217;t see.</p><p>Not their job title. Not what gets measured. The other work. The rerouting, the translating, the remembering why a decision was made three years ago when the documentation doesn&#8217;t say. When the people who made it have moved on. The judgment calls that keep a handoff from failing. The quiet compensation for a system that was never designed for the speed it&#8217;s operating at.</p><p>I&#8217;ve been calling this the quiet work. The people who do it rarely name it that way. They just know that if they stop, something breaks. And that when it works, no one notices.</p><p>This is structural work. It doesn&#8217;t appear on dashboards. It doesn&#8217;t show up in capacity models. It lives in the people who carry it, and it disappears when they&#8217;re no longer in the loop. Not because they left, but because the system stopped asking them.</p><p>Organizations have always run this way. Not because they chose to, but because this is the <a href="https://dripractice.com">physics of systems at scale</a>. Complexity generates ambiguity faster than any organization can resolve it. The gap between how the process is documented and how the work actually moves is absorbed by humans. Every day. In every function. Without acknowledgment, because acknowledging it would mean acknowledging that the system doesn&#8217;t work the way the org chart says it does.</p><p>This was sustainable when the speed of the organization was governed by the speed of the people inside it. The structural work set the pace. Judgment took time. Memory required asking the person who was there. Translation happened in hallways and one-on-ones. The system moved at human speed because humans were load-bearing.</p><p>Now AI enters. And it doesn&#8217;t know any of that.</p><p>AI operates at the speed the documented process <em>claims</em> to move. Not the speed it <em>actually</em> moves. It reads the workflow as designed and executes it as written. It doesn&#8217;t know that step four only works because someone calls procurement directly instead of submitting through the portal. It doesn&#8217;t know that the escalation path on paper hasn&#8217;t been used in two years because the real path runs through a Slack channel and a specific VP who answers after hours. It doesn&#8217;t know that the retention logic depends on a judgment call that three people in the org can make and none of them were consulted when the model was trained.</p><p>The AI isn&#8217;t wrong. The process was never right. It just had people in it who made it work anyway.</p><p>These are the same people who train new hires by saying &#8220;don&#8217;t follow the doc, here&#8217;s how it actually works.&#8221; Who carry the institutional memory that never made it into a system of record. Who see the failure before the dashboard does.</p><p>They&#8217;ve been doing structural work. Holding truth when the system distorts it. Exercising authority the org chart never granted them. Maintaining continuity across decisions, handoffs, and leadership changes that would otherwise lose their thread. They do this in the spaces where the organization&#8217;s design falls short. Every organization has these spaces. Most have more than they realize.</p><p>When you automate a process that depends on this work, you don&#8217;t get efficiency. You get exposure. The dysfunction that was always there, the authority gaps, the broken handoffs, the decisions nobody remembers making, surfaces. Not because AI created it, but because the person who was absorbing it is no longer in the loop.</p><p>This is not a technology problem. It&#8217;s not a change management problem. It&#8217;s a structural visibility problem.</p><p>Most organizations making AI deployment decisions are evaluating processes. What can be automated, what can be augmented, where there&#8217;s throughput to gain. Those are reasonable questions. But they assume that the process, as documented, is the system. It isn&#8217;t. The system is the process plus every human judgment and adaptation that makes it functional. And if you can&#8217;t see that work, you can&#8217;t account for what happens when it&#8217;s removed.</p><p>The diagnostic question isn&#8217;t &#8220;is this process automatable?&#8221; It&#8217;s &#8220;what structural work is hidden inside this process that automation will expose?&#8221;</p><p>Answering that requires a different instrument. Not performance metrics. Not process maps. Something that surfaces where authority doesn&#8217;t match responsibility, where decisions have lost their rationale, where the distance between what the organization says and how it actually operates has grown so wide that only people, specific people doing quiet work, are holding it together.</p><p>That&#8217;s what <a href="https://dripractice.com">Decision &amp; Responsibility Infrastructure&#8482;</a> is built to see. Not the process. The structure underneath it. The work that was always there, carried by people who were never given language for what they were doing, and rarely given credit.</p><p>The organizations that navigate AI adoption well won&#8217;t be the ones with the best models or the fastest deployment timelines. They&#8217;ll be the ones that understood what their people were actually doing before they automated it away.</p><p>The quiet work was never a bonus. It was the infrastructure.<br><br>-JG</p>]]></content:encoded></item><item><title><![CDATA[The Coherence Record, Edition 2]]></title><description><![CDATA[Where things stand. What broke. What held.]]></description><link>https://writing.justingreenbaum.com/p/the-coherence-record-edition-2</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/the-coherence-record-edition-2</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Tue, 17 Feb 2026 18:02:16 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8536e100-7c7c-4205-b5dc-5d3b2ffbf5bd_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3><strong>What&#8217;s Happened</strong></h3><p>Edition 1 introduced coherence as a structural property: truth, authority, and continuity reinforcing each other over time. This edition is about what happened when I tried to make it measurable.</p><p>Since the last update, I&#8217;ve been building a diagnostic pipeline. Software that takes what a company says about itself (center) and what the world says back (edge), then measures the gap. The system scores organizations on three vertices: Truth, Authority, and Continuity. The output is a structural diagnostic, not an opinion.</p><p>The target entity is Coinbase. Not because they&#8217;re broken. Because they&#8217;re public, data-rich, and operating at a scale where coherence failures become visible.</p><p>Thirteen runs later, the pipeline works.</p><div><hr></div><h3><strong>The Pipeline</strong></h3><p>It runs on a local GPU. No cloud APIs for data processing. The models are open-weight (Qwen 32B). Every claim and observation traces back to a source document with a cryptographic hash. Provenance is not optional. It&#8217;s structural.</p><p>The pipeline has four steps:</p><ol><li><p><strong>Collect</strong> &#8212; Gather what the company says (press releases, job postings) and what the world sees (customer reviews, social media, news coverage).</p></li><li><p><strong>Extract</strong> &#8212; Pull structured claims and observations from raw documents.</p></li><li><p><strong>Score</strong> &#8212; Evaluate Truth (does center match edge?) and Authority (does the entity&#8217;s voice carry weight?). An adversarial Skeptic challenges every finding. Only sustained findings enter the evidence ledger.</p></li><li><p><strong>Synthesize</strong> &#8212; Produce a diagnostic summary with failure modes, field notes, and a watch list.</p></li></ol><div><hr></div><h3>Thirteen Runs</h3><p>The run history tells the real story. Not the polished version. The actual one.</p><p><strong>Runs 001-004</strong> were scaffolding. Rule-based extraction, heuristic scoring. Establishing that the data moved through the system correctly. Run 004 produced the first real output: 296 claims, overall coherence 0.621. But 68% of the signal was unclassified. The pipeline was mechanically sound but structurally shallow.</p><p><strong>Runs 005-007</strong> introduced agent-powered extraction. The model reads the documents and produces structured claims and observations directly. Unclassified rate dropped from 68% to under 5%. Claims jumped to 1,335. Observations to 2,680. The system was seeing things the heuristics missed entirely.</p><p><strong>Run 006</strong> was the first fully unattended execution. Twelve hours, no manual intervention, every validation check passed. Run 007 confirmed repeatability.</p><p><strong>Runs 008-011</strong> were premature optimization. I tried to make it faster. Smaller model, larger batches. It worked mechanically (6x speedup) but broke structurally. Schema validation failed on 100% of first attempts. Four runs, three configurations, two killed early. The compounding lesson:</p><p><em><strong>I was changing the engine while flying.</strong></em></p><p>The root cause turned out to be a prompt conflict: two contradictory schemas in the same instruction set. Run 011 proved it wasn&#8217;t the model. The model was doing exactly what the prompt told it to do.</p><p><strong>Run 012</strong> was the reset. Back to the proven configuration, with three specific fixes. Zero schema failures across 276 batches. All validation checks passed. The largest single-run improvement in Authority (+0.13) came from adding center sources. Press releases and job postings made the organizational voice audible for the first time. Truth dropped because 1,319 center claims met only 23 edge observations. The pipeline correctly identified the imbalance.</p><p><strong>Run 013 was the payoff.</strong> was the payoff. Full edge expansion. Social media and news coverage split into individual documents. The numbers:</p><ul><li><p>1,377 claims. 4,736 observations.</p></li><li><p>Truth: 0.60. Up from 0.33. Edge depth restored.</p></li><li><p>Authority: 0.70. Capped &#8212; more on this below.</p></li><li><p>Overall coherence: 0.636.</p></li><li><p>13.7 hours. 73% of findings sustained by the Skeptic (11 of 15, 4 rejected). FM-01 detected again.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q6gU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc90a0ce8-c927-48d0-a90d-f860330e7a20_2448x1202.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q6gU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc90a0ce8-c927-48d0-a90d-f860330e7a20_2448x1202.png 424w, https://substackcdn.com/image/fetch/$s_!q6gU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc90a0ce8-c927-48d0-a90d-f860330e7a20_2448x1202.png 848w, https://substackcdn.com/image/fetch/$s_!q6gU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc90a0ce8-c927-48d0-a90d-f860330e7a20_2448x1202.png 1272w, https://substackcdn.com/image/fetch/$s_!q6gU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc90a0ce8-c927-48d0-a90d-f860330e7a20_2448x1202.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q6gU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc90a0ce8-c927-48d0-a90d-f860330e7a20_2448x1202.png" width="728" height="357.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c90a0ce8-c927-48d0-a90d-f860330e7a20_2448x1202.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:715,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:144862,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/188285208?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc90a0ce8-c927-48d0-a90d-f860330e7a20_2448x1202.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!q6gU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc90a0ce8-c927-48d0-a90d-f860330e7a20_2448x1202.png 424w, https://substackcdn.com/image/fetch/$s_!q6gU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc90a0ce8-c927-48d0-a90d-f860330e7a20_2448x1202.png 848w, https://substackcdn.com/image/fetch/$s_!q6gU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc90a0ce8-c927-48d0-a90d-f860330e7a20_2448x1202.png 1272w, https://substackcdn.com/image/fetch/$s_!q6gU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc90a0ce8-c927-48d0-a90d-f860330e7a20_2448x1202.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Score evolution across thirteen runs. The dip through runs 008&#8211;009b is the optimization regression. The recovery at 013 is edge depth.</figcaption></figure></div><div><hr></div><h3>FM-01: Responsibility Compression at the Edge</h3><p>This failure mode has appeared in every agent-powered run. Confidence 0.4 to 0.6. Across every model configuration, every data mix, every prompt version.</p><p>The pattern: responsibility concentrates where authority does not. Coinbase&#8217;s center narrative claims ownership of security, efficiency, and user experience. The edge data shows those responsibilities dispersing. Accountability for operational performance lands downstream without corresponding decision-making power. The diagnostic doesn&#8217;t measure whether Coinbase is good or bad at customer support. It measures whether the structure that owns those outcomes has the authority to change them.</p><p>This is not a model artifact. When a signal persists across nine runs with different extraction methods, different model sizes, and different data compositions, you&#8217;re looking at structure, not noise.</p><div><hr></div><h3>What&#8217;s Capped</h3><p>Two binding constraints remain:</p><p><strong>Authority is capped at 0.70</strong> because the pipeline has no employee reviews. Customers and journalists see the outside. Employees see the inside. Without that source, the Authority scorer can&#8217;t fully assess whether power and accountability are aligned within the organization. Employee reviews are the single highest-leverage data gap.</p><p><strong>Overall confidence is capped at 0.50</strong> because all data comes from a single collection period. Continuity, the third vertex, requires temporal depth. One snapshot tells you the current state. Three tell you whether it&#8217;s getting better or worse.</p><p>These caps are not limitations of the model. They&#8217;re structural constraints designed into the system. The pipeline knows what it doesn&#8217;t know.</p><div><hr></div><h3>Why This Matters Economically</h3><p>The business case for coherence is not moral. It is mechanical. When truth, authority, and continuity fall out of alignment, the organization generates friction: in revenue (customers experience something different from what&#8217;s promised), in talent (employees absorb structural failures disguised as performance problems), and in capital (investors price in narratives that the edge doesn&#8217;t support). That friction has a cost, and the cost compounds. Coherence measurement makes the friction visible before it becomes a write-down or a reorg or an exodus.</p><div><hr></div><h3>What I Learned</h3><ul><li><p><strong>Schema failures are prompt-driven, not model-driven.</strong> When extraction breaks, check the instructions before blaming the model. The model does what you tell it to do, including when you tell it two contradictory things.</p></li><li><p><strong>Don&#8217;t optimize what isn&#8217;t stable. </strong>Runs 008-011 should have been one run. The detour cost four attempts and a week. The principle is simple: get it right, prove it works, then make it fast.</p></li><li><p><strong>Edge splitting was the highest-leverage change in pipeline history.</strong> Splitting consolidated staging files into individual documents produced a 206x increase in observations. Not a model change. Not a prompt change. A data format change.</p></li><li><p><strong>Center sources make Authority audible.</strong> Adding press releases and job postings didn&#8217;t just add data. It gave the scoring agent visibility into what the organization is actually claiming. You can&#8217;t score authority if you can&#8217;t hear the voice making claims.</p></li><li><p><strong>The record exists for the truth.</strong> Every run has a review. Every review documents what happened, what broke, and what was learned. The run reviews are not retrospective polish. They&#8217;re written the same day, before the lessons have time to soften.</p></li></ul><div><hr></div><h3>What&#8217;s Next</h3><p>Run 014 is staged. Employee reviews are being added to the collection. A normalizer tool now converts manually-sourced Glassdoor data into pipeline-compatible format. When that source comes online, Authority uncaps from 0.70.</p><p>The scoring agents are being tuned. Few-shot examples now show the Authority agent what structural differentiation looks like. Not just what score to produce, but how to reason about power, accountability, and organizational voice as distinct signals.</p><p>After 014, the focus shifts to repeatability. A second entity. A second collection period. The system needs to prove it measures coherence, not just Coinbase.</p><p>Decision &amp; Responsibility Infrastructure was filed with the USPTO on February 10, 2026. Serial number 99645812. It names the field, not a product.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TC4f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e309d9e-6be5-4d29-94d7-0dfc28f26666_2456x1116.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TC4f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e309d9e-6be5-4d29-94d7-0dfc28f26666_2456x1116.png 424w, https://substackcdn.com/image/fetch/$s_!TC4f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e309d9e-6be5-4d29-94d7-0dfc28f26666_2456x1116.png 848w, https://substackcdn.com/image/fetch/$s_!TC4f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e309d9e-6be5-4d29-94d7-0dfc28f26666_2456x1116.png 1272w, https://substackcdn.com/image/fetch/$s_!TC4f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e309d9e-6be5-4d29-94d7-0dfc28f26666_2456x1116.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TC4f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e309d9e-6be5-4d29-94d7-0dfc28f26666_2456x1116.png" width="1456" height="662" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e309d9e-6be5-4d29-94d7-0dfc28f26666_2456x1116.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:662,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:440797,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/188285208?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e309d9e-6be5-4d29-94d7-0dfc28f26666_2456x1116.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TC4f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e309d9e-6be5-4d29-94d7-0dfc28f26666_2456x1116.png 424w, https://substackcdn.com/image/fetch/$s_!TC4f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e309d9e-6be5-4d29-94d7-0dfc28f26666_2456x1116.png 848w, https://substackcdn.com/image/fetch/$s_!TC4f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e309d9e-6be5-4d29-94d7-0dfc28f26666_2456x1116.png 1272w, https://substackcdn.com/image/fetch/$s_!TC4f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e309d9e-6be5-4d29-94d7-0dfc28f26666_2456x1116.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Thirteen runs. Every score, every status, every lesson. The record is the proof.</figcaption></figure></div><div><hr></div><h3>AR-001 Still Holds</h3><p>Edition 1 established the first Accountability Record: <em>*<strong>Automation may observe, summarize, and suggest, but may not decide.*</strong></em></p><p>The pipeline embodies this. Every diagnostic output is a suggestion. Every finding requires human review. The Skeptic debate mechanism is adversarial, but the final judgment is not automated.</p><p>Thirteen runs in, the governance hasn&#8217;t changed. The capability has grown around it without eroding it.</p><p>That&#8217;s coherence in practice.</p><div><hr></div><p><strong>Justin Greenbaum</strong></p><p>Greenbaum Labs</p><p>February 2026</p>]]></content:encoded></item><item><title><![CDATA[The Coherence Record, Edition 1]]></title><description><![CDATA[An executive builds diagnostic infrastructure on owned hardware using public data.]]></description><link>https://writing.justingreenbaum.com/p/the-coherence-record-edition-1</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/the-coherence-record-edition-1</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Tue, 10 Feb 2026 14:52:37 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ec96171a-95d7-4d26-986b-786dc7381aba_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>on Owned Hardware<br>and Maintains a Public Build Record</p><p>Justin Greenbaum<br>Greenbaum Labs<br>February 2026</p><p><strong>Edition 1</strong></p><div><hr></div><p>This is not a product announcement or a finished framework. It is a build record. It documents what it takes for an executive to design, test, and trust diagnostic infrastructure using public data and owned hardware. The technical and the strategic are not separated here, because in the work itself they were never separate.</p><div><hr></div><h2>Forty Hours and Counting</h2><p>There is a moment in every build where you question everything. Not the architecture. Not the strategy. Not the market. Everything.</p><p>It is the kind of doubt that sits in your chest at 2 a.m. when the GPU has been running for ten hours, every request is timing out, and you cannot determine why.</p><p>I reached that moment more than once.</p><p>Over forty hours across six days, I debugged the same pipeline. A parameter in the wrong nesting level. A prompt that was too permissive. A working process I terminated because I could not observe progress. Each failure different. Each one invisible until it was not.</p><p>This is what building infrastructure actually looks like. Not a highlight reel. Not an architecture diagram. The real sequence of decisions, mistakes, and corrections that separates an idea from a system that can be trusted.</p><div><hr></div><h2>What Coherence Means in This Work</h2><p>This work does not begin with a framework. It begins with an observation.</p><p>Organizations routinely say one thing and produce another.</p><p>The distance between those two is not abstract. It appears in public records, regulatory actions, consumer complaints, job postings, and narrative shifts over time. That distance can be observed, compared, and measured.</p><p>I use the term coherence to describe the structural integrity of that relationship.</p><p>Not alignment, which implies a static state.<br>Not integrity, which carries moral weight.<br>Coherence is descriptive. It reflects whether an organization&#8217;s internal narrative holds when tested against external reality.</p><p>The diagnostic model used here evaluates three dimensions, each derived exclusively from public sources.</p><p><strong>Truth</strong><br>The relationship between an organization&#8217;s stated claims and what outside observers experience. A gap does not require falsehood. Omission is sufficient.</p><p><strong>Authority</strong><br>The relationship between responsibility and decision-making power as expressed through role design, escalation patterns, and ownership signals.</p><p><strong>Continuity</strong><br>Stability of narrative over time. Repeated shifts without acknowledgment indicate structural drift rather than adaptation. This dimension requires multiple collection periods and emerges longitudinally.</p><p>Together, these form the Coherence Triangle. The question was never whether coherence could be measured. The question was whether I could build the infrastructure to measure it independently, on hardware I control, using models I understand, and produce diagnostics I trust.</p><div><hr></div><h2>Why the System Runs on Owned Hardware</h2><p>This project could have been implemented using cloud inference and commercial APIs. That approach is faster, cheaper in the short term, and easier to scale.</p><p>I chose not to use it.</p><p>If a diagnostic claims to measure the gap between narrative and reality, the diagnostic itself cannot depend on infrastructure I do not control. A system designed to surface fragility cannot be built on rented dependencies that can change without notice.</p><p>The environment consists of:</p><ul><li><p>NVIDIA DGX Spark for inference</p></li><li><p>Synology NAS for raw data storage</p></li><li><p>Mac Studio for orchestration and development</p></li></ul><p>All models run locally. All data is stored locally. All processing occurs on a private network.</p><p>The trade-offs are explicit. Inference takes minutes instead of seconds. Full diagnostic runs take hours. Throughput is constrained.</p><p>The benefit is traceability.</p><p>When a score is produced, I can identify exactly which data sources, model weights, prompts, and parameters contributed to it. That is not a performance feature. It is the foundation of trust in the measurement.</p><p>The economics reinforce the decision. Run 005 processed approximately 1.4 million tokens. On commercial APIs, that range spans from tens of cents to double-digit dollars depending on provider and configuration. At scale, those costs compound quickly. On owned hardware, the marginal cost is power draw. Thirteen hours at sustained utilization produced no invoice.</p><p>That is not optimization. It is independence.</p><div><hr></div><h2>Center and Edge Data</h2><p>A diagnostic reflects only the data it examines.</p><p>This system separates sources into two categories.</p><p><strong>Center data</strong><br>Materials an organization publishes intentionally. Press releases. Job postings. Regulatory filings. Earnings transcripts. Official communications.</p><p><strong>Edge data</strong><br>Public responses to those claims. Consumer complaints. Regulatory actions. Employee reviews. Other externally observable signals.</p><p>For the first diagnostic subject, Coinbase, all data was collected from publicly accessible sources, including:</p><ul><li><p>job postings retrieved via the Greenhouse API</p></li><li><p>SEC EDGAR filings</p></li><li><p>consumer complaints filed with the CFPB</p></li></ul><p>No internal systems, non-public documents, or privileged access were used.</p><p>The dataset is incomplete by design. Employee review platforms, earnings transcripts, and social media signals were not included in this run. The diagnostic explicitly records those omissions.</p><p>A system that claims to measure coherence must be able to state what it does not know.</p><div><hr></div><h2>How the Pipeline Operates</h2><p>The pipeline runs in four stages, each validated before proceeding.</p><p><strong>Collect</strong><br>Documents are retrieved from configured public sources. File counts, formats, and accessibility are validated.</p><p><strong>Extract</strong><br>Each document is processed by an extraction agent that identifies diagnostically relevant claims and observations and returns structured JSON. In the agent-powered run, this stage processed 844 documents across 844 consecutive agent calls with near-zero failure.</p><p><strong>Score</strong><br>Extracted content is evaluated across Truth, Authority, and Continuity. In agent mode, this includes structured debate between specialized agents and a Skeptic that can sustain or reject findings.</p><p><strong>Synthesize</strong><br>The system produces a diagnostic summary, supporting field notes, a watch list, an overall coherence score, and an explicit data quality grade.</p><p>The pipeline runs in two modes:</p><ul><li><p>rule-based, which completes in under a second using pattern matching</p></li><li><p>agent-powered, which takes hours using multi-agent inference</p></li></ul><p>Both produce results. The purpose of this build was to determine whether the agent-powered architecture materially improves diagnostic quality.</p><p>It does.</p><div><hr></div><h2>What Forty Hours Teaches You</h2><p>Synthetic data validated the mechanics. Real data exposed reality.</p><p>The rule-based extractor classified only 31.7 percent of real content. Most material fell outside predefined patterns. This was expected.</p><p>The agent pipeline was intended to read context and apply judgment. Initially, it returned nothing.</p><p>The cause was a single misplaced parameter. A model behavior flag was passed in the wrong location. The API accepted the request. The model ran. The output buffer was consumed internally. No usable output returned.</p><p>One parameter. Wrong nesting level. No error. No warning.</p><p>After correcting that, the model responded exhaustively. Each document produced more than ten thousand characters of structured JSON. Perfectly formatted. Completely unusable.</p><p>The problem was not infrastructure or model capability. It was the prompt. The instruction asked for everything, and the model complied.</p><p>Constraining the request to the top five diagnostically important items per document stabilized output immediately.</p><p>The lesson is not about prompt technique. It is that failure can live at any layer of the system, and it does not announce which one.</p><p>A later run appeared to hang. No logs. No output. No visible progress. The pipeline was working the entire time. Logging was not configured. Progress was invisible. I terminated a process that was nearly halfway complete.</p><p>Two print statements resolved it.</p><p>This is the kind of failure that does not appear in summaries. Progress you cannot observe is progress you will eventually destroy.</p><div><hr></div><h2>What the Diagnostic Found</h2><p>The agent-powered diagnostic processed 844 documents over thirteen hours and classified 95.1 percent of all content.</p><p>The overall coherence score was <strong>0.609</strong>, down from <strong>0.621</strong> in the rule-based baseline.</p><p>This is not regression. It is honesty.</p><p>Truth improved slightly as agents found more evidence on both sides of the narrative. The dominant pattern remained omission rather than contradiction.</p><p>Authority decreased materially. The agents identified concentration and diffusion patterns that rule-based logic could not detect.</p><p>Continuity was not scored. It requires longitudinal measurement.</p><p>Data quality was graded <strong>C</strong> due to incomplete source coverage. The diagnostic states this explicitly.</p><p>Inside the data, 645 observations clustered around the product experience. That signal emerged only because agents could read context that patterns could not.</p><div><hr></div><h2>The Evidence Chain Failure</h2><p>One critical subsystem failed.</p><p>The scoring agents produced substantive findings, but could not reliably cite the specific claim and observation identifiers that supported them. The Skeptic rejected every finding.</p><p>The scores are valid. The evidence ledger is incomplete.</p><p>This is a wiring problem, not a capability problem. Shorter identifier aliases are being introduced to restore provenance integrity. The failure is documented here because the framework requires it.</p><p>If a system measures gaps, it must disclose its own.</p><div><hr></div><h2>The Executive Who Builds</h2><p>I am not an engineer.</p><p>I am an executive who decided that understanding infrastructure is now a leadership capability.</p><p>As AI compresses the distance between intent and consequence, leadership that operates only through delegation loses resolution. The value is no longer in deciding. It is in understanding what decisions actually require.</p><p>No vendor briefing explains where reality resists abstraction. You learn that by building.</p><div><hr></div><h2>Why This Record Is Public</h2><p>There is no established reference for this work.</p><p>This document exists as a record, not a guide.</p><p>It preserves decision context and holds the work accountable to its own standards. Coherence Diagnostics measures the gap between what organizations say and what they do. The build itself must be coherent.</p><p>This record is the edge data for the project&#8217;s own narrative.</p><div><hr></div><h2>The Record Begins</h2><p>This is Edition 1.</p><p>What exists now:</p><ul><li><p>an extraction engine that comprehends over 95 percent of content</p></li><li><p>a multi-agent scoring system with an active Skeptic</p></li><li><p>a synthesis layer that grades its own data quality</p></li><li><p>owned infrastructure with zero cloud dependency</p></li></ul><p>And an evidence chain that is not yet complete.</p><p>Future editions will document what changes, why, and whether those changes improve coherence measurement over time.</p><p>If coherence matters, it must be observable.<br>If diagnostics matter, they must be accountable.<br>If an executive claims to understand the infrastructure, there must be evidence.</p><p>This is that evidence.</p><div><hr></div><p><strong>Justin Greenbaum</strong><br>Founder, Greenbaum Labs</p><p>Building diagnostic infrastructure to measure the gap between what organizations say and what they do.</p><p>This is <strong>The Coherence Record, Edition 1</strong>, published at writing.justingreenbaum.com.</p>]]></content:encoded></item><item><title><![CDATA[Coherence]]></title><description><![CDATA[How the work moves through the world.]]></description><link>https://writing.justingreenbaum.com/p/coherence</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/coherence</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Tue, 03 Feb 2026 15:24:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Viyw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Coherence is a property of a system in which truth, authority, and continuity reinforce each other over time, reducing the need for corrective effort.</p><p>When systems lose coherence, people absorb the cost.</p><p>I didn&#8217;t find this in a book. I found it by living inside complex systems for twenty years, noticing the same patterns, and writing to make sense of them. The framework didn&#8217;t arrive fully formed. It crystallized slowly, through attention, through failure, through watching what held and what didn&#8217;t.</p><p>This is what I&#8217;ve been working toward. Not a theory to defend, but a way of seeing that finally has a name.</p><div><hr></div><h2>The Triangle</h2><p>Three vertices. Three things that have to hold together.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Viyw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Viyw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png 424w, https://substackcdn.com/image/fetch/$s_!Viyw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png 848w, https://substackcdn.com/image/fetch/$s_!Viyw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png 1272w, https://substackcdn.com/image/fetch/$s_!Viyw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Viyw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png" width="1456" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2687828,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://writing.justingreenbaum.com/i/186745500?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Viyw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png 424w, https://substackcdn.com/image/fetch/$s_!Viyw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png 848w, https://substackcdn.com/image/fetch/$s_!Viyw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png 1272w, https://substackcdn.com/image/fetch/$s_!Viyw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F766f42b3-ef0f-46a2-b48e-0d30e9f8bfa7_2574x1358.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Truth</strong> &#8212; what actually is, not what we wish or what we report. The signal, not the story we tell about the signal. In operations, this is the difference between what the dashboard shows and what&#8217;s actually happening on the floor. In medicine, it&#8217;s the difference between the chart and the patient.</p><p><strong>Authority</strong> &#8212; not positional power, but earned weight. The kind that comes from showing up, from being right when it mattered, from carrying responsibility visibly over time. A frontline supervisor who flags a process gap. A nurse who speaks up. A junior analyst who catches the error.</p><p><strong>Continuity</strong> &#8212; authenticity proven over time. The through-line that lets people trust the pattern will hold tomorrow because it held yesterday. This is why teams with shared history outperform teams with more talent. The pattern is known. The gaps are covered.</p><p>When all three vertices are honest, the triangle holds. When one drifts, the others compensate. When two drift, the structure strains. When all three are compromised, coherence breaks, and humans absorb the gap.</p><div><hr></div><h2>The Flow</h2><p>This is how the work moves:</p><p><strong>Lived Practice</strong> &#8594; The ground. Real operations, real pressure, real decisions. This is where truth originates, not in frameworks or models. You learn what works during the outage, not during the planning session.</p><p><strong>Presence &amp; Trust</strong> &#8594; Before any framework lands, you have to demonstrate attention. Show continuity. Establish safety. No teaching. No agenda. Just showing up. This is why the best leaders listen before they speak.</p><p><strong>Shared Language</strong> &#8594; Name what people already recognize. Reduce ambiguity. Create alignment without instruction. The aviation industry calls this &#8220;crew resource management.&#8221; The military calls it &#8220;commander&#8217;s intent.&#8221; In operations, it&#8217;s the playbook everyone actually trusts. The words should feel like relief, not education.</p><p><strong>Observation</strong> &#8594; Capture early signals. Notice drift before failure. The goal is to preserve ambiguity long enough to understand it, not to force conclusions. Good incident review does this. Bad incident review assigns blame.</p><p><strong>Diagnosis</strong> &#8594; Name the repeatable breakdowns. Create durable reference language. Support people under pressure with patterns they can recognize and address. A shared vocabulary for what&#8217;s going wrong.</p><p><strong>Orientation</strong> &#8594; The loop closes. When signals appear again, the language exists. The diagnosis is available. No rediscovery required. This is how institutions learn, when they learn at all.</p><div><hr></div><h2>The Diagnostic Lens</h2><p>What I watch for:</p><p><strong>Early signals</strong> (Functional Field Notes)</p><ul><li><p>Signal Misread &#8212; the data is there but the interpretation drifts. The dashboard is green but the phones are ringing.</p></li><li><p>Responsibility Compression &#8212; ownership narrows as pressure increases. One person becomes the load-bearing wall.</p></li><li><p>Context Collapse &#8212; decisions made without necessary background. The person deciding doesn&#8217;t know what the person executing knows.</p></li><li><p>Meaning Drift &#8212; words stay the same but their meaning shifts. &#8220;Escalation&#8221; means something different to everyone in the room.</p></li></ul><p><strong>Breakdowns</strong> (Failure Modes)</p><ul><li><p>Responsibility Without Authority &#8212; you own the outcome but not the decision. Common in operations, common in healthcare, common everywhere matrices exist.</p></li><li><p>Metric Shadowing &#8212; the measure becomes the target, obscuring what it was meant to illuminate. Goodhart&#8217;s Law in action.</p></li><li><p>Escalation Theater &#8212; the process exists but doesn&#8217;t function. Everyone followed the steps. Nothing was addressed.</p></li><li><p>Decision Displacement &#8212; choices pushed to people without context to make them. The frontline shouldn&#8217;t be setting policy during an outage.</p></li></ul><p>The crosswalk between early signals and failure modes is where intervention is possible. Once a signal hardens into a failure mode, the cost is already being paid.</p><div><hr></div><h2>Where This Shows Up</h2><p>These patterns aren&#8217;t theoretical. They appear in:</p><p><strong>Operations &amp; CX</strong> &#8212; when the escalation path exists on paper but collapses under pressure. When frontline teams see the problem but the structure doesn&#8217;t support surfacing it. When metrics tell one story and customers tell another.</p><p><strong>Healthcare</strong> &#8212; when a nurse sees something wrong but the hierarchy discourages speaking up. When handoff protocols exist but context doesn&#8217;t transfer.</p><p><strong>Aviation</strong> &#8212; when crew resource management fails and copilots defer to captains against their own judgment. When checklists become ritual instead of verification.</p><p><strong>Any system where stakes are high and information is incomplete.</strong></p><p>The pattern is the same. The cost is always human.</p><div><hr></div><h2>What This Is</h2><p>I spent a long time not knowing what to call this work.</p><p>I knew I was noticing something. I knew the patterns were real. I wrote to try to make sense of them, revised, threw things out, started again. It took years to find language that didn&#8217;t feel borrowed or forced.</p><p>Coherence is what I arrived at. A diagnostic framework. A way of seeing where systems are holding and where they&#8217;re drifting.</p><p>It applies to organizations. It applies to teams. It applies to decisions you make alone at 2am when the information is incomplete and the stakes are real.</p><p>The work is noticing where truth, authority, and continuity are pulling apart, naming it clearly, and creating the conditions where it can be addressed.</p><p>This is what I&#8217;m building now. Not because someone asked for it, but because the pattern finally has a name and I can&#8217;t unsee it.</p><p>Coherence is the cornerstone. Authenticity is continuity proven over time.</p><div><hr></div><p><em>Explore the interactive version: <a href="https://justingreenbaum.com/coherence.html">justingreenbaum.com/coherence</a></em></p><p><em>More on how I&#8217;m building this infrastructure: <a href="https://www.greenbaumlabs.com">greenbaumlabs.com</a></em></p>]]></content:encoded></item><item><title><![CDATA[Authenticity Over Time]]></title><description><![CDATA[The faces look right.]]></description><link>https://writing.justingreenbaum.com/p/authenticity-over-time</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/authenticity-over-time</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Tue, 27 Jan 2026 11:49:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-OJJ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9353dbd-7265-40d5-a2e1-b4245008ec4d_512x512.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The faces look right.</p><p>The timing of the laugh. The hand gestures on video. The small imperfections in a voice note. Everything lands where it should. The posts read fine too. Coherent. Confident. Plausible.</p><p>And yet something feels off.</p><p>Not wrong, exactly. Just unmoored. Like the words belong to a person who appeared last week, rather than someone who has been here all along.</p><p>For a long time, authenticity worked this way. Presence was proof. A real face, a real voice, the right timing in the moment were enough to establish trust. We believed what we could see live, hear unedited, or experience directly because those signals were difficult to fabricate.</p><p>That was the contract.</p><p>But the ground has shifted.</p><p>Today, presence is cheap. Faces can be generated. Voices can be cloned. Timing can be simulated. Even imperfection can be designed. What once served as evidence now functions as aesthetic.</p><p>The problem is not deception. It is continuity.</p><p>Authenticity is no longer a moment. It is a record.</p><p>What we are losing is not the ability to tell whether something looks real. We are losing the ability to tell whether it belongs to anything that existed before.</p><p><strong>The Collapse of Moment-Based Proof</strong></p><p>Most conversations about authenticity still focus on detection. How to tell if something is fake. How to watermark content. How to identify manipulation. These efforts matter, but they miss the deeper failure mode.</p><p>The real breakdown is not that false things can look true. It is that true things can appear without a past.</p><p>A convincing statement with no lineage. A confident voice with no history. A polished worldview that arrived fully formed.</p><p>When everything can be generated instantly, the question stops being &#8220;Is this real?&#8221; and becomes &#8220;Where did this come from?&#8221;</p><p>Without continuity, authenticity becomes indistinguishable from performance.</p><p><strong>Time as the Missing Signal</strong></p><p>What cannot be faked easily is not presence. It is persistence.</p><p>Real people contradict themselves, then correct course. They change slowly. They leave traces of earlier thinking behind. They accumulate context rather than replacing it.</p><p>Authenticity over time looks uneven. It contains gaps. It carries the marks of revision, restraint, and unfinished thought.</p><p>This is why time matters.</p><p>Time forces coherence. It reveals whether a belief holds under pressure, whether a posture survives fatigue, whether a system continues to behave the same way when attention drops.</p><p>Moment-based authenticity asks, &#8220;Does this feel right now?&#8221; Time-based authenticity asks, &#8220;Does this still make sense later?&#8221;</p><p>Only one of those questions scales.</p><p><strong>Why This Matters Now</strong></p><p>As AI systems become embedded in everyday communication, they will not just generate content. They will generate confidence. They will speak fluently, quickly, and convincingly. They will compress years of thinking into seconds.</p><p>In that environment, trust cannot be built on eloquence or speed. Those are now table stakes.</p><p>Trust will shift toward things that unfold slowly: ideas that recur before they evolve, frameworks that constrain themselves, people whose work remains legible months later without reinforcement.</p><p>The future of authenticity is not louder signals. It is quieter records.</p><p><strong>Continuity as Infrastructure</strong></p><p>Authenticity over time does not require constant broadcasting. In fact, it is often undermined by it. What matters is not how often something is said, but whether it behaves the same way when it is not being explained.</p><p>This is why continuity becomes a form of infrastructure.</p><p>Logs instead of proclamations. Artifacts instead of opinions. Patterns instead of moments.</p><p>When someone can point to what they thought last year, what they changed, and what remained stable, trust has something to attach to. When they cannot, no amount of present-moment confidence can substitute for the absence.</p><p>This is not about nostalgia or permanence. It is about traceability.</p><p><strong>The New Shape of Trust</strong></p><p>In the years ahead, authenticity will be less about expression and more about alignment across time.</p><p>We will trust people whose work evolves without disowning its past, holds up without constant explanation, and survives silence.</p><p>The signal will not be intensity. It will be consistency.</p><p>Not the consistency of branding, but the consistency of behavior.</p><p>Not the sameness of messaging, but the coherence of judgment.</p><p><strong>Closing</strong></p><p>Authenticity is no longer something you prove in a moment.</p><p>It is something you earn by staying recognizable to yourself over time.</p><p>In a world that can generate confidence on demand, the rarest thing will be a life, a body of work, or a way of thinking that can be followed backward without collapsing.</p><p>That is the new threshold.</p><p>Not &#8220;Does this feel real?&#8221; But &#8220;Has this been real long enough to trust?&#8221;</p>]]></content:encoded></item><item><title><![CDATA[Conditionally: Setting the Conditions for Change in Complex Systems]]></title><description><![CDATA[For a long time, leadership treated change as a communication problem.]]></description><link>https://writing.justingreenbaum.com/p/conditionally-setting-the-conditions</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/conditionally-setting-the-conditions</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Mon, 12 Jan 2026 13:03:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-OJJ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9353dbd-7265-40d5-a2e1-b4245008ec4d_512x512.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>For a long time, leadership treated change as a communication problem.</p><p>If people understood the rationale, if leaders showed up visibly, if the rollout was sequenced correctly, the organization would follow. Resistance could be managed. Adoption could be driven.</p><p>That approach worked in an earlier operating environment. Systems were slower. Authority was clearer. The distance between decision and execution was small enough to be bridged with words.</p><p>What leaders are encountering now is not resistance.</p><p>It is something quieter and more structural.</p><div><hr></div><h3>What leaders are actually experiencing</h3><p>Most organizations today are communicating more than ever.</p><p>There are more updates, more town halls, more decks, more alignment rituals. Leaders are present. Messaging is frequent. Intent is visible.</p><p>And yet clarity does not hold.</p><p>Decisions do not land cleanly. Escalation increases instead of resolving issues. Interpretation fragments across teams. Trust erodes without a visible breaking point.</p><p>The system looks active.<br>Very little integrates.</p><p>This is not a failure of effort or care.<br>It is a mismatch between what is being introduced and what the system can currently absorb.</p><div><hr></div><h3>What change management assumed</h3><p>Change management was built on a set of assumptions that were once reliable.</p><p>It assumed authority already existed and would be exercised. It assumed responsibility was routed clearly enough to support adoption. It assumed people needed persuasion more than preparation. It assumed communication could close the gap between intent and behavior.</p><p>When those conditions were present, managing change meant managing messaging, sequencing, and reaction.</p><p>Today, those conditions are often absent.</p><p>No amount of explanation compensates for unclear authority.<br>No volume of communication stabilizes meaning when interpretation is fragmented.</p><p>Under these conditions, communication often accelerates confusion rather than resolving it.</p><div><hr></div><h3>A note on where this comes from</h3><p>This perspective is not formed from the outside.</p><p>I have led change management teams. I have been trained and certified in formal models, including ADKAR. I have worked in environments where change was not theoretical. The stakes were high, visible, and immediate. When things failed, they did so publicly and at scale.</p><p>In those conditions, discipline mattered. Sequencing mattered. Communication mattered. For a long time, the methods worked as designed.</p><p>What changed was not the people, the intent, or the effort.</p><p>What changed was the system those methods were operating inside.</p><div><hr></div><h3>The real issue is not resistance. It is readiness.</h3><p>When leaders say an organization is not ready, that statement is often interpreted as emotional or cultural.</p><p>Readiness is not about mindset.<br>It is about system capacity.</p><p>An unready system misinterprets signals. It politicizes information. It overcorrects or freezes. It routes responsibility downward and escalates risk upward without resolution.</p><p>In these environments, even well intentioned communication becomes destabilizing. Messages arrive before the system has the conditions required to integrate them.</p><p>The issue is not how much information is shared.</p><p>The issue is whether the system is prepared to receive it without losing agency and judgment.</p><div><hr></div><h3>Conditioning is the leadership work now</h3><p>Conditioning is not persuasion.</p><p>It is preparation.</p><p>It is the work leaders do before introducing change, not after resistance appears.</p><p>Conditioning focuses on different questions.</p><p>Is authority clear enough to support this decision.<br>Is interpretation stable enough to hold new information.<br>Will responsibility route correctly or compress at the edge.<br>Is the system calm enough to absorb change without distortion.</p><p>When the answer to those questions is no, better messaging is not the responsible move.</p><p>Setting the conditions is.</p><p>That may require slowing expression.<br>Reducing noise.<br>Clarifying decision rights.<br>Naming unresolved conditions instead of smoothing over them.<br>Allowing silence to do work communication cannot.</p><p>This is not avoidance.<br>It is stewardship.</p><div><hr></div><h3>Where change management and communications fit</h3><p>This shift does not make change management or communications work obsolete.</p><p>It places them correctly.</p><p>Change management remains a valuable discipline once the conditions for change are present. When authority is clear, interpretation is stable, and responsibility is correctly routed, change management does what it has always done well. It helps organizations introduce decisions, support adoption, and guide people through transition.</p><p>The problem is not the discipline.</p><p>The problem is when it is asked to compensate for conditions that are not yet in place.</p><p>In many modern systems, communication and change practices are asked to smooth over unresolved authority, stabilize meaning that has not yet settled, and create reassurance in environments that remain structurally uncertain.</p><p>That is not a failure of communication.<br>It is a signal that the system is not yet ready for communication to carry that load.</p><p>Conditioning does not replace these disciplines. It precedes them. It creates the conditions that allow communication and change management to operate without distortion or overreach.</p><p>When the system is conditioned, communication becomes lighter. Messages land with less effort. Change practices feel supportive rather than compensatory. Trust holds because expression matches reality.</p><p>This is not a critique of the people doing the work. It is an acknowledgment of how much has been asked of them as systems have grown more complex.</p><p>Conditioning restores the boundary. It ensures that communication and change management are applied when they can do what they are best at, rather than being forced to absorb structural gaps they cannot resolve.</p><div><hr></div><h3>What effective leaders are doing differently</h3><p>Leaders who navigate change well today are not louder or more charismatic.</p><p>They are more restrained.</p><p>They delay expression until decisions are real.<br>They speak less, but with consequence.<br>They resist premature reassurance.<br>They prepare systems before introducing change.<br>They treat silence as a legitimate leadership act.</p><p>They do not manage change.</p><p>They condition systems so that when change arrives, people retain agency and judgment rather than losing them.</p><div><hr></div><h3>A closing recognition</h3><p>Change management was not wrong.</p><p>It was built for a different operating environment.</p><p>What replaced it is not resistance or apathy.</p><p>It is complexity.</p><p>Complexity does not respond to persuasion.<br>It responds to preparation.</p><p>The organizations that change well today are not convinced.</p><p>They are conditioned.</p>]]></content:encoded></item><item><title><![CDATA[Before Ownership, There Must Be Language]]></title><description><![CDATA[Most organizations are not failing because they lack accountability.]]></description><link>https://writing.justingreenbaum.com/p/before-ownership-there-must-be-language</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/before-ownership-there-must-be-language</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Mon, 05 Jan 2026 13:04:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-OJJ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9353dbd-7265-40d5-a2e1-b4245008ec4d_512x512.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most organizations are not failing because they lack accountability.</p><p>They are failing because they lack a shared way to interpret what is happening.</p><p>This gap shows up in familiar ways. Customer experience scores fluctuate. Escalations rise. Teams debate ownership. Leaders ask who is responsible. Everyone senses something is off, yet progress stalls or conversations turn defensive.</p><p>What looks like resistance is usually something quieter.</p><p>People are being asked to respond to signals the system has not made safe to interpret.</p><div><hr></div><h3>Responsibility Lives in Infrastructure</h3><p>In complex systems, responsibility rarely lives in a person.<br>It lives in infrastructure.</p><p>Routing logic. Incentives. Metrics. Escalation paths. Decision rights. And critically, <strong>interpretive infrastructure</strong>: the shared language used to describe what is happening.</p><p>These structures shape outcomes long before intent or effort enter the picture.</p><p>When responsibility is framed as personal, systems get defended.<br>When responsibility is framed as infrastructural, systems can be examined.</p><p>Without interpretive infrastructure, failure cannot be made observable without blame.</p><p>This distinction matters because experience is not owned.<br>It is produced by how systems behave under real conditions.</p><div><hr></div><h3>Systems Built for Motion</h3><p>Most modern organizations were built for mobility.</p><p>They optimized for speed, scale, channel shifts, and flexibility. Work could move quickly. Decisions could route around friction. Teams adapted rapidly. That design made sense and it worked.</p><p>But mobility was never meant to be permanent.</p><p>When systems never settle, instability becomes ambient. Teams remain alert. Customers feel friction. Leaders experience noise instead of clarity.</p><p>Sensitivity increases. Pain points surface. CX metrics spike.</p><p>This is not incompetence.<br>It is a system under continuous motion.</p><p>The problem is not that organizations are seeing more failure.<br>The problem is that they are seeing failure without a stable way to describe it.</p><div><hr></div><h3>Why Defensiveness Appears</h3><p>When language fragments, every signal feels personal.</p><p>Different teams use the same words to mean different things. Ownership. Accountability. Quality. Efficiency. Customer experience. Everyone recognizes their importance, but few agree on what they describe in practice.</p><p>So when a metric rises, someone feels accused.<br>When an escalation occurs, someone feels exposed.<br>When feedback surfaces, someone prepares a defense.</p><p>Defensiveness is not cultural weakness.<br>It is a rational response to ambiguity.</p><p>People do not resist fixing systems.<br>They resist being blamed for systems they did not design and cannot clearly see.</p><p>Until failure is observable without blame, organizations will continue to defend instead of learn.</p><div><hr></div><h3>A Familiar Example</h3><p>Consider a large healthcare organization.</p><p>Patients move across appointments, portals, specialists, labs, billing systems, and follow ups. Each function performs its role competently. Each team cares. Each speaks the language of its discipline.</p><p>Access teams talk about availability.<br>Clinicians talk about quality of care.<br>IT talks about reliability and security.<br>Billing talks about accuracy and compliance.<br>Experience teams talk about satisfaction scores.</p><p>When patient complaints rise, everyone sees it.</p><p>But they do not see the same thing.</p><p>Each group explains the issue accurately from its own vantage point. No one is wrong. And the patient experience still degrades.</p><p>Meetings follow a predictable pattern. Data is reviewed. Context is added. Intent is defended. Ownership is debated. The room grows tense, not because people do not care, but because the system is asking them to respond without shared interpretation.</p><p>Each group leaves believing someone else needs to fix it.</p><p>What is missing is not effort or accountability.<br>What is missing is a way to observe how the system behaves across boundaries.</p><div><hr></div><h3>The Limits of Ownership Conversations</h3><p>This is why ownership debates rarely resolve experience problems.</p><p>Statements like &#8220;everyone owns CX&#8221; acknowledge shared responsibility but offer no shared interpretation. Accountability dissolves instead of clarifying.</p><p>Calls for singular ownership often escalate tension. They imply control without resolving the underlying confusion about what is actually happening.</p><p>Both approaches skip the same prerequisite.</p><p>Before responsibility can be assigned, failure must be observable without blame.</p><div><hr></div><h3>Language as Stabilization</h3><p>Stability does not begin with reorganization.<br>It begins with shared language.</p><p>When people agree on terms, patterns emerge.<br>When patterns emerge, blame loses traction.<br>When blame fades, systems become safe to examine.</p><p>Shared language allows organizations to talk about breakdowns without moral judgment. To distinguish signal from noise. To describe how systems behave under pressure rather than who failed.</p><p>This is not semantics.<br>It is infrastructure.</p><p>Without shared language, failure remains anecdotal.<br>With it, failure becomes diagnostic.</p><div><hr></div><h3>The Turn Toward Creation</h3><p>Many organizations try to create growth while standing in unstable environments. They ask teams to innovate while interpretation remains fragmented. They ask for better outcomes without first stabilizing how the system is understood.</p><p>Creation does not emerge from noise.<br>It emerges from coherence.</p><p>Before asking people to own outcomes, we must give them a way to see the system together.</p><p>Before optimizing experience, we must stabilize the environment the experience occurs in.</p><p>Once failure is observable without blame, teams can say &#8220;this is the same breakdown pattern we saw last quarter&#8221; instead of &#8220;CX is failing again,&#8221; and begin learning rather than reacting.</p><div><hr></div><h3>Why This Matters</h3><p>When failure can be named without shame, it can be studied.<br>When it can be studied, it can be understood.<br>When it is understood, it can be addressed.</p><p>This is the work beneath the work.</p><p>And it starts with language that makes failure observable without blame.</p>]]></content:encoded></item><item><title><![CDATA[Why Do We Fail?]]></title><description><![CDATA[Failure as evidence, not event]]></description><link>https://writing.justingreenbaum.com/p/why-do-we-fail</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/why-do-we-fail</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Fri, 26 Dec 2025 17:34:05 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/6815fd6a-737d-4ee7-8c0b-345def7c07ee_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Failure isn&#8217;t a moment.<br>It&#8217;s evidence.</p><p>Most organizations treat failure as an event. Something that happens, gets explained, and then assigned. A launch misses. A customer escalates. A system breaks. Someone is closest when it surfaces, and responsibility collapses onto them.</p><p>That&#8217;s why failure feels personal.<br>Context disappears.<br>The system compresses into a name.</p><p>But failures don&#8217;t arrive suddenly. They accumulate quietly. Conditions build over weeks, months, sometimes years, until the system runs out of room to compensate. When the outcome finally appears, it feels abrupt. The path was already set.</p><p>I&#8217;ve watched failures that looked like &#8220;bad weeks&#8221; trace back to decisions made long before anyone felt discomfort:<br>a signal removed to move faster,<br>a dependency assumed but never owned,<br>a metric optimized past its breaking point.</p><p>By the time the issue surfaced, the system had already made the outcome highly likely.</p><p>This is where most organizations go wrong.</p><p>They fix outcomes instead of understanding conditions.<br>They search for causes instead of patterns.<br>They run post-mortems that sound rigorous but leave the system unchanged.</p><p>None of this means accountability disappears.</p><p>There are moments, safety events, security incidents, ethical breaches, where speed matters, where access must be removed, where someone has to be stood down immediately. A systems lens doesn&#8217;t replace decisive action. It clarifies it.</p><p>The difference is whether responsibility is assigned after understanding conditions, or used as a substitute for understanding them.</p><p>Failure becomes dangerous when we don&#8217;t know how to talk about it.</p><p>When the only language available is success or fault, people optimize for hiding signals instead of surfacing them. Noise increases. Trust erodes. The same failures recur under new names.</p><p>One of the most consequential gaps in modern organizations isn&#8217;t intelligence, effort, or technology.</p><p>It&#8217;s the absence of shared language that allows people to observe what&#8217;s actually happening before outcomes harden, and to act on those observations without blame becoming the primary currency.</p><p>We say we want learning, but we punish visibility.<br>We say we want ownership, but separate it from authority.<br>We say we want resilience, but optimize systems until they can no longer bend.</p><p>Failure isn&#8217;t proof that people didn&#8217;t care.<br>It&#8217;s proof that the system relied on compensation it never made visible.</p><p>Some failures are true shocks, events outside a system&#8217;s design horizon. But many of the ones that do the most damage are not. They are signals ignored, pressures normalized, and risks carried quietly by individuals until the system can no longer absorb them.</p><p>What if failure wasn&#8217;t treated as a verdict, but as data?</p><p>What if early signals were captured instead of explained away?</p><p>What if learning didn&#8217;t require someone to absorb the cost personally before the system paid attention?</p><p>This doesn&#8217;t require more dashboards or louder retrospectives. It requires a different posture: observation before judgment, conditions before conclusions, structure before story.</p><p>In the pieces that follow, I&#8217;ll start breaking failures down into observable conditions, early warning signals, and decision points. Not as theory, but as field notes from inside complex systems.</p><p>Less narrative.<br>More structure.</p><p>Because the real question isn&#8217;t whether we fail.</p><p>It&#8217;s whether we know what we&#8217;re looking at when we do.</p>]]></content:encoded></item><item><title><![CDATA[Responsibility Is Infrastructure]]></title><description><![CDATA[A Manifesto for Human-Aligned, Responsibility-First AI]]></description><link>https://writing.justingreenbaum.com/p/responsibility-is-infrastructure</link><guid isPermaLink="false">https://writing.justingreenbaum.com/p/responsibility-is-infrastructure</guid><dc:creator><![CDATA[Justin R. Greenbaum]]></dc:creator><pubDate>Sun, 21 Dec 2025 14:15:43 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/306a08bd-3b47-44ca-836f-5e5ccb2a5b13_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>AI is no longer just a tool.<br>It is becoming <strong>decision infrastructure</strong>.</p><p>As artificial intelligence moves from automating tasks to mediating judgment, the central question is no longer whether systems are accurate, fast, or scalable. The question is whether <strong>responsibility survives scale</strong>.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://writing.justingreenbaum.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Justin's Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Most AI failures do not come from malicious intent or broken models. They emerge when authority becomes diffuse, incentives become misaligned, and systems optimize what is measurable while eroding what is meaningful. Harm accumulates quietly through feedback loops, proxy metrics, and overconfidence until accountability is impossible to reconstruct.</p><p>This manifesto asserts a simple truth:</p><ul><li><p><strong>Prediction is not judgment.</strong></p></li><li><p><strong>Automation is not responsibility.</strong></p></li><li><p><strong>Efficiency is not legitimacy.</strong></p></li></ul><div><hr></div><h2>1. Responsibility must be designed, not assumed</h2><p>Responsibility does not emerge from policies, training, or good intentions. It emerges from defaults, permissions, escalation paths, and failure handling. Systems shape behavior. If responsibility is not embedded into how AI operates, it will dissolve under scale.</p><div><hr></div><h2>2. AI should support judgment, not replace it</h2><p>AI systems should surface uncertainty, highlight risk, and expand human understanding, not quietly substitute for deliberation. In high-consequence domains, judgment must remain human, explicit, and owned.</p><div><hr></div><h2>3. Deployment is a responsibility event</h2><p>The moment an AI system is deployed, it becomes an actor within a social and organizational system. Decisions made during problem framing, data selection, and workflow integration determine who bears risk and who benefits. Deployment without ownership is abdication.</p><div><hr></div><h2>4. Robustness is alignment over time</h2><p>Real-world environments change. Drift is not a nuisance to be retrained away. It is a signal that context has shifted and authority must be reassessed. Systems that slow down under uncertainty are safer than those that remain confidently wrong.</p><div><hr></div><h2>5. Fairness is legitimacy, not a metric</h2><p>Bias is not just statistical imbalance. It is a failure of representation and authority. Fairness metrics can inform, but they cannot decide. When AI-mediated outcomes undermine trust or equity, restraint, not optimization, is the responsible response.</p><div><hr></div><h2>6. Governance must be infrastructure</h2><p>Ethics that rely on vigilance will fail at scale. Responsible AI requires named ownership, traceability, escalation authority, and the ability to stop systems when alignment degrades. Governance works when responsible behavior is routine and irresponsible behavior is difficult.</p><div><hr></div><h2>7. Some decisions should not be automated</h2><p>Not all problems are optimization problems. AI must refuse autonomy where harm is irreversible, contested, or morally non-fungible. Restraint is not a limitation of intelligence. It is a mark of it.</p><div><hr></div><h2>The future of AI is not smarter models</h2><p>It is <strong>better systems</strong>.</p><p>Systems that acknowledge the limits of prediction.<br>Systems that preserve human agency.<br>Systems that learn without erasing accountability.</p><p>AI&#8217;s highest value lies not in replacing judgment, but in <strong>strengthening it</strong>, provided responsibility remains explicit, continuous, and owned.</p><p>This is not AI 1.0 scaled up.<br>This is AI designed to live among humans, deliberately, humbly, and responsibly.</p><div><hr></div><h3>Context</h3><p>These ideas emerged from building and operating large-scale decision systems, observing how AI fails in practice, and formal study in AI deployment and governance. The focus here is not theory, but what happens when systems scale faster than accountability.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://writing.justingreenbaum.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Justin's Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>