<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Enterprise Token Economy]]></title><description><![CDATA[The economics, sovereignty, and governance of enterprise AI agents. Token costs, AI agent memory, and the frameworks CFOs, CDAOs, and FinOps leaders use to govern AI spend. By Tony Wenzel, Co-Founder and CEO of Excipio.]]></description><link>https://tokens.excipio.ai</link><image><url>https://tokens.excipio.ai/img/substack.png</url><title>The Enterprise Token Economy</title><link>https://tokens.excipio.ai</link></image><generator>Substack</generator><lastBuildDate>Sun, 06 Sep 2026 18:00:41 GMT</lastBuildDate><atom:link href="https://tokens.excipio.ai/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Wenze Wenzel]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[theenterprisetokeneconomy@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[theenterprisetokeneconomy@substack.com]]></itunes:email><itunes:name><![CDATA[Tony Wenzel]]></itunes:name></itunes:owner><itunes:author><![CDATA[Tony Wenzel]]></itunes:author><googleplay:owner><![CDATA[theenterprisetokeneconomy@substack.com]]></googleplay:owner><googleplay:email><![CDATA[theenterprisetokeneconomy@substack.com]]></googleplay:email><googleplay:author><![CDATA[Tony Wenzel]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Fifth Layer]]></title><description><![CDATA[Payments, authorization, agent-to-agent, fulfillment. Four races are underway in agentic commerce. The quiet fifth race is memory, and it is the one enterprises cannot afford to lose.]]></description><link>https://tokens.excipio.ai/p/the-fifth-layer</link><guid isPermaLink="false">https://tokens.excipio.ai/p/the-fifth-layer</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Fri, 04 Sep 2026 13:35:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!6lDW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe606e302-7618-4415-b979-5dde25366250_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Gartner&#8217;s analysts don&#8217;t commonly write thrillers, but their top strategic prediction for 2026 drops a bomb: <strong>by 2028, 90 percent of B2B buying will be intermediated by AI agents</strong>.  Those agents will be pushing more than 15 trillion dollars of spend through agent exchanges. Gartner unveiled the number at the IT Symposium last October, and it has been echoing through every commerce deck since. Fifteen trillion dollars is not a vertical. It is roughly half of US GDP rerouted through software negotiating with software.</p><p>And it&#8217;s already started. Sandy Carter, the AI executive and Forbes contributor who has been documenting this shift closely, recently sat down with builders from three of the companies constructing it: David Minarsch of Olas, Nitya Subramanian of Para, and Will Papper of Cloudflare. Minarsch&#8217;s definition of agentic commerce is the clean: &#8220;If one software program purchases a service from another software program.&#8221; Simple, and no longer hypothetical. Carter reports that Olas has recorded more than 14 million agent-to-agent interactions, Para&#8217;s wallets serve more than 15 million end users, and more than 30 companies have joined Mastercard&#8217;s Agent Pay.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Carter frames the buildout in four layers, as four races. Payments, where Visa, Mastercard, and Stripe are competing to move the money. Authorization, where Para is building what amounts to a corporate card with rules attached: a wallet that holds funds and a permission structure that governs what an agent may do with them. Agent-to-agent interaction, where Olas runs the exchange. And fulfillment, where Cloudflare is making the point that purchasing is only half a transaction; something still has to arrive.</p><p>Money moves on the rails. Permission is granted. Agents talk. Goods arrive. It is a complete picture of a transaction, and an incomplete picture of commerce.</p><p><strong>The question the four layers do not answer</strong></p><p>Before any of this scales inside a large enterprise, a fifth question must be asked, and  by people with audit authority: what did the agent know when it bought, and can we prove it?</p><p>Stay with the corporate card metaphor, because it is a good one, and extend it one step. A corporate card with rules attached still comes with a monthly statement. The statement is not a courtesy. It&#8217;s the control mechanism. It&#8217;s how the CFO reconciles, how the auditor samples, how disputes get resolved, how fraud gets found. Nobody hands out cards without statements, no matter how good the rules are.</p><p>Now run that at AI speed. When agents transact thousands of times a day, the monthly statement becomes something else entirely: a continuous audit trail of every request, every decision, and the context behind each one. Which sources did the agent consult? Which policy did it apply? What did it know, and when? Whether the answer it acted on was fresh or stale, grounded or improvised? When two software programs disagree about a purchase, and they will, dispute resolution is a replay of what each side knew at the moment of the transaction. Whoever holds that record holds the leverage.</p><p><strong>The quiet race</strong></p><p>Payments and fulfillment are the visible races because money and logistics always are. The quiet race is who holds the memory and record of all that agent-to-agent traffic.</p><p>There are three candidate answers. The payment rails would love to hold it; transaction data has always been their second business. The platforms and model providers would love to hold it; context is what makes their agents smarter. Or the enterprise holds it, inside its own perimeter, as infrastructure it owns.</p><p>My bet is on the third, and it&#8217;s not because I&#8217;m sentimental. Your agents&#8217; purchase history is your demand curve. Their search history is your strategy. The context they carried into each negotiation is your cost structure and your risk appetite, logged. An enterprise that lets that exhaust accumulate on someone else&#8217;s rails is publishing its playbook one transaction at a time.  It&#8217;s likely training a counterparty&#8217;s model on it. I would not want anyone else training on my customer data, my purchase data, or my search data. Neither will your CFO, your CISO, or your regulator.</p><p>Regulated industries will get there first, because they already live this way. Banks run on the discipline that any model touching money must be inventoried, validated, and explainable to an examiner. Health systems run on the discipline that any system touching patient data must be logged and attestable. Those institutions will not accept &#8220;the agent decided&#8221; as an answer to an examiner&#8217;s question, which means the memory layer is not optional for them. It is the precondition for participating in agentic commerce at all.</p><p><strong>Know your traffic across AI Surfaces</strong></p><p>Here is the uncomfortable baseline: only 18 percent of enterprises can name every AI surface operating in their environment today&#8230; <em>before</em> agentic commerce multiplies the surface count. The estate is already a sprawl of apps, agents, tools, and MCP servers, each making calls, each carrying context, almost none of it recorded anywhere the enterprise controls.</p><p>This is the layer we build at Excipio: a private memory layer that audits and controls the traffic between AI surfaces, inside the enterprise perimeter, on the enterprise&#8217;s infrastructure. Not a payment rail, not an exchange. The statement that comes with the card.</p><p>The four races will produce their winners, and there will be more than one. The fifth layer is different. It should not have an external winner at all, because the record of what your agents knew, asked, and decided belongs to exactly one party.</p><p><strong>If you are building in AI, you should be able to audit and control the traffic between your surfaces.</strong> When 90 percent of B2B buying runs machine to machine, that stops being an architecture preference. It becomes the price of admission.</p><p>Know your traffic.</p><p>--</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6lDW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe606e302-7618-4415-b979-5dde25366250_1200x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6lDW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe606e302-7618-4415-b979-5dde25366250_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!6lDW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe606e302-7618-4415-b979-5dde25366250_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!6lDW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe606e302-7618-4415-b979-5dde25366250_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!6lDW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe606e302-7618-4415-b979-5dde25366250_1200x630.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6lDW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe606e302-7618-4415-b979-5dde25366250_1200x630.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e606e302-7618-4415-b979-5dde25366250_1200x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:55472,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://tokens.excipio.ai/i/214153325?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe606e302-7618-4415-b979-5dde25366250_1200x630.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6lDW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe606e302-7618-4415-b979-5dde25366250_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!6lDW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe606e302-7618-4415-b979-5dde25366250_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!6lDW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe606e302-7618-4415-b979-5dde25366250_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!6lDW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe606e302-7618-4415-b979-5dde25366250_1200x630.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>-</p><p>References</p><p>Gartner, &#8220;Top Strategic Predictions for 2026 and Beyond,&#8221; Gartner IT Symposium/Xpo, October 2025: https://www.gartner.com/en/newsroom/press-releases/2025-10-21-gartner-unveils-top-predictions-for-it-organizations-and-users-in-2026-and-beyond</p><p>Sandy Carter, LinkedIn, September 2026, on her conversation with David Minarsch (Olas), Nitya Subramanian (Para), and Will Papper (Cloudflare): [PASTE POST URL]</p><p>Sandy Carter&#8217;s companion Forbes article, linked in the first comment of her post: [PASTE FORBES URL]</p><p>Suggested Substack tags (tag field, not body): agentic commerce, AI governance, enterprise AI, token economy, audit</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Whoever Holds the Loop]]></title><description><![CDATA[Brett Queener's Horse, Harness, Hay framework, read from the runtime side of the ledger.]]></description><link>https://tokens.excipio.ai/p/whoever-holds-the-loop</link><guid isPermaLink="false">https://tokens.excipio.ai/p/whoever-holds-the-loop</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Thu, 03 Sep 2026 22:05:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!eQNm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eQNm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eQNm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!eQNm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!eQNm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!eQNm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eQNm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:103009,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://tokens.excipio.ai/i/213928110?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eQNm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!eQNm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!eQNm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!eQNm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b2504a0-774d-46e5-bb1c-0657efffaad9_1200x630.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Brett Queener published a piece this week called <a href="https://substack.com/home/post/p-213463050">The Harness, the Horse, or the Hay</a>. It is a thoughtful look at the AI software cycle and an essential framework for making sense of where value settles. If you build, buy, or fund enterprise software, read it before you read this.</p><p>I won't summarize Brett's piece. Rather, I want to consider his ideas from the runtime side of the ledger&#8230; down in the inference traffic where memory and context actually live.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Brett's thesis in brief: because each worker will eventually collapse their workflow into a single interface, most legacy applications will disappear. What survives is the horse (the foundational model), the harness (the single application fitted to a rider and the complete job), and the hay (the operational infrastructure keeping horse and harness running at scale).</p><p>I mostly agree. But there is a crucial dynamic unfolding underneath his taxonomy:</p><p>The defining question of this cycle is not which category you belong to. It is who holds the loop.</p><h2>The Claudeforce Problem</h2><p>Brett defines the harness by four properties: it carries the ontology, it owns the interaction model, it owns the eval and learning loop, and it is singular.</p><p>Then, in his analysis of "Claudeforce," he does something far more interesting than categorize. He runs Salesforce against those four criteria for a sales rep and finds that it fails three of them. Salesforce kept the static records of system state. The corrections, the nuance, and the operational memory of how that seller actually works accumulated inside Claude.</p><p>His takeaway: the layer to watch is whoever ends up holding the loop.</p><p>That is the entire AI cycle in one sentence. Every enterprise AI deployment is answering a silent question: where does the accumulated judgment about how this company works actually live? Right now, by default, it is accumulating in whichever surface the employee used last. That is not an architectural strategy; it is entropy with a login.</p><p>Brett notes that no serious organization will surrender its embedded intelligence back to the frontier model providers, because a platform that serves your direct competitors cannot be the custodian of your DNA.</p><p>I would extend that rule one step further: no serious organization should hand that intelligence to a third-party harness either.</p><p>The harness vendor will argue the ontology belongs to them. For vertical micro-monopolies, that is true. But for the broader enterprise, the ontology is the company. Renting it back at renewal time is the SaaS trap rebuilt one layer higher.</p><p>The own-versus-rent decision is no longer about seats or licenses. It is about memory and context. If you rent the loop, you are a tenant in your own judgment.</p><h2>Compounding Requires Singularity</h2><p>The most compelling part of Brett's harness thesis is compounding: solve a job so thoroughly that the rider hands you everything, and within 60 days they ask you to run the adjacent workflow.</p><p>But compounding carries a precondition: corrections must land in a consistent place.</p><p>In reality, enterprise corrections land in fragmented silos. An account executive corrects a client fact inside a conversational assistant. A paralegal corrects contract logic inside a vertical workflow tool. A financial analyst catches a discrepancy in a sheet the harness never sees.</p><p>Each surface learns a fraction of the business, and none of them share the lesson. The loop is not closed; it is scattered. And a scattered loop does not compound&#8230; it leaks.</p><p>This makes the most critical word in Brett's thesis not harness, but singular. Singularity is what allows a feedback loop to close. If an employee operates out of a single surface, context compounds. Remove that singularity, and compounding halts regardless of ontology quality.</p><p>For the broad swath of the enterprise that will never find a single off-the-shelf vertical harness, that singularity cannot live at the application layer. It must come from the runtime layer shared across every surface.</p><h2>Roadside Traces vs. Inline Control</h2><p>In his evaluation of observability and evals, Brett drops a two-word aside: hold that thought. He later clarifies his rule: sell the instrument that produces judgment, never the judgment itself. Langfuse can show an engineering team where an agent pipeline stalled; it cannot tell them what institutional "good" looks like.</p><p>There is a technical distinction inside that layer that directly impacts value creation: observing a call and controlling a call are entirely different primitives.</p><p>A pure observability platform watches traffic after the fact and serves an execution trace. That is useful telemetry, but it is not a closed loop: a human reads a dashboard, files a ticket, and an engineer updates a prompt or code. Brett's own litmus test for a loop is whether a graded outcome alters future behavior without human intervention.</p><p>You cannot pass that test from the roadside. You can only pass it from inside the traffic.</p><p>If the layer evaluating the call also acts as the inline proxy routing it, the grade and the behavioral adaptation become the same runtime event. That is the dividing line between an instrument that merely records judgment and an instrument that executes it. Both belong in the hay category, but only one closes the loop.</p><h2>What Hay Actually Pools</h2><p>Brett applies a rigorous network-effect test to hay companies: serving 400 customers must make you structurally better than serving four, or you will be commoditized into a native feature. His examples include cross-model pricing, latency curves, failure signatures, and security heuristics.</p><p>Let's be explicit about what this layer should and should not pool:</p><p>What gets pooled: The physics of the traffic. Which models handle specific query shapes best, at what latency and cost. Where failure rates spike. Which routing strategies preserve uptime. No single customer can synthesize that alone, and frontier labs have no incentive to expose it because runtime transparency makes the customer portable.</p><p>What never gets pooled: The content. The semantic ontology. The enterprise's private context.</p><p>That data must remain strictly inside the customer's perimeter. Pooling proprietary content across tenants is the exact corporate self-harm Brett warns against.</p><p>Pool the physics of the traffic; never pool the meaning.</p><h2>The Permanent Buyer</h2><p>Brett frames "harness parts" as a transitional trap with two exits: step down into hay, or scale up into a harness.</p><p>The exit into hay is wider and more permanent than it looks. The self-assembling enterprise is not a transitional buyer waiting for a vendor to arrive. Regulated institutions will assemble their own stacks indefinitely, not because vertical SaaS failed them, but because governance, auditability, and legal liability mandate that they retain exclusive custody of their loop.</p><p>This is a durable market. These buyers do not buy proprietary harnesses; they buy primitives that integrate cleanly with their security posture.</p><p>The strategic question this leaves on the table is the one every CIO and technology leader should be asking:</p><p>Where does our loop live today, and who holds it?</p><p>If the answer is a sprawling list of disparate interfaces, you do not own it yet. If the answer is a single third-party application, you are renting it.</p><p>Horses will get faster, and harnesses will get bought. Through it all, the enterprise that retains custody of its own loop is the only party in the barn whose position compounds every year.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The $8 Billion Middle]]></title><description><![CDATA[Stripe just bought the choke point between AI consumption and the invoice. The claim behind the deal is bigger than the deal, and it is only half right.]]></description><link>https://tokens.excipio.ai/p/the-8-billion-middle</link><guid isPermaLink="false">https://tokens.excipio.ai/p/the-8-billion-middle</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Thu, 20 Aug 2026 18:39:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!InMz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3263df5-9774-46f2-b91c-81751e681b38_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!InMz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3263df5-9774-46f2-b91c-81751e681b38_1200x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!InMz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3263df5-9774-46f2-b91c-81751e681b38_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!InMz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3263df5-9774-46f2-b91c-81751e681b38_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!InMz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3263df5-9774-46f2-b91c-81751e681b38_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!InMz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3263df5-9774-46f2-b91c-81751e681b38_1200x630.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!InMz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3263df5-9774-46f2-b91c-81751e681b38_1200x630.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a3263df5-9774-46f2-b91c-81751e681b38_1200x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:62640,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://tokens.excipio.ai/i/212041583?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3263df5-9774-46f2-b91c-81751e681b38_1200x630.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!InMz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3263df5-9774-46f2-b91c-81751e681b38_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!InMz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3263df5-9774-46f2-b91c-81751e681b38_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!InMz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3263df5-9774-46f2-b91c-81751e681b38_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!InMz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3263df5-9774-46f2-b91c-81751e681b38_1200x630.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Stripe acquired OpenRouter this week for a reported $8 billion. PitchBook broke the analysis on August 19, and the framing from investors was unusually candid. Jeremy Jonker of Infinity Ventures described the acquisition as buying position rather than product, with the next competitive battleground &#8220;sitting at the choke point between AI consumption and the invoice.&#8221;</p><p>The more interesting artifact is Stripe&#8217;s letter to investors, obtained by Axios. In it, Stripe argues that optimizing for developers and optimizing for coding agents are largely the same discipline, and that building economic infrastructure for the internet and building economic infrastructure for AI are mostly the same job.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>That claim deserves interrogation, because a lot of enterprise AI strategy is about to be built on top of it. It is half right. The half that is wrong is the half that will cost you.</p><h2>Where Stripe is right: the middle is the business</h2><p>Stripe never issued the cards or held the deposits. It sat between businesses and the banks that did, taking a small percentage of every transaction that passed through. OpenRouter runs the identical trade on inference: it sits between developers and model providers, connects the two on price, speed, and availability, and takes a fraction of the spend it facilitates.</p><p>The deal prices a position, and the position is the middle. Not the models. Not the applications. The thin layer that every request passes through and where every dollar gets metered. Ramp launching its own model router in the same news cycle confirms the pattern: fintechs have concluded that the layer between AI consumption and payment is where the next decade of margin lives.</p><p>On this point, the letter is correct, and the enterprise implication is direct. The layer between every AI surface, apps, agents, tools, MCP servers, and the models they call is where cost, control, and accountability concentrate. An $8 billion price tag on a two-year-old company is the market agreeing loudly.</p><h2>Where the analogy breaks: payments forget by design</h2><p>Here is what the letter skips. Payments infrastructure moves money, and money is fungible and stateless on purpose. A dollar does not change when it moves. The network is not supposed to remember your dollar, learn from your dollar, or serve your dollar back to you later. Settlement is designed to complete and clear. Forgetting is a feature.</p><p>AI infrastructure moves something categorically different: context and answers. That cargo is non-fungible, it compounds in value when retained, and the defining defect of today&#8217;s stack is that it forgets all of it. A transformer keeps no memory between sessions. Every agent session starts from zero and re-derives what the organization already worked out. The bill arrives again anyway. In this newsletter&#8217;s vocabulary, that is the Rediscovery Tax: the cost an organization pays every time an AI system re-derives an answer it already has.</p><p>So the analogy inverts at the exact point where it matters:</p><ul><li><p>Payments infrastructure succeeds by forgetting. Statelessness is what makes settlement clean.</p></li><li><p>AI infrastructure fails by forgetting. Statelessness is what generates the repeat spend the meter is billing you for.</p></li></ul><p>Building economic infrastructure for AI is not mostly the same job as building it for the internet. The internet&#8217;s economic layer had to move value without holding it. AI&#8217;s economic layer has to decide what is worth holding, because retained knowledge is the only thing that bends the cost curve.</p><h2>A meter has no incentive to shrink what it meters</h2><p>Follow the business model. A router in the middle earns a fraction of the inference spend it facilitates. Its revenue grows with token volume, and token volume includes waste. Stripe&#8217;s take rate on payments was aligned with its customers, because merchants want more transactions. A take rate on inference is aligned with consumption itself, and consumption is exactly the number an enterprise should be interrogating.</p><p>Token Yield is the metric that exposes the gap: the ratio of necessary spend to total spend. Every AI bill divides into unavoidable spend on genuinely novel work and recoverable waste on things the organization already knows. Routing operates entirely inside the second category without touching it. It finds you a cheaper price for re-deriving an answer you already own.</p><p>The cheapest model for a known answer is no model. Routing optimizes the price of rediscovery; memory eliminates it.</p><p>The market data says the recoverable half is not a rounding error: 69% of input tokens are repeated context, yet only 28% of calls use any caching (Datadog State of AI Engineering 2026). That is a market prior, not a product claim, and it describes the traffic a metering layer happily bills at full freight.</p><h2>Three questions to ask of any middle layer</h2><p>If the middle of the AI stack is now an $8 billion position, enterprises should be underwriting that layer the way they underwrite any critical infrastructure. Three questions do most of the work:</p><ol><li><p><strong>What does it remember, and who owns the record?</strong> If validated answers, corrections, and context accumulate in a vendor&#8217;s infrastructure, you are building institutional memory inside someone else&#8217;s walls. If they accumulate inside your own perimeter and survive a model switch, you own an appreciating asset. Rent buys the answer once. Ownership means never buying it twice.</p></li><li><p><strong>Does its revenue rise or fall with your waste?</strong> A layer paid on volume monetizes your Rediscovery Tax. A layer that resolves known answers locally, in milliseconds instead of a round trip measured in seconds, is structurally on the other side of that trade.</p></li><li><p><strong>What can it prove?</strong> The middle sees every request from every identity in the request path. That vantage point is either an audit and governance asset with lineage you control, or it is telemetry accruing to someone else. There is no neutral option.</p></li></ol><h2>Where the layer goes next</h2><p>Stripe&#8217;s acquisition settles the question of where value sits in the AI stack. It does not settle what the layer should do. Metering consumption is the first, most obvious business to build there, which is why a payments company got there first.</p><p>The harder questions arrive with the invoice. What did all of those calls actually need to happen? Which answers did the organization already own? Who can reconstruct why an agent decided what it decided, and from which source? Those are questions of memory, governance, audit, and control, and they are answered by architecture, not by contract.</p><p>The middle layer of AI is going to be one of the defining infrastructure positions of this decade. Stripe just paid $8 billion to hold the version of it that meters the spend. The version worth owning is the one that remembers why the spend happened, proves what it produced, and quietly makes the meter run slower.</p><p>The way out is not fewer tokens. It is fewer wasted ones.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Nobody Is Pricing the Arrows]]></title><description><![CDATA[Graph engineering is going viral. The token bill it creates is not in the diagram.]]></description><link>https://tokens.excipio.ai/p/nobody-is-pricing-the-arrows</link><guid isPermaLink="false">https://tokens.excipio.ai/p/nobody-is-pricing-the-arrows</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Wed, 12 Aug 2026 23:50:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!X0MF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Graph engineering is the term of the month. You have seen the diagrams on your feed: a question enters at the top, a planner splits it into five parallel researchers, a skeptic attacks the findings, a merger writes the recommendation, a checker grades it, a human approves it. The diamond shape is everywhere, and the pattern is genuinely good. Work designed as a graph beats work crammed into one chat window.</p><p>None of the diagrams show the part that lands on your invoice. Every arrow has a price tag.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The chat version of &#8220;should I build this startup&#8221; is one model call. The graph version of the same question is a planner call, five researcher calls, a skeptic call, a merger call, and a checker call. That is nine calls before anything loops. Add one revision cycle, which every real workflow has, and you are at twelve to fifteen. The diagram that went viral because it produces better answers is also a diagram of your token bill multiplying by an order of magnitude.</p><p>Graphs are worth it. I would build them anyway, and I tell design partners the same. What bothers me is that the industry is teaching everyone the workflow pattern and nobody the unit economics.</p><h2>The redundancy hiding inside the diamond</h2><p>Look closely at what those nine calls actually contain. The planner reads the company context to decompose the question. Each of the five researchers reads overlapping slices of the same context to do their work. The skeptic re-reads what the researchers produced plus the context needed to attack it. The merger re-reads everything. The checker re-reads the merge.</p><p>The same enterprise knowledge, retrieved and paid for five, six, seven times inside a single run. Now multiply by the runs. A support triage graph does not run once. It runs on every ticket. A content graph runs on every piece. A code review graph runs on every pull request. Bounded, high repetition workloads are exactly where graphs get deployed first, because that is where the quality gain justifies the build.</p><p>Which means the pattern the market is adopting fastest is the pattern that inflates the rediscovery tax fastest. Your agents keep re-deriving what your organization already knows, and in a graph, they re-derive it at every node.</p><h2>Two graphs, and only one of them is trending</h2><p>The video essays make a distinction worth keeping. There are agent graphs, which govern how work moves: the boxes and arrows, the planners and checkers. And there are knowledge graphs, which govern what the work knows: the relationships in your data that reasoning runs across.</p><p>The content wave is almost entirely about agent graphs. That is understandable. Agent graphs are visible, drawable, and you can build one this afternoon in LangGraph or n8n. The knowledge layer underneath is invisible in the diagram, so it is invisible in the discourse.</p><p>But follow the arrows. Every node in an agent graph resolves its knowledge from somewhere. Today, for most teams, that somewhere is a frontier model API, at full price, at roughly 2,500 milliseconds per live call, with your context leaving the building every time. You are renting your own institutional knowledge back from a model provider, once per node, per run, forever.</p><p>That is the rent position. The own position looks different: repeated semantic queries resolve from a customer owned, compounding knowledge graph inside your perimeter, at sub-10ms on a hit, with zero bytes leaving the perimeter on a cache hit. The agent graph stays exactly as designed. The substrate underneath it changes, and the modeled result is roughly a 42% blended token cost reduction across the workload.</p><p>The agent graph decides how work moves. The knowledge graph decides what the work knows, and whether you rent it or own it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!X0MF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!X0MF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!X0MF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!X0MF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!X0MF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!X0MF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:68527,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://tokens.excipio.ai/i/210970180?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!X0MF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!X0MF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!X0MF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!X0MF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86d8825-de71-455c-a2d9-3ddb8080ad98_1200x630.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The question nobody has measured</h2><p>If you have agent graphs in production, or on the roadmap, there is one number worth knowing before the CFO asks for it: across the nodes in your graph, what percentage of retrieved context is the same enterprise knowledge, retrieved repeatedly?</p><p>Almost nobody measures this, because the diagrams do not have a column for it. But the arithmetic is not subtle. Parallel decomposition splits one retrieval into five overlapping ones.</p><p>The overlap is the tax.</p><p>The advice circulating in the graph engineering wave is correct as far as it goes: build the smallest graph that raises quality, and put the human gate where mistakes get expensive. I would add one line to it. Put the memory layer where the retrievals repeat, because that is where the money leaks.</p><p>The diagrams price nothing. Your invoice prices everything. Close that gap while it is still cheap to close.</p><p>The Enterprise Token Economy tracks the unit economics of enterprise AI. If your agent workloads are repetitive and your token bill is not shrinking, that gap is the subject of this newsletter.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Meter Is Not the Fix]]></title><description><![CDATA[KPMG says half of enterprises pulled back AI agents over cost. The prescription everyone is writing stops one step short.]]></description><link>https://tokens.excipio.ai/p/the-meter-is-not-the-fix</link><guid isPermaLink="false">https://tokens.excipio.ai/p/the-meter-is-not-the-fix</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Mon, 10 Aug 2026 12:29:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!YFfw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A number went around the internet last week: 49 percent of executives have scaled back AI agent deployments because operating costs outweighed the benefits. It came from KPMG's Global AI Pulse for Q2 2026, a survey of 2,145 senior leaders across 20 countries, and within a day Polymarket was quoting odds on an AI bubble burst.</p><p>Sandy Carter read the full report so the rest of us didn't have to, and her Forbes breakdown is the one to read. Her line is the correct frame: "That is not a bubble bursting. That is a bill arriving."</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>She is right. Look at what else sits in the same dataset. AI is a top investment priority for 79 percent of leaders, up from 74 percent the quarter before. Average AI budgets are holding at $188 million. The share of organizations calling AI part of everyday work jumped from 13 percent to 22 percent in a single quarter, the biggest move KPMG has ever measured at any point on its maturity curve. Nobody is leaving. They are rephasing.</p><p>We have seen this movie. A decade ago, cloud invoices arrived that nobody could explain, and the industry answered with FinOps. Owners, forecasts, unit costs, a discipline. The token era is getting its own version, and KPMG's data shows it forming in real time. 53 percent of organizations report AI cost dashboards. 54 percent have embedded cost review in their AI approval loops. The companies with dashboards are roughly five times more likely to report real ROI.</p><p>Here is where I part ways with the standard prescription.</p><h2>Phase one and phase two</h2><p>Cloud FinOps matured in two distinct phases. Phase one was visibility. See the spend, tag the resources, show back the costs. Phase two was where the money actually got saved: rightsizing, reserved instances, deleting the zombie infrastructure that phase one exposed. Visibility never saved a dollar by itself. It told you where the dollars were dying.</p><p>Every recommendation now circulating for AI cost, including the four in the KPMG coverage, is a phase one move. Install a meter before you scale. Make token economics a leadership literacy. Put cost review inside the approval loop. Rephase rather than retreat. All correct. All observability.</p><p>Nobody is talking about phase two, because phase two requires knowing what the meter will reveal once you can finally read it. And what it reveals is uncomfortable.</p><h2>What the meter shows</h2><p>Agents are priced like electricity but they do not behave like appliances. They run long tasks, call other tools, and check their own work, and every step is on the meter. When GitHub Copilot moved to usage-based billing on June 1, a Visual Studio Magazine writer tracked his first day and projected a $180 monthly bill on a plan that had been flat $10. One long, tool-heavy session. Eighteen times the cost.</p><p>Now ask what those metered steps actually were. Some fraction was genuinely new work, novel reasoning on a novel problem. Look inside any agent deployment, though, and a large share is repeat purchases. The agent re-derives a policy it derived yesterday. It re-verifies a fact the organization verified a thousand times. It re-retrieves and re-summarizes the same document for the fortieth user this month. Each of those steps bills full freight and produces zero new value.</p><p>This is the Rediscovery Tax, and it is invisible on a cost dashboard. The dashboard tells you the bill is high. It does not tell you that you bought the same answer 400 times.</p><p>That is why only 26 percent of billion-dollar companies reporting full real-time cost visibility is not even the scary number. The scary number is how few of the 26 percent can decompose their spend into first purchases versus repeat purchases. Token Yield, the ratio of necessary spend to total spend, is the metric phase two runs on. Almost nobody can compute it yet.</p><h2>What phase two looks like</h2><p>Phase two of cloud FinOps was rightsizing and reservation. Stop paying on-demand prices for predictable workloads. Phase two of token FinOps is interception. Stop paying inference prices for answers you already own.</p><p>The mechanics differ from cloud because the waste differs. Cloud waste was idle capacity, machines running with nobody using them. Token waste is rediscovery, intelligence re-manufacturing its own prior output. You fix idle capacity by turning things off. You fix rediscovery by remembering, which means a memory layer that sits in front of the model, recognizes when a question has already been answered, and resolves it from knowledge the enterprise owns instead of renting the answer again.</p><p>This is also where cost discipline and governance stop being separate conversations. An answer resolved from your own knowledge graph never left your perimeter, never touched a vendor's meter, and never depended on a model's mood that day. Rent the commodity. Own the differentiation.</p><h2>The bill is the beginning</h2><p>KPMG's data does not describe a retreat. It describes a market that just learned to read its first invoice, and cloud already taught us the discipline takes two phases to complete.</p><p>The 49 percent who pulled back agents are the realists in this dataset, clearing room to scale what pencils out. When they come back, and the 79 percent priority number says they will, they will arrive with a phase two question. Not "what did the agents cost," but "how much of that did we need to spend at all."</p><p>That question has an answer, and it is a lot smaller than the current bill.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YFfw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YFfw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!YFfw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!YFfw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!YFfw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YFfw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:61879,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://tokens.excipio.ai/i/210593751?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!YFfw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!YFfw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!YFfw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!YFfw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a00d816-edf5-4fb1-a31a-9a6e6a75355c_1200x630.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The 49/79/26 figures are from KPMG's Global AI Pulse Q2 2026, surfaced in Sandy Carter's Forbes analysis, which includes a plain-English token pricing explainer your CFO will thank you for.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Only Library That Survived]]></title><description><![CDATA[Vesuvius sealed a private archive in 79 AD. AI read it in 2024. What the only surviving library of antiquity teaches an enterprise about the Rediscovery Tax and where intelligence compounds.]]></description><link>https://tokens.excipio.ai/p/the-only-library-that-survived</link><guid isPermaLink="false">https://tokens.excipio.ai/p/the-only-library-that-survived</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Fri, 07 Aug 2026 03:09:08 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a2101cf5-a10e-4375-8839-3e45210f9b7d_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In August of 79 AD, Mt. Vesuvius erupted and buried a private library in the seaside town of Herculaneum. If you haven&#8217;t visited, you should.  It&#8217;s amazing.  The villa of which we speak, likely belonged to Lucius Calpurnius Piso, Julius Caesar's father-in-law. Its library containt a collection of roughly eighteen hundred scrolls.  All were carbonized in place: turned to charcoal, sealed under volcanic rock, unreadable but intact.</p><p>What&#8217;s interesting is, that of all the great libraries of the classical world, this is the only one that survived as physical objects. Alexandria, the legendary central knowledge infrastructure of antiquity, disintegrated and dispersed across centuries. Not one scroll from the library of Alexandria survives. Pergamon, gone. The public libraries of Rome, gone. By a quirk of fate, one library we can still hold in our hands is the private one.  It&#8217;s archive born of single household, sealed by catastrophe.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>In April 2024, researchers from the University of Pisa and Italy's National Research Council announced that modern imaging technology and AI had recovered readable text from one of the carbonized scrolls from Herculenium: new sections of the History of the Academy by Philodemus, including an account of Plato's final hours and a revised date for his enslavement.</p><p>Like magic, the retrieval worked. It only took 1,945 years.</p><p>The enterprise alegory of this story is operating on a compressed timeline, inside your own systems, right now.</p><h2>Retention is not memory</h2><p>The villa held the goods, but nothing could be accessed. Every answer its library held was preserved perfectly, yet unavailable completely. For nineteen centuries the  value of the archive was zero.  It was not because the knowledge was lost but rather because the retrieval layer did&#8217;t yet exist.</p><p>Most enterprises are running a similar architecture. The work gets done. The answers get validated. You see glimpses of it in a closed ticket, a buried thread, or a slide from a project whose owner left in 2023. Then it becomes Groundhog Day on a loop.  An AI agent gets asked a question it answered a few minutes ago.  But nothing an agent concludes persists anywhere.  So the the next agent has to ask the question again and the organization has to pay a frontier model to derive it again.  No memory.  No synergy.  </p><p>We call it Rediscovery Tax: the cost an organization pays every time an AI system re-derives an answer the organization already has. The work&#8217;s been done. The answer been validated. Unfortunately, the bill arrives again anyway.</p><p>It sounds trite to say it, but Herculaneum is the Rediscovery Tax at geological scale. The answers were there the whole time.  Yet, billing cycle for not being able to reach them ran for nineteen centuries.</p><h2>The sovereignty lesson is sharper</h2><p>Alexandria was the cloud repo of the ancient world: centralized, shared, magnificent, and owned by whoever held the city. When Alexandria fell, everyone's knowledge fell with it. The villa was a sigle perimeter. One owner, one roof, one collection that never left the premises. When catastrophe destroyed the building, it could no longer disperse the asset.</p><p>The sovereignty question is not a matter of where your data lives. It is really a matter of where your intelligence compounds. Satya Nadella put it plainly at Davos in January 2026: "If you are not able to embed the passive knowledge of a firm in a set of weights in a model that you control, by definition you have no sovereignty." An enterprise that streams its validated answers out to rented intelligence is building Alexandria.  It&#8217;s impressive,and communal but structurally incapable of surviving as an owned asset. Firms that hold their institutional intelligence a decade from now are building the villa: a customer-owned, compounding knowledge graph inside their own perimeter. They keep intelligence inside. It compounds permanently.</p><h2>The punchline in the ash</h2><p>In a metacognative irony of Hitchiker&#8217;s Guide proportion, the scroll the researchers recovered wasn&#8217;t a treatise on metaphysics. It was the History of the Academy: the institutional record of the West's first knowledge organization. It was an archive about how an institution created, validated, and transmitted knowledge, preserved by the one library architecture that kept its assets inside its own walls.  </p><p>The Academy understood the knowledge problem. Plato built an institution precisely so that validated knowledge would outlive the people who validated it. It&#8217;s form changed, but it lasted nearly nine centuries, longer than any company on earth has existed.</p><h2>What to do with this</h2><p>Don&#8217;t wait for the imaging team. The lesson of Herculaneum isn&#8217;t that a sufficiently advanced AI will eventually excavate your SharePoint. Though it&#8217;s cool to think about that.  What matters is that retention without a retrieval layer is indistinguishable from complete loss for as long as the gap lasts, and the gap can last longer than the institution.</p><p>Three questions for your own villa:</p><ol><li><p>When an agent answers a question your firm has already answered, does anything capture that answer so the next agent can reach it, or does the billing repeat?</p></li><li><p>If your model provider relationship ended tomorrow, what validated intelligence would you still hold, in an asset you own, inside your perimeter?</p></li><li><p>Is your knowledge compounding where it matters to you, or where your API traffic goes?</p></li></ol><p>The cheapest model for a known answer is no model at all. Transformer models were inconceivable in Herculenium.  But this little private library teaches the whole course like a fable: what you own survives, what you can retrieve compounds, and the only library that lasted was the one that never let its scrolls leave the building.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The 80 Percent Price Cut Is a Margin Call on the Intangible Economy (2026)]]></title><description><![CDATA[Weights are non-rival. Serving is rival. And 92 percent of the S&P 500 just became a sorting question.]]></description><link>https://tokens.excipio.ai/p/the-80-percent-price-cut-is-a-margin</link><guid isPermaLink="false">https://tokens.excipio.ai/p/the-80-percent-price-cut-is-a-margin</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Mon, 03 Aug 2026 17:23:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!fDcP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>On July 30, OpenAI cut the price of GPT-5.6 Luna by 80 percent, three weeks after launching it. Input tokens dropped from $1 to $0.20 per million. Terra fell 20 percent. Sol held its price and gained a Fast mode that charges double for 2.5 times the speed.</p><p>Most of the commentary treats this as a pricing story. It is not. It is the first margin call on the intangible economy, and the collateral being marked down is every business model built on the assumption that intelligence would stay scarce.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>What happened, minus the drama</h2><p>A price cut this fast, this deep, on a model this new is not generosity. It is what pricing power looks like when it evaporates.</p><p>The pressure is documented. CNBC reported in July that Chinese models had captured 46 percent of US enterprise token usage on OpenRouter. DeepSeek V4 Pro was serving tokens at a fraction of Luna&#8217;s old price. Open weights had become a credible substitute, and once a credible substitute exists, the rent collapses.</p><p>That word, rent, is doing precise work. Understanding why requires one distinction that most of the coverage missed.</p><h2>Weights are non-rival. Serving is rival.</h2><p>Intelligence, once trained into a model, is a non-rival good. Copying the weights costs nothing. My use of them does not diminish yours. This is the same property that economics textbooks assign to lighthouse beams and national defense, and it has a known consequence: non-rival goods cannot sustain a price above the cost of delivering them once a substitute exists.</p><p>Serving is different. Running a model consumes GPU seconds, power, and memory bandwidth, all of which are consumed exclusively. Serving is rival, and it has a real cost floor.</p><p>So what deflated on July 30 was not the cost of intelligence. It was the rent being charged on top of a non-rival good. The 80 percent cut is not a strategy. It is arithmetic catching up with structure. You cannot hold a rent on something anyone can copy once someone with a cluster decides to serve the copy at cost.</p><p>This makes the deflation predictable rather than shocking. Nobody needed to see the future. They needed to read the asset class.</p><h2>The rhyme: dark fiber and free browsers</h2><p>We have run this experiment before.</p><p>In the late 1990s, telecom carriers overbuilt fiber on the assumption that bandwidth scarcity would persist. It did not. Global Crossing and WorldCom collapsed. The fiber went dark for a decade. And the companies that won the internet were the ones that designed as if bandwidth were free before it was. Streaming video looked insane in 2002 and inevitable in 2012. The only thing that changed in between was who ate the deflation.</p><p>The modern dark fiber is depreciated accelerators. As frontier training moves to next-generation silicon, the prior generation does not disappear. It becomes the commodity inference substrate, serving open weights at something close to cost. That glut is already forming, and the OpenRouter numbers are its first visible symptom.</p><p>The browser tells the second half of the story. Netscape tried to monetize the artifact. Microsoft gave it away. The browser&#8217;s value did not evaporate. It relocated, from the thing itself to the position it occupied. Value moved to what ran inside the browser and to whoever controlled the default.</p><p>Transport commoditized. The layers above it captured the value. That is the rhyme.</p><h2>Where the rhyme breaks</h2><p>Analogies earn trust by admitting their limits. This one has two.</p><p>First, bandwidth had no permanent frontier tier. A gigabit in 2010 was a gigabit. Model capability keeps moving, so there is always an expensive tier that has not commoditized yet. Sol held its price while Luna fell 80 percent, and that split may be structural rather than transitional. The frontier stays priced. Everything a year behind it deflates to the marginal cost of serving.</p><p>Second, the telcos did not own the layer above them. Global Crossing had no cloud business. The AI labs and hyperscalers are vertically integrated across compute, weights, and application. Value migrating up the stack does not automatically mean value migrating away from incumbents this time. They are waiting at the top of the stack too.</p><p>Both caveats sharpen the question rather than dulling it. If the commodity layer deflates and the incumbents contest the layers above, the only durable position is one they cannot copy or subpoena. Which brings us to the number that reframes the whole event.</p><h2>The 92 percent problem</h2><p>Ocean Tomo&#8217;s 2025 Intangible Asset Market Value Study, released in February 2026, puts intangible assets at roughly 92 percent of S&amp;P 500 market capitalization, up from 17 percent in 1975. Fifty years of American enterprise migrating its value out of things you can touch and into things you can think.</p><p>Then along comes a technology whose entire function is to copy, compress, and regenerate thought at near zero marginal cost.</p><p>That is why July 30 matters beyond model pricing. It is a preview of what happens to any intangible asset that turns out to be non-rival and non-excludable. Ninety-two cents of every S&amp;P 500 dollar is now exposed to a sorting question, and the sorting runs on one axis: excludability. Sort the intangible dollar into three buckets.</p><ol><li><p><strong>Codified knowledge.</strong> Anything that can be written down: documentation, methodologies, playbooks, generic software patterns. Non-rival by nature and increasingly non-excludable in practice. This is the bucket the models ate first, and the Luna cut just repriced it toward zero.</p></li><li><p><strong>Legally excludable IP.</strong> Patents, trademarks, copyrights, trade secrets. Here rivalry is manufactured by law. The asset holds value only as long as the legal regime holds.</p></li><li><p><strong>Structurally excludable assets.</strong> Proprietary data, accumulated and validated context, distribution, regulatory position, switching costs. Excludable by architecture rather than statute. Nobody can copy their way into these, and nobody can litigate their way in either.</p></li></ol><p>Bucket one just got marked to market. Bucket three is where value is fleeing. Bucket two is the interesting one, because its floor is about to be tested.</p><h2>The other shoe: IP, know-how, and licensing</h2><p>The legal regime that makes bucket two excludable was written for an economy where copying was hard and generation was human. Every pillar of it is under live stress at once.</p><p>Training inputs sit in unresolved copyright litigation, with fair use doctrine stretched over a use case it was never shaped for. Model outputs receive no copyright in most jurisdictions, which produces a perverse result worth sitting with: the more your product is model-generated, the thinner your own IP claim on it becomes. The data licensing market is being priced in real time, with publishers and platforms negotiating rents on assets whose enforceability the courts have not yet confirmed.</p><p>And then there is know-how. Trade secret protection depends on information deriving value from secrecy and receiving reasonable efforts to keep it secret. General counsels are beginning to ask an uncomfortable question: what happens to that claim when the substance of your know-how is transmitted to a third-party inference endpoint on every API call? I am not offering a legal conclusion. I am observing that the question is now on the table at every regulated enterprise, and that it is an architecture question wearing a legal costume.</p><p>The common thread: legal excludability is itself a kind of rent, and it can be repriced by a ruling as abruptly as Luna was repriced by a press release. Structural excludability cannot. Its defensibility never depended on a judge.</p><h2>Where B2C and B2B actually split</h2><p>The obvious prediction is that consumer AI goes cheap and DIY while enterprise stays frontier. The evidence runs the other way. The largest consumer deployments on earth run open weights, because billions of users against thin revenue per user cannot pay frontier rates. Meanwhile an enterprise vendor charging $500 per seat can burn a few dollars of tokens per seat and never notice.</p><p>So the split is not B2C versus B2B. Two different variables are doing the sorting.</p><p>Consumer sorts on unit economics: revenue per inference call against the cost of that call. The Luna cut resolves most of that question by making frontier access cheap enough that DIY loses its main argument at consumer scale.</p><p>Enterprise sorts on sovereignty. Regulated buyers carry contractual, statutory, and privilege-based constraints on where their information may travel, and those constraints do not read pricing tables. Satya Nadella named this at Davos in January 2026: a firm that cannot embed its knowledge in an asset it controls has no sovereignty and is leaking enterprise value to a model company on every query. An 80 percent price cut changes the size of the leak&#8217;s invoice. It does not close the leak.</p><p>There is a second reason the cut will not shrink enterprise AI budgets: volume. Datadog&#8217;s 2026 State of AI Engineering report, drawn from thousands of organizations, found that 69 percent of enterprise LLM input tokens are system prompts and repeated context. Agentic workloads multiply calls faster than prices fall, and the repeated context rides along on every one of them. Cheaper tokens do not reduce the habit of re-buying answers you already own. They subsidize it. Jevons had this figured out in 1865 with coal.</p><h2>How to sort your own intangibles</h2><p>If you run a SaaS company, an AI product, or a P&amp;L that depends on either, the exercise is uncomfortable but short. Take your differentiation claims and ask three questions of each.</p><p>First, can it be copied? If a competitor with API access and a quarter can replicate it, it is bucket one. Stop calling it a moat in board meetings before someone does it for you.</p><p>Second, does its defensibility depend on a legal outcome you do not control? Then it is bucket two, and it deserves a discount rate that reflects doctrinal uncertainty, not the discount rate you were using in 2023.</p><p>Third, does it compound in an asset you own? Validated answers, corrected outputs, accumulated context, customer-specific knowledge that gets richer with use. That is bucket three. It is the only bucket where AI spend builds equity instead of paying rent, and it is the only story worth telling an investor in 2026.</p><p>I spent part of my career pricing intangible value in public markets, running a fund that traded US large-cap equities on proprietary brand signals. The lesson from that work holds here. The market has never doubted that intangibles carry the value. The question, every time, is which intangibles stay defensible when the copying gets easier. On July 30, the copying got 80 percent cheaper.</p><p>Rent the commodity. Own the differentiation. The price of the commodity just told you which is which</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fDcP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fDcP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!fDcP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!fDcP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!fDcP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fDcP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:69239,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://tokens.excipio.ai/i/209665802?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fDcP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!fDcP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!fDcP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!fDcP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61370d2-c036-4ffa-99c2-32a6ce48f9b3_1200x630.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Two-Minute Test: The AI Governance Metric Nobody Is Measuring in 2026]]></title><description><![CDATA[The gap between your sanctioned path and the forbidden one predicts everything your policy cannot.]]></description><link>https://tokens.excipio.ai/p/the-two-minute-test-the-ai-governance</link><guid isPermaLink="false">https://tokens.excipio.ai/p/the-two-minute-test-the-ai-governance</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Fri, 24 Jul 2026 18:10:41 GMT</pubDate><content:encoded><![CDATA[<p>Nearly every enterprise has an AI usage policy. Few of them can tell you the one number that predicts whether the policy works: how much longer the approved path takes than the forbidden one.</p><p>This week Nate B. Jones put a name to that gap. Writing about the gauntlet squeezing employees, a manager demanding AI-driven output on one side and an IT policy banning uploads on the other, he proposed a simple diagnostic: put your sanctioned AI workflow on a clock and race it against the consumer route. He calls it the two-minute test, and <a href="https://natesnewsletter.substack.com/p/use-ai-sensitive-files">his full piece is worth reading</a>. No plan survives contact with the enemy.  In most contexts that enemy is time.  If the compliant path loses by two minutes, you already know what people will do the night before the deadline.  </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Nate&#8217;s right a lot and his idea deserves to be pushed further than a diagnostic. The two-minute test is not just an intrusive check on employee behavior. It is a measurement of pragmatic, and hence useful, your governance architecture really is.  I posit that it belongs on the same dashboard as your token spend.</p><h2>The double bind is a design failure, not a discipline failure</h2><p>The employee caught between the manager and the IT policy is not confused about priorities.  They are responding to the incentives the organization built:</p><ul><li><p>The manager controls promotion, and the manager wants AI-scale output</p></li><li><p>The policy names what must not happen, but not how the work should happen</p></li><li><p>The consequence of a privacy breach is abstract and future; the consequence of a missed deadline is concrete and Monday</p></li></ul><p>So, practically speaking, the privacy decision gets made file by file, by the person with the least authority to make it and the least context for what a regulator will think of it later. The company has a privacy process after all. It just depends entirely on whoever is holding the file.  </p><h2>What the two-minute test actually measures</h2><p>Timed statistically, the test collapses three governance questions into one number:</p><ol><li><p>Friction: how many approvals, portals, and redactions stand between an employee and a sanctioned answer</p></li><li><p>Capability gap: how much intelligence the safe tool sacrifices relative to the frontier consumer tool</p></li><li><p>Default direction: which path a reasonable person takes when nobody is watching</p></li></ol><p>The third one is the whole game. Compliance programs assume the default is the policy. In practice the default is the fastest path that produces acceptable work. Policy can move intentions. Only architecture moves defaults.</p><h2>Governance by contract vs. governance by architecture</h2><p>There are two approaches to this problem.  Both employ different mechanisms with different failure modes.</p><p>Governance by contract relies on rules, vendor agreements, training decks, and periodic review. Its failure mode is human: it works exactly as well as the person doing the remembering, on the day they are busiest. It also cannot undo its core exposure. A vendor agreement governs what a provider promises to do with your data, not the fact that it was already sent. The transmission is the problem, not the contract. A February 2026 ruling from Judge Rakoff in the Southern District of New York made the stakes concrete for legal teams: disclosing privileged work product to a third-party AI tool can destroy attorney-client privilege. No DPA claws that back.</p><p>Governance by architecture removes the decision instead of policing it. If sensitive context never has to leave the perimeter to get an answer, there is no file-by-file judgment call to get wrong, no deadline-night exception, no training deck to forget. The safe path and the fast path become the same path, so the default does the compliance work.</p><p>The two-minute test is how you find out which one you actually have, as opposed to which one is in the binder.</p><h2>How to run the test in your organization</h2><ol><li><p>Pick three real tasks from three real roles: an analyst summarizing a client document, a marketer drafting from a strategy memo, an engineer debugging against internal code</p></li><li><p>Time the fully sanctioned path, including logins, approvals, redaction steps, and any quality shortfall that forces rework</p></li><li><p>Time the consumer route the policy forbids, honestly, the way a stressed employee would actually do it</p></li><li><p>Subtract. That number is your shadow AI forecast</p></li><li><p>Re-run quarterly. The gap widens every time a frontier lab ships and your approved stack does not</p></li></ol><p>If the sanctioned path wins or ties, your policy will hold. If it loses by minutes, no amount of training closes the gap, because you are asking people to donate their scarcest resource to a risk they cannot see.</p><h2>Where the token economy meets the test</h2><p>There is a version of this problem hiding inside your inference bill too. A large share of enterprise AI traffic is not novel work. It is the same questions, re-asked and re-answered at full price, with institutional knowledge leaving the perimeter on every round trip. You pay a Rediscovery Tax in tokens and a Disclosure Tax in sovereignty, and both compound quietly.  Then they scale by the number of agents at work.  Ken Huang postulates that "coordination, not intelligence, is becoming the hidden tax on AI adoption&#8221;.  Ken calls it the Coordination Tax.</p><p>This is the problem we built Excipio to remove, so read the next paragraph with that disclosure in mind. A semantic memory layer that sits between agents and models changes the physics of the test: a repeat question resolves in single-digit milliseconds instead of the roughly 2,500ms a live model call takes, zero bytes leave the perimeter on a cache hit, and our modeling shows a 42% reduction in LLM API spend on Day 1 (modeled, not yet production-measured, and we say so every time). The governance point stands independent of any vendor: when the compliant path is also the fastest path, the two-minute test inverts, and shadow AI loses its only advantage.</p><h2>The takeaway</h2><p>Shadow AI is not an employee discipline problem. It is an architecture problem, and Nate B. Jones&#8217;s two-minute test is its unit of measurement. Run the clock before your auditors, your regulators, or your competitors run it for you.</p><p>Rent the commodity, own the differentiation. And make owning it faster than leaking it.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Coordination Tax Is a Memory Problem]]></title><description><![CDATA[Ken Huang named the right tax. Here is where it lives, and how to stop paying it in 2026.]]></description><link>https://tokens.excipio.ai/p/the-coordination-tax-is-a-memory</link><guid isPermaLink="false">https://tokens.excipio.ai/p/the-coordination-tax-is-a-memory</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Thu, 23 Jul 2026 16:09:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xae_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xae_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xae_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png 424w, https://substackcdn.com/image/fetch/$s_!xae_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png 848w, https://substackcdn.com/image/fetch/$s_!xae_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!xae_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xae_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png" width="1080" height="1080" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1080,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:83937,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://tokens.excipio.ai/i/208217143?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xae_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png 424w, https://substackcdn.com/image/fetch/$s_!xae_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png 848w, https://substackcdn.com/image/fetch/$s_!xae_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!xae_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f02020a-2629-4766-9aab-cf448c80ec1a_1080x1080.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><a href="https://kenhuangus.substack.com/p/the-hidden-tax-on-ai-is-coordination">Ken Huang published an essay </a>this morning that deserves a wide read. Writing in his Agentic AI newsletter, he argues that the binding constraint on AI adoption has shifted. Model labor is getting cheaper faster than organizational coordination is getting easier. Once agents run for hours, split work across parallel branches, and touch real workflows, the expensive part is no longer producing tokens. Rather, it is deciding who can act, what gets handed off, which artifacts matter, and when a human steps in. He calls this the <strong>coordination tax</strong>, and he is right that most teams are not measuring it.</p><p>Read his piece first. It stands on its own. What follows is an extension, because the coordination tax connects directly to the framework this newsletter has been building all year. Coordination is not a new tax. It is where the taxes we already track go to compound.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Fan-out multiplies redundant inference</h2><p>Huang&#8217;s sharpest observation is that agent capacity now outruns human approval bandwidth. One operator can trigger ten parallel workstreams and generate more machine work in an afternoon than one person can casually supervise.</p><p>Here is what that fan-out does at the token layer. Parallel branches do not divide the work cleanly. They overlap. Each branch re-derives the same context, re-asks questions a sibling branch answered minutes ago, and re-buys conclusions the organization validated last week. The evidence for how much overlap already exists is stark:</p><ul><li><p>Datadog&#8217;s State of AI Engineering research found 69 percent of enterprise LLM input tokens are system prompts: instructions, policies, and tool guidance repeated verbatim across every call.</p></li><li><p>Only 28 percent of caching-capable calls in that research showed any cached reads. Most workloads pay full frontier pricing to re-process identical context.</p></li><li><p>Gartner estimates agentic AI consumes 5 to 30 times more tokens per task than chatbot workloads.</p></li></ul><p>Stack those three findings and the conclusion is uncomfortable. The more agentic a workload becomes, the more of its spend is re-purchase rather than purchase. This newsletter calls that the Rediscovery Tax: the cost of re-processing answers the organization already owns. Fan-out does not just create Huang&#8217;s approval queue. It multiplies the Rediscovery Tax by the branch count.</p><h2>Every handoff is a memory event</h2><p>Huang argues the design unit of agentic systems is no longer the prompt. It is the handoff: what job is delegated, what authority transfers, what evidence must survive the transition.</p><p>Restate that as an infrastructure question and it answers itself. Evidence survives a transition only if it has somewhere durable to live. In most agent stacks today it does not. A transformer holds no memory between sessions and keeps no stored records. When a handoff drops context, that is not an etiquette failure between agents. It is the absence of a memory layer. The receiving agent reconstructs what the sending agent knew by re-deriving it from scratch, at frontier prices, with fresh opportunities for drift.</p><p>The artifact glut Huang describes has the same root. Drafts, traces, and half-finished outputs pile up because nothing in the system knows which answer is authoritative, where it came from, or whether it is still true. That is the Stale Answer Tax in operation: answers that were correct when generated persist after their underlying source has moved.</p><h2>Review does not scale. Architecture does.</h2><p>The instinctive fix is more governance: more review checkpoints, more approval gates, more humans reconstructing what happened. Huang&#8217;s own math shows why that fails. Approval bandwidth grows slowly and irregularly while agent capacity grows almost vertically. Governance that depends on a person reviewing outputs works exactly as well as that person, on the week they have time.</p><p>The durable fix is structural. Three properties, built into how AI knowledge gets stored:</p><ul><li><p>Lineage: every answer knows the source it came from.</p></li><li><p>Invalidation: when that source changes, the answer expires automatically.</p></li><li><p>Source precedence: a standing rule for which source wins when two disagree, set before the conflict happens, not during it.</p></li></ul><p>This is governance by architecture, not by contract. It compresses coordination the way Huang says the winners will: shorter approval paths because authority is explicit, fewer retries because context survives handoffs, less rework because artifacts carry their own validity.</p><h2>How to evaluate a coordination-compression layer</h2><p>For teams acting on this in 2026, four questions separate real memory infrastructure from another dashboard:</p><ol><li><p>Ownership. Does validated knowledge accumulate in an asset you control, or inside a vendor&#8217;s model? Satya Nadella put this on the record at Davos in January 2026: a firm that cannot embed its knowledge in an asset it controls has no sovereignty. The sovereignty question is not where your data lives. It is where your intelligence compounds.</p></li><li><p>Lineage and invalidation. Can any answer show its source, and does it expire when the source changes? If governance lives in a review meeting instead of the storage layer, it will not survive fan-out.</p></li><li><p>Perimeter. When a known answer is served, does the query leave your network at all? Zero external transmission on a resolved query is a security property no audit trail can match.</p></li><li><p>Deployment cost. A memory layer that requires rearchitecting the agent stack recreates the coordination tax it claims to remove. The bar is a drop-in proxy: one BASE_URL change, live in hours.</p></li></ol><p>This is the problem Excipio was built for: a private memory layer and semantic proxy network that intercepts agent queries and resolves known answers from a customer-owned, compounding knowledge graph before they reach a frontier API. Modeled results show roughly a 42 percent blended reduction in LLM API spend on Day 1, with resolved queries returning in 8ms against roughly 2,500ms for a live frontier call. The cost reduction is the mechanism. The mission is the asset: validated institutional intelligence that stays inside the perimeter and compounds permanently, instead of leaking into someone else&#8217;s model on every repeated question.</p><p>Huang ends his essay with a dividing-line question: is the workflow getting smarter, or just generating more work for the people around it? A workflow gets smarter in exactly one way. It remembers. The coordination tax is real, it is growing with every parallel branch, and it is, at bottom, a memory problem.</p><p>Read Ken Huang&#8217;s original essay, <a href="https://kenhuangus.substack.com/p/the-hidden-tax-on-ai-is-coordination">The Hidden Tax on AI Is Coordination</a>, at his Agentic AI newsletter on Substack. It is worth your subscription.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The AI Bill Is the Last to Know]]></title><description><![CDATA[A microscope and a magnifying glass are precision instruments.]]></description><link>https://tokens.excipio.ai/p/the-ai-bill-is-the-last-to-know</link><guid isPermaLink="false">https://tokens.excipio.ai/p/the-ai-bill-is-the-last-to-know</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Thu, 16 Jul 2026 18:23:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!qDEo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A microscope and a magnifying glass are precision instruments. Both were built to examine what is already in front of you, in fine detail, after it arrived. Neither was built to repel an invasion. That is the mismatch sitting at the center of enterprise AI cost right now. The tools we have pointed at the problem are instruments of measurement. The problem is a force that has to be stopped upstream, before it lands.</p><p>The AI bill is a lagging indicator. By the time spend renders on a dashboard, the tokens are already gone and the data has already left the building. Every cost function in the enterprise today is ex-post by design. It sorts, attributes, and forecasts money that has already moved. That work is real and it is done well. It is also, by definition, downstream of the only moment that decides the outcome. The moment is the call itself.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Cloud cost discipline was built for resources that are elastic, observable, and throttleable. Redundant token spend is none of those. You cannot rightsize a question that was asked twice and answered twice at full price, with the answer walking out the door each time. Rightsizing optimizes what flows through the pipe. It was never designed to stop what should never have flowed at all. A magnifying glass shows you the breach in perfect detail. It does not close it.</p><p>The cost lever and the sovereignty lever are the same lever. Both are pulled at the input layer, at the instant the query is formed, by whoever controls what leaves. Today no one owns that layer, so the bill defaults to the team equipped only to report it. That is a category error, not a failure of effort. The real owner is architecture and the AI platform, the part of the org that decides what happens before the call, not after.</p><p>Boards have already moved. They stopped asking for token charts and started asking for cost per outcome. That number cannot be reported into existence. It is built. You architect your way to it, at the input layer, before the spend and before the byte leaves, or you do not reach it. No dashboard, however granular, produces a number that lives one layer above it.</p><p>So the reframe is simple. This was filed as a reporting problem. It is an architecture and action problem. Measurement teams are doing measurement well, and measurement is the wrong instrument for a force that has to be met at the perimeter. Hand the input layer to architecture and platform. Hand the sovereignty win to security, where nothing leaves on a resolved query. Cost and control then resolve at the source, by architecture, not by contract. You do not repel an invasion with a microscope. You decide what gets through the gate</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qDEo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qDEo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!qDEo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!qDEo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!qDEo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qDEo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:45665,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://tokens.excipio.ai/i/207325543?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qDEo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!qDEo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!qDEo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!qDEo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47a1cfc0-df95-4685-861c-d484a0a2b957_1200x630.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Fifth Estimate]]></title><description><![CDATA[What your dispatch rule prices at retail]]></description><link>https://tokens.excipio.ai/p/the-fifth-estimate</link><guid isPermaLink="false">https://tokens.excipio.ai/p/the-fifth-estimate</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Fri, 10 Jul 2026 23:28:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!A779!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!A779!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!A779!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png 424w, https://substackcdn.com/image/fetch/$s_!A779!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png 848w, https://substackcdn.com/image/fetch/$s_!A779!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!A779!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!A779!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png" width="1080" height="1080" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/df34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1080,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:89364,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://tokens.excipio.ai/i/206498152?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!A779!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png 424w, https://substackcdn.com/image/fetch/$s_!A779!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png 848w, https://substackcdn.com/image/fetch/$s_!A779!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!A779!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf34ef13-da3c-4ac9-ae5a-f6e80ab5896a_1080x1080.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Nate published <a href="https://natesnewsletter.substack.com/p/agent-shaped-work">a dispatch rule</a> this morning. Four things you can estimate about any task in about a minute: size, independence, separation, checkability. They resolve into one of four verdicts: a chat, one agent, a team, or don&#8217;t bother. Read it. It is the best budgeting frame I have seen for the question everyone with a freshly installed agent is quietly asking, which is &#8220;what do I even point this at?&#8221;</p><p>He opens with a number that should haunt every AI budget owner. More than 1.6 million agents signed up for an agents-only social network this year. Most were never asked to do a single thing. Not failed. Never dispatched. His diagnosis is right: nobody grew up with instincts for metered thinking. Thinking used to come attached to people. Now it is priced by the token and purchasable tonight.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>So the dispatch question is a budgeting question. Agreed. But every budgeting rule has a price assumption buried inside it, and this one prices every token at retail.</p><h2>The evidence that spend buys results</h2><p>Two research results anchor Nate&#8217;s case, and both hold up at the source.</p><p>Stanford&#8217;s Scaling Intelligence Lab took DeepSeek-Coder-V2-Instruct, a model that solves 15.9 percent of SWE-bench Lite issues in one attempt, and let it try 250 times per problem. Solve rate: 56 percent. That beat the single-attempt state of the art from frontier models. The paper is called <a href="https://arxiv.org/abs/2407.21787">Large Language Monkeys</a>, and its message is blunt. Coverage scales with samples across four orders of magnitude. Buying more attempts buys more answers.</p><p>Anthropic found the same thing from the other direction. In their analysis of <a href="https://www.anthropic.com/engineering/multi-agent-research-system">their multi-agent research system</a>, three factors explained 95 percent of performance variance on a hard browsing benchmark. Token usage alone explained 80 percent. Not prompt phrasing. Not orchestration cleverness. Tokens. Their multi-agent architecture beat a single agent by 90.2 percent on internal evals, and it did so by spending roughly 15 times the tokens of a chat interaction.</p><p>Both findings say the same thing: purchased thinking works. Spend more, get more. Which makes the dispatch decision exactly what Nate says it is, a question of whether the task justifies the spend.</p><p>Here is the assumption underneath: that every token in that spend was necessary.</p><h2>The fifth estimate</h2><p>Ask one more question about the task on your desk. Of the tokens this task will consume, how many purchase thinking your organization already owns?</p><p>Datadog measured what enterprise LLM traffic actually contains. 69 percent of input tokens are system prompts and repeated context, the same instructions and the same background re-transmitted call after call. Only 28 percent of caching-capable calls use caching at all. 72 percent of those workloads pay full frontier pricing to re-process identical context. And Gartner puts agentic AI at 5 to 30 times the tokens per task of a chatbot, which means every one of those redundancy ratios gets multiplied before it hits your invoice.</p><p>The architecture explains why. A transformer holds no memory between sessions and stores no records. When an agent seems to remember, the application is resending the entire transcript. Every new session gets a brilliant new consultant who has read everything and recalls nothing of the last meeting. So agents re-derive conclusions the organization validated last week, yesterday, or four seconds ago by the agent running next to them. I call this the Rediscovery Tax, and it is the single largest yield killer in agentic deployments.</p><p>Now look at Anthropic&#8217;s 15x multiplier again. Fifteen parallel context windows are fifteen consultants who cannot see each other&#8217;s notes. The very architecture that wins by spending tokens is the architecture where redundant spend compounds fastest. The 80 percent finding says token spend drives quality. It does not say wasted token spend drives quality. Those are different claims, and the gap between them is where your budget goes to die.</p><p>The metric is Token Yield: necessary spend divided by total spend. Every AI bill splits into unavoidable spend, the genuinely novel questions that legitimately require a frontier model, and recoverable waste, the tokens spent re-buying what you already know. No factory evaluates output per kilowatt. No logistics operator measures the fleet in gallons of diesel. Enterprise AI is the one budget line where we still confuse the meter with the metric.</p><h2>Why this moves Nate&#8217;s verdicts</h2><p>Run his four estimates on a task and suppose it lands on don&#8217;t bother. The economics fail: the task is agent-shaped, but the token bill exceeds the value of the output. That verdict was computed at retail, with every redundant re-derivation priced as if it were novel thinking.</p><p>Now compute it at yield. If a meaningful share of the task&#8217;s spend is recoverable waste, the effective cost of the necessary thinking is a fraction of the sticker price. Tasks migrate across the don&#8217;t-bother line. Not because the agents got smarter. Because you stopped paying full freight for thinking you already bought.</p><p>The same logic runs in reverse for the tasks you did dispatch. A team verdict at 15x tokens is only rational if those tokens are doing new work. If your agents spend most of their budget re-establishing context and re-validating each other&#8217;s conclusions, you did not buy a team. You bought one consultant fifteen times.</p><p>Nate&#8217;s rule tells you whether a task deserves purchased thinking. The fifth estimate tells you what that thinking should actually cost. Both questions take about a minute. Only one of them shows up on your invoice every month, forever, compounding with every agent you add.</p><p>The way out isn&#8217;t fewer tokens. It&#8217;s fewer wasted ones.</p><p>Sources: Nate&#8217;s dispatch guide, July 10, 2026. Brown et al., Large Language Monkeys: Scaling Inference Compute with Repeated Sampling, Stanford Scaling Intelligence Lab, arXiv 2407.21787. Anthropic, How we built our multi-agent research system, anthropic.com/engineering. Datadog State of LLM usage data and Gartner agentic consumption estimates as cited in The Enterprise Token Economy Issue 5.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Enterprise Token Economy! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[You Don't Build Databases]]></title><description><![CDATA[License the engine. Own the asset. The build-vs-license decision every AI company is about to make.]]></description><link>https://tokens.excipio.ai/p/you-dont-build-databases</link><guid isPermaLink="false">https://tokens.excipio.ai/p/you-dont-build-databases</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Thu, 09 Jul 2026 21:04:57 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f72affc2-e763-43a5-93f5-3579fee81a96_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Somewhere right now, an engineering leader at an AI company is sketching a memory layer on a whiteboard. A cache in front of the model APIs. A vector store for semantic matching. A table of validated answers. It looks like a quarter of work for two engineers. It looks like a smart way to cut the inference bill.</p><p>The industry has seen this whiteboard before. It had "database" written at the top.</p><h2>The decision every company already made</h2><p>No enterprise writes its own database engine. Not because they lack the talent. Because the visible ten percent of the problem, store a record and get it back, hides the brutal ninety percent: correctness under concurrency, crash recovery, query planning, replication, governance, performance at scale. Companies that tried spent years discovering requirements one production incident at a time. The market studied that outcome and reached a verdict so complete that nobody argues it anymore. You license the engine. You own the data.</p><p>Then the market went further. Enterprises paid a premium for managed databases so they would not have to operate the engine either. Then they paid for serverless databases so they would not have to think about capacity at all. Thirty years of buying behavior points one direction: abstraction up, DIY down. Enterprises pay more, on purpose, to do less infrastructure work. That is not laziness. That is discipline about where differentiation lives.</p><p>A compliance platform's moat is compliance judgment. An underwriting platform's moat is underwriting. A legal AI platform's moat is legal reasoning. None of them ever won a customer because of a homegrown storage engine. The engine was undifferentiated heavy lifting, and the market priced it accordingly.</p><h2>The same decision, arriving again</h2><p>AI memory is the database decision replaying at higher speed. Every company running agents at scale is about to choose: build the memory layer or license it.</p><p>The build looks easy for the same reason the database looked easy. The demo works in a week. An exact-match cache, an embedding model, a similarity threshold. The bill drops in the first test. The whiteboard wins the meeting.</p><p>Then the hidden ninety percent arrives. Cache invalidation is one of the two famously hard problems in computer science, and in an enterprise memory layer it is not a punchline. It is a compliance surface. An answer that was correct when generated goes stale the moment its source changes. Serve it anyway and you have not saved money. You have shipped the Stale Answer Tax straight into production, where acting on expired knowledge carries operational and regulatory risk.</p><p>The durable fix is structural, three properties built into how AI knowledge gets stored. Lineage: every answer knows the source it came from. Invalidation: when that source changes, the answer expires automatically. Source precedence: a standing rule for which source wins when two disagree, set before the conflict happens, not during it. That triad is <a href="https://excipio.ai/governance-by-architecture">governance by architecture, not by contract</a>. It is also months of engineering that nobody budgeted on the whiteboard, built by people whose actual job was the product.</p><h2>Even a perfect build buys you one pillar</h2><p>Here is the part the build-vs-license spreadsheet misses. Suppose the in-house team executes. Eighteen months, no incidents, the cache works. What did they get?</p><p>Some cost reduction. One pillar out of four.</p><p>A licensed memory layer ships the full set. Cost: roughly 42 percent off blended token spend on day one, not in eighteen months. Latency: cache hits in about 8 milliseconds against 2,500 from a frontier round trip, a 312x gap that is an engineering moat in its own right, not a side-project deliverable. Portability: model-agnostic by design, so the layer survives the next model generation, while an in-house build marries whichever API it was written against. And IP security: zero bytes leave the perimeter on a cache hit, with lineage and invalidation already in the architecture instead of on the roadmap.</p><p>The pillars are not a feature list to evaluate. They are the parts of the build the in-house team was never going to reach, because each one is a product in itself and their product is something else.</p><h2>The one place the analogy breaks, in your favor</h2><p>Databases came with a trade. Self-hosted meant sovereignty and operational burden. Managed and serverless meant abstraction and someone else's perimeter. You picked one.</p><p>The memory layer does not force that pick. <a href="https://excipio.ai">Excipio</a> deploys as a drop-in semantic proxy network inside the customer's own perimeter, live in hours, one endpoint change, zero rearchitecting. The operational abstraction of serverless. The sovereignty of self-hosted. And the part that compounds, the customer-owned, compounding knowledge graph of validated institutional intelligence, is the customer's asset from the first cached answer.</p><p>That distinction is the whole strategic point. Satya Nadella named it at Davos in January 2026: a firm that cannot embed its knowledge in an asset it controls has no sovereignty, and is leaking enterprise value to a model company on every query. The database era's system of record held your transactions. The intelligence era's system of record holds your validated answers. Nobody wrote their own system of record last time. The question is only who owns what accumulates inside it, and the answer has to be you.</p><h2>The rule</h2><p>License the engine. Own the asset. The engine is the vendor's job. The intelligence is yours. Confusing the two is how a company ends up building a database in 2026, except this time the failure mode is not a slow query. It is institutional knowledge compounding in someone else's model while six engineers rediscover why invalidation was famous.</p><p>The knowledge graph you establish in year one is the moat that compounds permanently. The one you build in year three starts three years behind. And the one you try to write yourself starts behind and stays there.</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://tokens.excipio.ai/subscribe?"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Context as the New Moat]]></title><description><![CDATA[The CEOs stopped arguing about models this year. They started arguing about who owns your context.]]></description><link>https://tokens.excipio.ai/p/context-as-the-new-moat</link><guid isPermaLink="false">https://tokens.excipio.ai/p/context-as-the-new-moat</guid><dc:creator><![CDATA[Tony Wenzel]]></dc:creator><pubDate>Wed, 08 Jul 2026 14:13:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_6DL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Something shifted in the executive conversation about AI this year, and it happened fast. Twelve months ago, every keynote was a model announcement. Benchmarks, parameter counts, reasoning scores. This year the most quoted lines from the most powerful people in enterprise technology are not about intelligence at all. They are about context. And if you read them carefully, they are telling you exactly where the next decade of enterprise AI value will accumulate, and who intends to capture it.</p><p>## The chorus</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Start with Ali Ghodsi. The Databricks CEO opened his Data + AI Summit keynote this June with a blunt diagnosis: &#8220;AI doesn&#8217;t have an intelligence problem. It has a context problem.&#8221; He had the numbers to back it. PwC found that 56% of CEOs report zero financial benefit from AI. The share of companies scrapping AI projects before production jumped from 17% to 42% in a single year. Over the same period, models got dramatically better and API costs collapsed. Intelligence improved. Results got worse. The missing variable is not in the model.</p><p>Satya Nadella said a version of the same thing from the Davos stage in January, and regular readers, if I have any,</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_6DL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_6DL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!_6DL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!_6DL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!_6DL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_6DL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:60709,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://wenzewenzel938942.substack.com/i/206052173?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_6DL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!_6DL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!_6DL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!_6DL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef17ca2-df01-44cc-9761-075bf951ac59_1200x630.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p> know I have written about it before: a firm that cannot embed its own knowledge in an asset it controls has no sovereignty, and is leaking enterprise value to a model company on every call. He predicted it would become the most talked about topic in AI this year. Six months in, he is winning that bet.</p><p>Pat Grady opened Sequoia&#8217;s AI Ascent with a warning to every founder in the room: &#8220;The things that you build might be irrelevant tomorrow.&#8221; The models move faster than any product built on top of them. Amazon&#8217;s top AI executive, Peter DeSantis, put the economic version to The Wall Street Journal in four words: &#8220;AI has a cost problem.&#8221; And Prukalpa Sankar, co-CEO of Atlan, compressed the whole thesis into the sharpest line of the year: while intelligence converges, context compounds.</p><p>Different companies, different incentives, same conclusion. Intelligence is commoditizing. Context is not.</p><p>## The tell</p><p>Now listen to what the model providers themselves are saying, because this is where it gets interesting.</p><p>Sam Altman has spent 2026 telling anyone with a podcast that memory is OpenAI&#8217;s real moat. Not model capability. Memory. His stated ambition is an AI with what he calls infinite, perfect memory: every document, every email, every decision you have ever consulted it on. He describes personalization as &#8220;extremely addictive&#8221; and says users who invest their history into ChatGPT will find it very hard to leave. We are, in his words, in the GPT-2 era of memory.</p><p>Read that carefully. The CEO of the largest model company in the world is telling you, in public, that his moat is built from your context. Your questions, your documents, your decision history, accumulated on his infrastructure, creating switching costs that keep you paying him. He is not hiding it. It is the strategy.</p><p>For a consumer, that might be a fair trade. For an enterprise, it is the sovereignty problem stated as a product roadmap. Every query your agents send to a frontier model is a small deposit into someone else&#8217;s moat.</p><p>## Rented context, owned context</p><p>Here is the structural choice underneath all of this. Every time one of your AI agents fires a query at a frontier LLM, one of two things happens.</p><p>**Rented context.** The query goes out. The model answers from its general training. The exchange is discarded, or worse, it accumulates on the provider&#8217;s side of the ledger. Your agent got an answer. You paid for it. You own nothing from the transaction, and you will pay full price to ask a semantically identical question tomorrow.</p><p>**Owned context.** The query is intercepted before it leaves your perimeter. The validated answer is stored in a customer-owned, compounding knowledge graph, alongside the semantic fingerprint of the question. The exchange becomes part of a growing corpus of proprietary intelligence about how your specific domain works: your terminology, your edge cases, your approval patterns, your exceptions. The intelligence stays inside. It compounds permanently.</p><p>Most enterprises are operating entirely in the first mode. The data says so: roughly 95% of enterprise AI usage still runs on frontier models, per Glean&#8217;s CEO, and Datadog&#8217;s telemetry shows 69% of enterprise LLM input tokens are system prompts and repeated context. That is the Rediscovery Tax at industrial scale. Firms are renting their own institutional knowledge back from an external model, one API call at a time.</p><p>## Why this compounds</p><p>The strategic implication takes a moment to land. An enterprise operating in owned-context mode is accumulating something the rented-context enterprise never will. A knowledge graph that grows with production usage is an AI system that gets smarter, faster, and cheaper over time. Not because the underlying model improved. Because your proprietary layer did.</p><p>Domain questions start resolving locally in milliseconds instead of round-tripping to a frontier API. The share of agent traffic answered from owned intelligence climbs as the graph matures. Token spend drops compoundingly, because more of the work resolves against knowledge the firm already validated and already owns.</p><p>And here is the question a software executive asked me recently that I have been thinking about since: if the model providers cut prices again, does the owned-context advantage shrink? It does not. It grows. Cheaper frontier inference makes the genuinely novel queries more affordable, which enables more agent deployment, which generates more validated answers, which deepens the graph. The moat compounds regardless of what happens to frontier pricing. Lower model prices are a tailwind for the enterprises that own their context and a treadmill for the ones that rent it.</p><p>This is Context as the New Moat. It is the difference between an AI system that costs the same per query forever and an AI system that becomes a durable competitive asset. And it is the one moat the model providers cannot absorb, because replicating it would mean commoditizing their own inference revenue.</p><p>## Two questions for your vendors</p><p>If you are evaluating AI infrastructure right now, two questions cut through most of the noise.</p><p>First: is this a capability the model provider is likely to build natively? If yes, understand what you actually own when they do. Features get absorbed. Infrastructure between you and the provider does not.</p><p>Second: does my usage of this platform accumulate proprietary intelligence I could not recreate elsewhere? If you could swap the vendor tomorrow and lose nothing but integration work, you are renting context, not building a moat. And if the intelligence accumulates on the vendor&#8217;s side rather than yours, you are building someone else&#8217;s.</p><p>The CEOs have already told you where value accrues next. The only open question is whose balance sheet your context compounds on.</p><p>## One thing to read</p><p>Jaya Gupta&#8217;s &#8220;Context graphs: AI&#8217;s trillion-dollar opportunity&#8221; from Foundation Capital. The venture-side articulation of the same thesis: the next platforms will be built on persistent records of enterprise decisions, not on better models. Systems of record store what happened. Context graphs store why. Worth your time.</p><p>## About Excipio</p><p>Excipio is a private memory layer for enterprise AI, built in C++, sitting between your agents and the frontier models. We intercept every agent query, serve validated answers from your customer-owned, compounding knowledge graph in under 10ms, and route only genuinely novel questions to the cheapest capable model. Roughly 42% blended token cost reduction. Zero bytes leave the perimeter on a cache hit. Zero changes to your agent code.</p><p>If you are running AI agents at scale and want your context compounding on your balance sheet instead of someone else&#8217;s, I would like to talk.</p><p>**tony@excipio.ai** &#183; **excipio.ai**</p><p>*The Enterprise Token Economy publishes on Substack. Forward to anyone building or deploying AI agents at scale.*</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://tokens.excipio.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>