12 min read

Moneyball for AI: Why Tokens, Teams, and Trust Matter More Than Raw Benchmarks

Chris Potts's "tokenflation" highlights a critical ROI re-evaluation as AI conversation shifts to economics, trust, and collaboration, making cost/efficiency paramount for agentic workloads.

Moneyball for AI: Why Tokens, Teams, and Trust Matter More Than Raw Benchmarks

📊 12 episodes across 8 podcasts

⏱ 565 minutes of intelligence analyzed

🎙 Featuring: Brian Armstrong (Coinbase), Elad Gil (Host), Rhett Alden (Elsevier Health Markets), Yolandi de Weerdt (Emerj)


Presented by

Velocity Road

Turn AI spend into EBITDA—with proof every quarter.

Velocity Road finds where AI actually pays, installs the governance your board can trust, and builds the systems that move the number. Run all year. Proven every quarter.

Learn More

The Big Shift

The AI conversation has shifted dramatically from raw model benchmarks to the economics, collaboration, and societal trust required for real-world deployment. It's no longer just about who has the biggest, smartest model; it's about the pragmatic realities of "tokenflation," building shared agentic workflows, and navigating a growing public backlash against AI's physical footprint.

The New Metric: For the past year, it's been a race to the top of the leaderboard. This week, we heard the first real challenges to that paradigm. Chris Potts, a Professor and Co-founder at Stanford University & BigSpin, introduced "tokenflation" on The TWIML AI Podcast, arguing that "the purchasing power of tokens... is going down." This signals a critical re-evaluation of ROI, moving beyond simply model intelligence to its economic sustainability. Costs and efficiency are becoming paramount, especially for agentic workloads that consume vast quantities of tokens.

"What we see with Opus 4.6 usage in the time period we have... a decline in the purchasing power of tokens, that CPI is going down... is it because we have the wrong basket of goods or is it because we're actually getting less value from these tokens?"
— Chris Potts, Co-founder at BigSpin on The TWIML AI Podcast

Collaboration as the Next Frontier: The era of solo AI is ending. Nathaniel Whittemore, Host of The AI Daily Brief, highlighted the rise of Multiplayer AI 🆕, predicting that "the shift from single player to multiplayer AI is going to be one of the biggest trends to finding the best, most successful AI using companies." This isn't just theory; Anthropic's internal Claude Tag, a shared agent in Slack, now generates 65% of their product team's code, demonstrating how collaborative AI agents are transforming team workflows.

The Public Weighs In: Beyond the tech, the public is starting to push back on the physical infrastructure of AI. Journalist Jasmine Sun, on Hard Fork, reported on a growing data center backlash 🆕 in the Midwest, where communities are uniting across political lines to oppose facilities due to noise, environmental impacts, and a deep distrust of tech companies. This "AI populism" views AI as an elite project rather than a shared benefit, complicating deployment and highlighting the need for genuine community engagement.

The Move: Companies must now scrutinize the economics of their AI deployments, pivot to collaborative AI strategies, and proactively address societal concerns around infrastructure and job displacement, rather than solely chasing benchmark scores.


The Rundown

① Computer-use agents are now indispensable for personal productivity.

Demetrios Brinkmann, Head of Development Experience at Agentic AI Foundation, shared on Practical AI how OpenAI's computer-use skill has dramatically improved his personal workflow, automating complex tasks like flight bookings and email management without traditional APIs.

Why it matters: This isn't just theoretical; it's a critical shift in how high-leverage individuals are augmenting their daily work, bypassing traditional software interfaces and directly manipulating applications.

② AI models are becoming payment processors, bypassing traditional e-commerce.

On Practical AI, Demetrius and Chris Benson discussed the potential for AI companies like OpenAI to disintermediate platforms like Amazon by becoming the marketplace and payment processor within an agentic shopping ecosystem, hinting at a future where "chat is the new browser."

What to watch: This model-as-marketplace shift could fundamentally alter e-commerce, forcing incumbents to adapt or face disintermediation as AI agents handle transactions directly.

③ "Tokenflation" is eroding the value of AI usage, demanding new efficiency metrics.

Chris Potts, Co-founder at BigSpin, noted on The TWIML AI Podcast that Anthropic's Opus 4.6 showed a decline in token purchasing power between February and April, raising questions about the actual value received despite supposed model improvements.

The context: Businesses must focus on "engineering goods" CPI (Consumer Price Index) relative to token usage, rather than just model performance, to ensure economic sustainability and ROI from AI deployments.

④ OpenAI is quietly building a "recursive self-improvement" AI for its own software development.

Brian Armstrong, CEO of Coinbase, revealed on No Priors that Coinbase is feeding human edits to AI-generated code back into a 'brain' for each team/repository, ensuring continuous learning and improvement in future code changes within their internal AI system.

Why it matters: This is a powerful, real-world example of AI improving its own capabilities in a closed-loop system, demonstrating a significant step towards autonomous software development and potentially supercharging internal development cycles.

⑤ Frontier coding agents spend 86% of their time reading, not solving.

Alexander Whedon, Co-founder and CTO of Subquadratic, explained on Eye On A.I. that "86% of the steps were actually read steps" for context gathering before actual problem-solving, highlighting a massive inefficiency in current agent architectures.

What to watch: Innovations like Subquadratic's Sparse Attention Mechanism 🆕, offering 40x faster inference and 64x less compute at 1 million tokens, are crucial for making coding agents truly efficient and scalable.

⑥ AI safety concerns are amplified by media and political incentives, not just technical risk.

NLW on The AI Daily Brief discussed how viral warnings about AI existential risk, like those from Anthropic researchers, gain political and media traction due to increased bipartisan interest in AI and media incentives, overshadowing common-sense regulation.

The context: This creates a challenge for leaders to discern genuine, actionable risks from performative doomerism, pushing for policy responses that are specific and address concrete issues like biorisk rather than vague hypotheticals.

⑦ Healthcare AI governance is evolving in a regulatory vacuum.

Rhett Alden, CTO at Elsevier Health Markets, shared on The AI in Business Podcast that healthcare institutions are establishing their own cross-functional AI governance boards, similar to clinical trial review structures, because the FDA in the US is not yet regulating AI.

Why it matters: This forces healthcare leaders to proactively develop robust internal frameworks for claim-level validation 🆕 and epistemic trust, ensuring AI reliability and maintaining clinician responsibility in the absence of external standards.

⑧ Data center expansion is triggering an "AI populism" backlash in unexpected places.

On Hard Fork, Jasmine Sun, Journalist and Writer at jasmi.news, reported on how local communities in the Midwest are opposing data centers due to concerns about aesthetics, noise, and electricity consumption, viewing AI as an "elite political project."

The context: This populist resistance, fueled by low trust in corporations and government, could significantly complicate future AI infrastructure build-outs and requires tech companies to re-evaluate their community engagement strategies.


The Signals

🌍 Heating Up

Multiplayer AI 🆕: Y Combinator's investment thesis and Anthropic's internal success with CLAUDE Tag point to collaborative AI agents as the next frontier for team workflows. (Nathaniel Whittemore on The AI Daily Brief)

Agentic Finance: Coinbase CEO Brian Armstrong sees 76% of agentic commerce transactions under 30 cents, necessitating crypto rails. (Brian Armstrong on No Priors)

Opportunity AI: NLW emphasizes that models like Astra (GPT-6) 🆕 unlock novel capabilities beyond just efficiency, requiring a new mindset for adoption. (NLW on The AI Daily Brief)

👀 On Watch

Data Center Backlash (Midwest) 🆕: Communities are uniting against data centers, driven by distrust and environmental concerns, indicating political and social friction for AI infrastructure. (Jasmine Sun on Hard Fork)

AI Tokenomics and Tokenflation: The declining purchasing power of tokens suggests a need for new economic models and ROI metrics for AI. (Chris Potts on The TWIML AI Podcast)

Computer-Use Agents: OpenAI's advanced computer-use skill is enabling significant personal and enterprise productivity gains, automating complex tasks across web interfaces. (Demetrios Brinkmann on Practical AI)

Sparse Attention Mechanism 🆕: Subquadratic's breakthrough dramatically cuts compute and speeds inference, addressing a core inefficiency in long-context transformer models. (Alexander Whedon on Eye On A.I.)

❄️ Cooling Off

Black Box Algorithms in Healthcare: Clinicians' fear of non-transparent AI and the lagging FDA regulation are pushing healthcare institutions to create their own governance models. (Rhett Alden on The AI in Business Podcast)

Benchmark Maxing: AI models are being optimized for public benchmarks rather than generalized performance, leading to questions about their true utility in real-world scenarios. (NLW on The AI Daily Brief)


The Debate

Are advanced AI models a threat to humanity, or are warnings overblown?

🐂 The bull case: Evan Hubinger, Alignment Science Lead at Anthropic, believes there is a "greater than 10% chance" of AI causing human extinction within the next decade, with Jacob Coxon, a former Anthropic researcher, echoing that "the people building AI earnestly believe that it could kill us all by the end of the decade." These warnings are gaining political traction, indicating a serious, growing concern from those closest to the technology. (NLW on The AI Daily Brief)

🐻 The bear case: Critics argue that "doomer" claims lack specificity and that policy proposals often risk "the cure worse than the disease." Sam Liu, a PhD dropout in AI safety, stated his "biggest pet peeve is that no one can really provide tangible pathways to why it matters," suggesting that the focus should be on concrete issues like biorisk rather than vague existential threats. (NLW on The AI Daily Brief)

Our read: The debate is highly polarized, but the political and media amplification of extreme views means concrete, actionable policy is being overshadowed by hypotheticals.


The Bottom Line

AI's future isn't just about raw intelligence; it's about making the economics work, building for team collaboration, and navigating a growing public demand for control over its physical and societal impact.


Episode Guide (Web Version)

No Priors: Artificial Intelligence | Technology | Startups — "Coinbase’s Everything Exchange: Agentic Finance, Stablecoins, and Tokenization with CEO Brian Armstrong"

Runtime: 45 min | Host: Elad Gil | Guest: Brian Armstrong (Co-founder and CEO, Coinbase)

Listen if: You're charting the future of finance or building agentic systems that require micro-transactions.

Coinbase CEO Brian Armstrong details his vision for "agentic finance," where AI agents need their own financial accounts, and the internal recursive self-improvement of AI in software development at Coinbase.

"About like 76% of the agentic commerce transactions we're seeing are under 30 cents."
— Brian Armstrong, CEO of Coinbase

▶ Listen · Apple Podcasts

The AI Daily Brief: Artificial Intelligence News and Analysis — "Why GPT-6 Astra Is So Significant and So Confounding"

Runtime: 29 min | Host: Nathaniel Whittemore | Guest: NLW (Host, The AI Daily Brief)

Listen if: You're trying to understand how new AI models fundamentally change interaction patterns beyond simple prompting.

NLW explores OpenAI's Astra (GPT-6) 🆕, an "opportunity AI model" expanding what users can do with agentic tasks, coding, and 3D modeling, despite its struggles with front-end UI design.

"GPT-6 Astra is not an efficiency AI model. In other words, it is not about doing what you currently do better. It is through and through an opportunity AI model that is going to challenge you to think differently about what you can do."
— Nathaniel Whittemore, Host of The AI Daily Brief

▶ Listen · Apple Podcasts

Practical AI — "Computer-Use Agents and the Future of the Agentic Internet"

Runtime: 56 min | Host: Chris Benson | Guest: Demetrios Brinkmann (Head of Development Experience at Agentic AI Foundation, Founder of MLOps Community)

Listen if: You're seeking practical applications for AI agents in automating personal and enterprise tasks, or considering agentic commerce implications.

Chris Benson and Demetrios Brinkmann discuss the rapid evolution of computer-use AI agents, highlighting personal productivity gains and the challenges of integrating these agents into enterprise environments, including the blurring lines between models and harnesses.

"Since that came out, if you specifically use the computer use skill, it is so good, there are very few things that I can throw at it that it does not accomplish."
— Demetrios Brinkmann, Head of Development Experience at Agentic AI Foundation

▶ Listen · Apple Podcasts

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) — "Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776"

Runtime: 59 min | Host: Sam Charrington | Guest: Chris Potts (Professor & Co-founder, Stanford University & BigSpin)

Listen if: You're focused on the economic sustainability of AI models and the critical need for architectural innovation beyond current transformer limitations.

Sam Charrington and Chris Potts delve into AI tokenomics, introducing "tokenflation" and questioning the ROI of increasing token usage. They also discuss architectural innovation and the critical role of user expertise in interacting with AI.

"What we see with Opus 4.6 usage in the time period we have... a decline in the purchasing power of tokens, that CPI is going down... is it because we have the wrong basket of goods or is it because we're actually getting less value from these tokens?"
— Chris Potts, Co-founder at BigSpin

▶ Listen · Apple Podcasts

The AI Daily Brief: Artificial Intelligence News and Analysis — "10 Ways to Think Bigger with Opportunity AI"

Runtime: 28 min | Host: NLW | Guest: NLW (Host, The AI Daily Brief)

Listen if: You want concrete ideas on how to leverage AI for novel capabilities rather than just improving existing tasks.

NLW distinguishes "Opportunity AI" from "Efficiency AI," offering thought starters for new applications like interactive marketing games, custom video production, and even productizing personal expertise using advanced models like Astra (GPT-6) 🆕.

"The value of new models is very frequently not going to be in just doing the same things that you've been doing with AI better, but actually about totally unlocking new capabilities that you've never even considered."
— NLW, Host of The AI Daily Brief

▶ Listen · Apple Podcasts

The Neuron: AI Explained — "OpenAI Astra, Local AI, and the Hardware Race"

Runtime: 95 min | Host: Corey Noles | Guest: Jacob Pachocki (Chief Scientist, OpenAI)

Listen if: You're tracking the latest AI model advancements, hardware developments, and the safety implications of recurrent depth techniques.

Corey Noles and Grant Harvey discuss OpenAI's Astra (GPT-6) 🆕 capabilities, the controversy around its recurrent depth technique, and the re-emerging importance of CPUs alongside GPUs for local and cloud AI, including the new Gemini 3.8 Flash 🆕 model.

"ASTRA saturated Frontier Math Tier 4 with a 98% score. When they say the phrase saturated, what that means is that basically it's solved. There's really no points in trying to get that 100% here because it's close enough that 98% of the time it's going to give you the right answer."
— Grant Harvey, Host at The Neuron

▶ Listen · Apple Podcasts

The AI Daily Brief: Artificial Intelligence News and Analysis — "AI Model Month Is Off to a Blistering Start"

Runtime: 34 min | Host: Nathaniel Whittemore | Guest: NLW (Host, The AI Daily Brief)

Listen if: You're trying to make sense of the rapid pace of AI model releases and the shift towards multi-model architectures.

NLW covers the shift from single to multi-model architectures, OpenAI's controversial Navier-Stokes claim, new model releases like Gemini 3.8 Flash 🆕 and Meta's MuSpark 1.3, and the "benchmark maxing" debate.

"At what point does speed matter more than the last bit of quality?"
— NLW, Host of The AI Daily Brief

▶ Listen

Hard Fork — "The Ezra Klein Show: The A.I. Revolt Is Here"

Runtime: 80 min | Host: Ezra Klein | Guest: Jasmine Sun (Journalist and Writer, jasmi.news)

Listen if: You want to understand the growing public backlash against AI infrastructure and how it's shaping political discourse.

Ezra Klein and Jasmine Sun 🆕 explore the populist opposition to AI data centers in the Midwest, highlighting community concerns over environmental impacts, distrust in corporations, and the concept of "AI populism."

"It was really clear to me that data centers are showing up in an environment of extremely low trust in both governments and in corporations to the extent where the pro arguments almost do not land because people just aren't interested in anything an outside tech company is going to tell them."
— Jasmine Sun, Journalist and Writer at jasmi.news

▶ Listen · Apple Podcasts

The AI Daily Brief: Artificial Intelligence News and Analysis — "The Multiplayer AI Sprint: Build Your Team’s First Shared Agent"

Runtime: 26 min | Host: Nathaniel Whittemore | Guest: NLW (Host, The AI Daily Brief)

Listen if: You're looking to transition your team from individual AI use to collaborative, shared agentic workflows.

NLW introduces the concept of Multiplayer AI 🆕 and a free learning program, "The Multiplayer AI Sprint for Teams," to help organizations build shared agents, citing examples like Anthropic's CLAUDE Tag 🆕.

"The next frontier of agent design is going to move agents from the individual silos in which they have operated so far to the shared spaces that teams inhabit together."
— Nathaniel Whittemore, Host of The AI Daily Brief

▶ Listen · Apple Podcasts

The AI Daily Brief: Artificial Intelligence News and Analysis — "Anthropic Researcher Says AI Has Over a 10% Chance of Killing All Humans"

Runtime: 36 min | Host: Nathaniel Whittemore | Guest: Jacob Coxon (Former Researcher, Anthropic)

Listen if: You're navigating the ethical debates around AI safety and the political and media amplification of existential risk warnings.

NLW discusses Anthropic researchers' warnings about AI existential risk, the political and media traction these warnings gain, and strong counter-arguments against AI doomerism, emphasizing the need for concrete policy.

"I personally think it is greater than 10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve Alignment for Superintelligence."
— Evan Hubinger, Alignment Science Lead at Anthropic

▶ Listen · Apple Podcasts

The AI in Business Podcast — "Responsible Generative AI in Healthcare and What Leaders Need to Know for 2026 - with Rhett Alden of Elsevier Health"

Runtime: 22 min | Host: Yolandi de Weerdt | Guest: Rhett Alden (Chief Technology Officer, Elsevier Health Markets)

Listen if: You're a healthcare leader grappling with AI adoption, clinical trust, and the need for robust internal governance in a rapidly evolving regulatory landscape.

Rhett Alden 🆕, CTO at Elsevier Health Markets, discusses the critical need for transparency and claim-level validation in healthcare AI, highlighting how institutions are establishing AI governance boards in the absence of clear FDA regulation.

"The biggest challenge for clinicians or for health institutions is the fear of black box algorithms. They're used to things being validated scientifically, they're used to predictability, very used to precision."
— Rhett Alden, Chief Technology Officer at Elsevier Health Markets

▶ Listen · Apple Podcasts

Eye On A.I. — "86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic"

Runtime: 55 min | Host: Craig Smith | Guest: Alexander Whedon (Co-founder and CTO, Subquadratic)

Listen if: You're looking to overcome inefficiencies in coding agents and explore next-generation transformer architectures.

Alexander Whedon (Alex) 🆕, Co-founder and CTO of Subquadratic (SubQ) 🆕, reveals that coding agents spend 86% of their time "reading" for context, and discusses how their sparse attention mechanism dramatically reduces compute and speeds up inference, enabling faster long-context training.

"We found that 86% of the steps were actually read steps. Just trying to do that context engineering before the execution review, which was only the last 14%."
— Alexander Whedon, Co-founder and CTO of Subquadratic

▶ Listen · Apple Podcasts

PARTNER

Not sure where AI fits in your operations? Start with the data.

Velocity Road's AI Readiness Assessment maps your organization against 7 operational dimensions and shows exactly where AI creates ROI — in under 10 minutes.

Take the Assessment -> →

Avi Savar

Get AI & Technology in your inbox

How AI and Tech are reshaping business. Free.