11 min read

Simile AI Models Human Flaws at 85% Accuracy

Simile AI advances AI adoption, modeling human-like behavior, biases, and "social physics" to simulate populations accurately for more realistic predictions.

Simile AI Models Human Flaws at 85% Accuracy

The next frontier isn't just about bigger models; it's about making AI models behave like us, even with our flaws, and getting them off the cloud into our hands and our workflows.


📊 12 episodes across 9 podcasts

⏱ 583 minutes of intelligence analyzed

🎙 Featuring: Eric Porres (Logitech), Henrik (Beyond The Prompt), Jeremy (Beyond The Prompt), Marilie Fouché (The AI in Business Podcast)


Presented by

Velocity Road

Turn AI spend into EBITDA—with proof every quarter.

Velocity Road finds where AI actually pays, installs the governance your board can trust, and builds the systems that move the number. Run all year. Proven every quarter.

Learn More

The Big Shift

Forget the dream of perfectly rational AI agents. This week, we're hearing a powerful signal that the real game-changer in AI adoption isn't about perfect intelligence, but about modeling human-like behavior, including our mistakes and biases. Companies like Simile AI are making significant progress in simulating individuals and even entire populations with surprising accuracy by focusing on "social physics" over pure logic.

Why it matters: While frontier LLMs are optimized for rationality, they often fall short in predicting human actions because they miss the nuances of human biases and 'social physics.' This focus on behavioral modeling opens up entirely new avenues for AI application, especially in areas like drug discovery, policy-making, and even marketing, where understanding human irrationality is key.

"Simile actually doesn't care about any of this. The models that we're talking about here, what we're trying to create are models that are as dumb as I am. Right. So if I make such mistakes, the model has to make the same kind of mistake."
— Joon Sung Park, Co-founder/CEO of Simile AI on Latent Space: The AI Engineer Podcast

The context: This approach isn't just theoretical. Simile AI is achieving 85% accuracy in replicating individual behaviors, vastly outperforming current frontier models (20-60%) in this domain (Joon Sung Park on Latent Space). This capability allows for rapid iteration and hypothesis testing in virtual environments, compressing R&D timelines from months to days, as seen in drug discovery with Turbine (Kristóf Szalay on The AI in Business Podcast).

The move: Prioritize understanding and leveraging AI's capacity for human behavior modeling. Explore how behavioral simulation can de-risk new product launches, test market reactions, or even optimize internal processes by predicting employee responses, rather than solely relying on generalized 'smart' AI. The shift from pure intelligence to contextual, behavioral understanding is where the next wave of value lies.


The Rundown

① The AI compute race is moving beyond just chips to infrastructure financing.

NVIDIA is directly financing a 4.25 gigawatt AI data center for OpenAI in Ohio as part of a massive $600 billion compute deal, signaling a shift where chip makers are securing their own demand by funding customer infrastructure (AI Breakdown).

Why it matters: This isn't just selling hardware; it's a strategic move to ensure compute capacity for key partners, deepening integration and potentially locking in future demand, while also making AI adoption more accessible for large-scale projects.

② AI is shaking up high-level mathematics, sparking an "existential crisis" for the field.

OpenAI's Astra model solved multiple long-standing math problems, and while AI still struggles with basic arithmetic, its proficiency in abstract math has shocked the mathematical community, raising questions about academic grants and the role of human mathematicians (Robert Hart on Decoder with Nilay Patel).

The context: The speed of AI advancement in complex problem-solving is outstripping expectations, challenging traditional academic structures and forcing a re-evaluation of what constitutes a "human" contribution in intellectual fields.

③ Logitech's Chief AI Officer mandates internal AI 'Board Advisor' use for executive presentations.

Eric Porres revealed that Logitech leadership is expected to run board presentations through an AI "Board Advisor" gem for feedback before meetings, embedding AI directly into critical executive workflows (Eric Porres on Beyond The Prompt).

What to watch: This indicates a maturation of AI adoption, moving from optional experimentation to a mandatory tool for strategic decision-making, setting a precedent for how organizations can institutionalize AI's role in high-stakes environments.

④ Public opposition to AI data centers is escalating, becoming a bipartisan political issue.

Pennsylvania Governor Josh Shapiro's executive order on data centers, though pragmatic, highlights how animosity towards these facilities now exceeds that for nuclear power plants among voters (Nathaniel Whittemore on The AI Daily Brief).

What to watch: The political landscape for AI infrastructure is hardening; expect stricter regulations, community benefit requirements, and potentially higher costs for data center development as local communities demand more from these energy-intensive projects.

⑤ A "premium token economy" is emerging where users pay 10x more for low-latency AI inference.

Sid Sheth of d-Matrix explains that for instant responses and high interactivity, users are willing to pay significantly more, a market segment where traditional GPU architectures are not optimally suited due to memory bandwidth limitations (Sid Sheth on Eye On A.I.).

Why it matters: This bifurcation in the AI chip market signals specialized demand for architectures optimized for inference, driving innovation in areas like Low Latency Compute 🆕 and creating opportunities for companies that can deliver speed and interactivity over brute-force compute.


The Signals

🔥 HEATING UP

Simile AI: Their human behavior simulation models are achieving 85% accuracy, promising to revolutionize R&D and decision-making by predicting human-like mistakes and social physics. (Joon Sung Park on Latent Space)

Low Latency Compute 🆕: A premium token economy is emerging, with users willing to pay 10x more for instant, highly interactive AI responses, driving demand for specialized chip architectures that can deliver speed. (Sid Sheth on Eye On A.I.)

Deletions as a metric for innovation: Logitech's Chief AI Officer uses 'deletions' (of old, inefficient processes) as a key indicator of successful AI-driven organizational transformation. (Eric Porres on Beyond The Prompt)

👀 ON WATCH

rack scale 🆕: New rack scale deployments are reducing loss from grid to chip by a factor of two, offering significant energy efficiency gains for AI data centers. (Drew Baglino on Gradient Dissent)

Unsloth Desktop 🆕: This tool allows users to run powerful quantized AI models locally on consumer hardware, enabling personal superintelligence without cloud dependence. (The Neuron on The Neuron: AI Explained)

Generative Agents 🆕: Companies like Simile AI are building on the "Smallville" paper, creating deep models of human behavior to power more realistic and impactful simulations. (Joon Sung Park on Latent Space)

Buzz (collaboration platform) 🆕: Emerging as a "Slack for agents," this platform enables human-AI collaboration and shared agent interactions within team environments. (The Neuron on The Neuron: AI Explained)

❄️ COOLING OFF

Anthropic Claude AI watermarking: The watermark was quickly bypassed, raising concerns about its effectiveness for AI detection and the practicality of EU AI Act mandates. (AI Breakdown on AI Breakdown)

GPU-centric architectures for inference: While still dominant, traditional GPU designs are not optimally suited for the emerging premium, low-latency inference market due to memory bandwidth caps. (Sid Sheth on Eye On A.I.)

General-purpose biological AI: Virtual cell models are proving to be powerful narrow tools (like AlphaGo) but are not yet "general purpose discovery engines" or LLMs of biology. (Kristóf Szalay on The AI in Business Podcast)


The Debate

Should large AI labs push for government regulation of AI research?

🐂 The bull case: A letter from 1,100 AI employees, including those from OpenAI and Google, requested government intervention to pace AI research, arguing that rapid progress necessitates safety guardrails. Nathaniel Whittemore noted that OpenAI voluntarily paused Frontier RL training due to safety concerns, aligning with the idea that "model capabilities were outstripping the pace of safety and alignment" (Nathaniel Whittemore on The AI Daily Brief).

🐻 The bear case: This push for regulation is viewed skeptically as a potential attempt at regulatory capture by large companies, effectively "pulling up the ladder" on smaller competitors (AI Breakdown on AI Breakdown). Juan Ladano of the Cato Institute critiques the White House's AI testing regime, stating, "This approach contravenes the rule of law, risks becoming as prescriptive as a licensing regime, and most importantly, fails to fulfill its most basic objective to build trust in the population." (Juan Ladano on The AI Daily Brief).

Our read: While safety is paramount, the timing and source of calls for regulation raise valid questions about competitive motives versus genuine concern, making explicit transparency and independent oversight crucial.


The Bottom Line

The smartest AI isn't just about raw intelligence; it's about deeply understanding human behavior and localizing that power, even as the infrastructure to run it gets exponentially bigger and more complex.


📖 Want the full episode breakdowns, guest details, and listen links?

Read the Episode Guide →

Episode Guide

1. AI Breakdown — "Cracked Watermarking and Google's Chip Moves"

Runtime: 20 min | Host: AI Breakdown | Guest: AI Breakdown

For the Strategic Technologist: Unpacking the real-world implications of AI watermarking failures and Google's aggressive push into custom AI silicon.

The episode highlights how Anthropic's AI watermark was quickly broken, raising questions about detection reliability. It also covers Google's $12.2 billion investment in Marvell for AI chips, signaling hyperscaler moves into custom hardware.

"Google is going to buy up to 58 million Marvel shares for up to $12.2 billion... Broadcom, which is Google's incumbent custom chip partner, actually fell 5%."
— AI Breakdown, Host at AI Breakdown

▶ Listen · Apple Podcasts

2. The AI Daily Brief: Artificial Intelligence News and Analysis — "The AI Backlash Is Getting Stupider. But Also Smarter."

Runtime: 30 min | Host: Nathaniel Whittemore | Guest: Nathaniel Whittemore, NLW

For the Policy-Minded Executive: A nuanced take on the evolving public and political scrutiny of AI, especially around data center development.

NLW discusses the dual nature of the AI backlash, becoming both meme-driven and more productive. He analyzes Governor Shapiro's pragmatic (though strict) executive order on AI data centers and scrutiny on OpenAI/Anthropic revenue figures.

"The anti AI conversation is somehow getting dumber and more productive at the same time."
— Nathaniel Whittemore, Host of The AI Daily Brief

▶ Listen · Apple Podcasts

3. Beyond The Prompt - How to use AI in your company — "What Happens When AI Adoption Actually Works? - with Eric Porres, Chief AI Officer of Logitech"

Runtime: 65 min | Host: Henrik, Jeremy | Guest: Eric Porres (Chief AI Officer, Logitech)

For the Transformational Leader: Insights into how a large enterprise is moving from AI experimentation to full-scale, workflow-embedded adoption with concrete examples.

Eric Porres shares Logitech's journey to widespread AI integration, including mandatory "Board Advisor" AI use for executives and a focus on "deletions" as a metric for transformation. He details personal AI stacks and organizational champion programs.

"When you embed AI into the workflow of an organization or of a process, then everything follows. If you add it on at the end, then it's not as necessarily as powerful and impactful."
— Eric Porres, Chief AI Officer of Logitech

▶ Listen · Apple Podcasts

4. The AI in Business Podcast — "Determining Virtual Cell Impact for Drug Discovery - with Kristóf Szalay and Gerold Csendes of Turbine"

Runtime: 41 min | Host: Marilie Fouché | Guest: Kristóf Szalay (CTO and Co-Founder, Turbine), Gerold Csendes (Scientist, Turbine)

For the Biotech Innovator: Deep dive into the real capabilities and limitations of virtual cell models in accelerating drug discovery and the importance of human-AI alignment.

Kristóf Szalay and Gerold Csendes discuss how virtual cell models are precise tools for specific predictions in drug discovery, not general AGI. They highlight the challenge of unified evaluation frameworks and Turbine's virtual lab approach to compress R&D timelines.

"We don't really have the LLMs of biology that are generally applicable problem solvers. We have say, the Deep Blue or Alphago kind of machines, which are very good in one neural task."
— Kristóf Szalay, CTO and Co-Founder of Turbine

▶ Listen · Apple Podcasts

5. Decoder with Nilay Patel — "Welcome to the AI crisis in math"

Runtime: 41 min | Host: Nilay Patel | Guest: Robert Hart (London-based AI Reporter, The Verge), Olivia Lanes (Global Lead for Content and Education, IBM Quantum)

For the Intellectually Curious: Exploring the profound impact of AI's rapid advancements on high-level mathematics and its academic future.

Robert Hart discusses the "existential crisis" in mathematics caused by AI's sudden proficiency in abstract math, despite its struggles with basic arithmetic. The episode also touches on quantum computing cooling and the skepticism around AI's generalization capabilities.

"If a human had solved these, we would be impressed. If a human had solved all 10, we probably wouldn't believe it."
— Robert Hart, London-based AI Reporter at The Verge

▶ Listen · Apple Podcasts

6. AI Breakdown — "Revenue Growth: Anthropic and AI Investments"

Runtime: 14 min | Host: AI Breakdown | Guest: AI Breakdown

For the Investment Strategist: A quick overview of significant financial shifts and deals in the AI sector, highlighting competitive dynamics and infrastructure plays.

The host details NVIDIA's $600 billion compute deal with OpenAI, financing a massive AI data center. It also covers Anthropic's 14x Q2 revenue surge to $11.5 billion, surpassing OpenAI's run rate, and Groq's pivot to cloud with NVIDIA funding.

"Nvidia is backing a 4.25 gigawatt Ohio AI data center for OpenAI in a 600 billion dollar compute deal. OpenAI has locked in $600 billion worth of Nvidia chips through 2030."
— AI Breakdown, Host of AI Breakdown

▶ Listen · Apple Podcasts

7. Gradient Dissent: Conversations on AI — "Elon's Former Battery Chief: AI Data Centers Will Make Electricity Cheaper"

Runtime: 94 min | Host: Lukas Biewald | Guest: Drew Baglino (Founder & CEO, Heron Power), Anon (Former Battery Chief at Tesla; Founder of Heron Power), Patty Poppy (CEO, PG&E)

For the Infrastructure Innovator: A counter-intuitive deep dive into how AI data centers can drive down electricity costs and the challenges of grid modernization.

Drew Baglino, former Tesla exec, argues that properly designed AI data centers can be "grid positive" by acting as baseload customers, potentially making electricity cheaper. He discusses the stagnation of the US electricity grid and Heron Link's technology to reduce power loss.

"We can reduce the loss from grid to chip by about a factor of two. Our core first product, the Heron Link, the transformer part is 100 times smaller."
— Drew Baglino, Founder & CEO of Heron Power

▶ Listen · Apple Podcasts

8. The Neuron: AI Explained — "BONUS: AI Tool Roundup: What’s Actually Worth Trying?"

Runtime: 112 min | Host: The Neuron, Grant, Corey | Guest: Host-led discussion

For the Practical Practitioner: A comprehensive review of new AI tools, models, and techniques, focusing on real-world utility, cost, and local execution.

The hosts discuss Qwen 3.8, Unsloth Studio, and DeepSeek Harness, emphasizing "cost-per-intelligence." They highlight the shift towards local AI for "personal superintelligence" and explore AI agent harnesses like OpenClaw and Buzz for collaborative AI.

"The future of your personal superintelligence is going to live on your computer. You can run it locally, meaning you don't need a data center to run it."
— The Neuron, Host of The Neuron: AI Explained

▶ Listen · Apple Podcasts

9. The AI Daily Brief: Artificial Intelligence News and Analysis — "9 AI Techniques You Probably Haven't Tried"

Runtime: 30 min | Host: Nathaniel Whittemore | Guest: Nathaniel Whittemore, NLW

For the AI Power User: Actionable tips and lesser-known AI techniques to enhance interaction and productivity with current models and platforms.

NLW shares nine advanced AI techniques, including "multiplayer AI" with Claude Tag, using Grokbot with a virtual computer, and simple "two-word prompts" like "now what." He also touches on personalized cancer vaccines and the White House's AI safety framework.

"This pattern of use of agents that are shared across teams is going to be one of the biggest and most important trends for AI."
— NLW, Host at The AI Daily Brief

▶ Listen · Apple Podcasts

10. Latent Space: The AI Engineer Podcast — "Simulation: the new Scaling Law — Joon Sung Park, Simile AI"

Runtime: 70 min | Host: Swyx, Vibhu | Guest: Joon Sung Park (Co-founder and CEO, Simile AI)

For the Forward-Thinking Strategist: A deep dive into human behavior simulation as the "new scaling law" for AI, with implications for solving complex societal challenges.

Joon Sung Park discusses Simile AI's focus on modeling "human-like mistakes" to achieve 85% accuracy in replicating individual behaviors. He highlights the potential for simulating billions of people to address wicked problems and the emerging "scaling law" for simulations.

"We basically could replicate people's behaviors and attitudes 85% as accurately as people would replicate their own."
— Joon Sung Park, Co-founder/CEO of Simile AI

▶ Listen · Apple Podcasts

11. AI Breakdown — "ChatGPT's iMessage Hookup Explained"

Runtime: 15 min | Host: AI Breakdown | Guest: AI Breakdown

For the Digital Marketer/Privacy Advocate: Examining the dual potential and pitfalls of AI integrating with personal communication platforms like iMessage.

The host covers ChatGPT's iMessage plugin, highlighting its automation potential but also severe privacy concerns, including OpenAI's warning against persistent approval. It also touches on AI training data markets and calls for government pacing of AI research.

"All of the OpenAI tools can access your iMessage if you allow it. They can sort it, it can read it, it can draft, and it can send iMessages."
— AI Breakdown, Host at AI Breakdown

▶ Listen · Apple Podcasts

12. Eye On A.I. — "Why People Are Paying 10x More for AI - and What That Means for the Chip Market | Sid Sheth, d-Matrix"

Runtime: 51 min | Host: Craig Smith | Guest: Sid Sheth (CEO, d-Matrix)

For the Hardware Investor: An essential listen on the emerging "premium token economy" in AI inference and its disruptive impact on chip architectures.

Sid Sheth discusses how a "premium token economy" is leading users to pay 10x more for instant AI responses, a market ill-suited for traditional GPUs. He explains d-Matrix's architecture, which addresses memory bandwidth limits, and the rise of "organizational AI."

"If I want that high level of interactivity, I'll pay $20 for a million tokens. And people are willing to pay that $20. They want that interactivity."
— Sid Sheth, CEO of d-Matrix

▶ Listen · Apple Podcasts

PARTNER

Not sure where AI fits in your operations? Start with the data.

Velocity Road's AI Readiness Assessment maps your organization against 7 operational dimensions and shows exactly where AI creates ROI — in under 10 minutes.

Take the Assessment -> →

Avi Savar

Get AI & Technology in your inbox

How AI and Tech are reshaping business. Free.