The world is too loud. Read what matters.

Latent Space

Simulating real people is not a reasoning problem: frontier models get only 20–30% accuracy

Joon Sung Park's argument: treating people as trainable objects does not call for a stronger reasoning model but for a model that makes the same mistakes people make — frontier models reach only 20–30% accuracy predicting the behavior of niche populations, and Simile has pushed that to 85% using two-hour deep interviews plus RCT data.

SimulationPost-trainingBehavioral dataSynthetic panelsScaling lawMarket research

The video won't play here. Listen to the audio instead:

Worth listening. It lays out a post-training route completely unlike coding or reasoning, plus a market structure that already has paying customers but no clearly stateable TAM.

The argument · tap a timestamp to hear it

5:59

Agents fail at your errands because they do not know who you are

When he ran a time machine exercise in 2022, simulation and ‘an extremely personalized personal assistant’ came out as the top two, side by side. Joon chose simulation not merely out of a taste for science fiction, but on a judgment about ordering: you ask a model to buy your dinner and it orders a Hawaiian pizza, and you do not eat pineapple. That failure cannot be repaired with stronger execution; only a deep understanding of ‘who I am’ fixes it. His bet is that an accurate representation of a person is causally prior to automated agents. His hot take is that no personal assistant to date has actually reached that level of ambition; ChatGPT and Claude do know a good deal about their users, and their output fits better, but the key ingredient is still not in place.

— Joon Sung Park
9:50

Public models have not learned the social physics of human beings

He offers a test. If a model has to learn the underlying physics of the world it operates in — a new social physics — then you must train or post-train it. If that physics is already inside the model and you only need it to react to a situation, a prompt is enough. He does not think public models have learned a complete map of human social physics, and the reason lies in what the data is made of: web data is essentially people's ‘self-disclosed attitudinal data’, with behavioral data only scattered through it in fragments. The blank space between what people say they do online and what people actually do in the real world is what he calls humanity's dark knowledge, and it has to be collected afresh and brought into the modeling process. OpenClaw's Markdown memory shares the same intuition as the 2022 generative agents work: clever, but with a ceiling.

— Joon Sung Park
11:21

Data describing why is the scarcest and also the most valuable

The first bucket is deep interviews: asking directly, ‘tell me your life story’ — childhood memories, trauma, first love. This long-tail material is what gives ‘the person as a model’ its texture, and it works in ways that are hard to predict. The second bucket is observational behavioral data: transaction data, scrapable web data. It supplies the base statistics of behavior and is also the easiest to obtain. The third bucket he considers the most important — data that describes causal mechanism, data that describes why — and it comes mainly from RCTs: same setup, move one variable, watch how behavior changes. That kind of data is both scarce and critical, because the world is ground truth, but it happens only once. Simile runs a large number of RCTs itself, and treats ‘the decision carries a real cost’ as the dividing line between attitude and behavior: in their experiments, what you buy actually gets shipped to your house.

— Joon Sung Park
14:08

Telling a client sales will collapse in two quarters says nothing at all

‘Your sales will collapse two quarters from now’ is of no use to a decision-maker; all they can say is ‘well, that's terrible’. What they need to know is what to do now to avoid that future — which is causal mechanism, and which is the boundary between simulation and prediction. The highest form of simulation takes as input not a question but a goal: as in Foundation, ‘compress the turmoil to a thousand years’, and then it outputs every step. And those steps are frequently counterintuitive — the first one turns out to be exiling the scientist who issued the warning to Terminus. Brought down to business: an automaker wants to know how to market EVs so as to push its stock price up, and the simulation might return ‘marketing it this way will change how consumers perceive your non-EV models, and actually pull total sales down’ — a conclusion you could never arrive at by watching EV sales alone.

— Joon Sung Park
26:02

The denominator of that 85% is a person replicating themselves

The method in Generative Agent Simulations of 1000 People: recruit 1000 people, a representative sample of the US population, into a virtual lab; collect two hours of data per person (the interview script taken from the American Voices Project, plus as much behavioral data as possible); send them away for two weeks and use that data to build digital twins; then bring the real people back after two weeks for a full battery of behavioral-economics games, the Big Five, the General Social Survey, and RCTs that have been published in PNAS, and have each twin predict what its person would do. The result is replication accuracy reaching 85% of a person replicating themselves. The denominator is a person's own test-retest consistency, not absolute correctness — that is the only correct way to read this number.

— Joon Sung Park
28:24

The more rational frontier models get, the worse they simulate real people

Frontier models are moving toward being ‘hyper-rational, objective machines’: data from Mercor and Scale, annotation by professional programmers and scientists, all aimed at reasoning ability. Simile's training objective runs the other way — it wants to build a model that makes the same mistakes he makes; whatever error he commits, the model has to commit it too. That is an entirely different set of data and training objectives, and it also explains the performance gap: on the niche populations and topics customers actually care about, frontier models' behavioral prediction accuracy falls to 20–30%, and on gen pop it is roughly 50–60%, not enough to make decisions on. Another post-training route is scraping pre-registered studies from the Open Science Framework — tens of thousands of high-quality RCTs that come with stated hypotheses and protection against p-hacking — and using them to significantly improve behavioral prediction. That model is part of open science and is not being commercialized.

— Joon Sung Park
37:30

Combinatorial prompts only retrieve knowledge the model already holds

Faced with Tencent's billion persona paper (a Cartesian product of occupations and backgrounds used as prompts), he first grants the scale, then gives the core of his objection: that approach depends heavily on statistics that have already entered the model's parameters — you are only retrieving knowledge the model already has. If that route worked, simulation would already be solved, and what the market shows is not that. Real people carry a great deal of fine-grained, niche knowledge; any single piece of it looks unremarkable, but assembled it is extremely rich, and that part can only come from bespoke data collection. The same logic determines which social data they want most: Facebook, rather than LinkedIn or Twitter. On LinkedIn people are professionally guarded; on Twitter people are performing a persona; Facebook is closer to a person's base state.

— Joon Sung Park
40:04

Scale is not for significance, it is for sieving out subpopulations

He says Simile is already seeing early signs of a scaling law in simulation: the more data about people and the more compute it takes in, the more predictable the gain in simulated-person performance. But be clear about what the scale is for. Today they collect data on tens of thousands of people a week, and panel partnerships reach tens of millions of people worldwide, while actual deployments use only tens of thousands to a few hundred thousand, because what customers want is not stronger statistical significance but the ability to sieve out the subpopulation they care about. And the same person can be reused across every subsequent study, because the model is domain-agnostic: what it learns is that person's social physics, and traits such as risk preference do not change over time in the first place. Looking forward, he judges that within some years there will be simulations that cost as much to build as training a foundation model; if a single simulation could solve climate change, he would go raise that money today.

— Joon Sung Park

In their own words · checked verbatim

I don’t think the models that are out in the open have yet learned the complete mapping of social physics of humanity. … it has yet to learn the really deep behavioral nature of people, not just what people say they do online, but what they do in real life.

Joon Sung Park9:50

Simile doesn’t care about any of this. The models that we’re talking about here, what we’re trying to create are models that are as dumb as I am, right? So if I make some mistakes, the model has to make the same mistake.

Joon Sung Park28:24

The thing that we’re seeing is the early glimpse of scaling law in simulations. The more data about humans and more compute you ingest, you start to get predictive and predictable gains of the model performance in simulating it, simulating people.

Joon Sung Park40:04

I do think in the next some number of years, we will start creating simulations that will cost as much as training a foundation model. But perhaps it’s going to be so valuable to the society that it would be a no-brainer.

Joon Sung Park48:35

So market research is a $100 billion industry. … But the thing about simulation is not a tool for market research. Simulation is a tool for human decision-making.

Joon Sung Park55:01

Figures

Google Scholar citations of the generative agents (Smallville) paper7,2001:46
Accuracy of digital twins replicating a person's behavior and attitudes (denominator: the person replicating themselves)85%26:02
Data collection and follow-up design of the 1000-person study2 hours of data collection per person; full battery of tests at a follow-up 2 weeks later26:02
Frontier models' behavioral prediction accuracy on niche populations and topics customers care about20–30%28:24
Frontier models' behavioral prediction accuracy on gen pop50–60%28:24
Number of people Simile newly collects data on each week / panel partnership reachtens of thousands of people per week; panels reach tens of millions of people worldwide45:31
Size of the market research industry$100 billion55:01
Simile team size and number of co-foundersabout 60 people, 4 co-founders1:05:16

Glossary

social physics
The underlying regularities that govern how a population behaves; a model lacking it must be trained, and a model that has it only needs a prompt.
attitudinal vs. behavioral
The line is whether the decision carries a real cost — it counts as behavior only if what you bought actually ships to your house.
concept testing
A marketing-side practice: taking different messages, products or ideas out to test how a population reacts.
synthetic panel
A model-generated pool of virtual respondents that can be queried over and over, standing in for a panel of real people.
agent-based modeling
A method dating to Schelling in the 1970s: give each agent one simple rule, then watch what emerges at the group level.
Open Science Framework pre-registration
Publishing hypotheses before running the experiment so they cannot be changed after the fact; it has accumulated tens of thousands of high-quality RCTs.

How to listen

Who it's for

People building consumer products, doing growth, or running user research; investors in the AI application layer looking for a non-coding post-training route; engineers who care about the cost structure of multi-agent simulation.

Skip

58:27–1:03:03, small talk about painting and life philosophy; no added judgment.