Behavioral digital twins, measured against the field's own yardsticks
Two systems, one bet: simulate people from observed behavior,
not prompted demographics. Every figure below is held out, carries a
confidence interval, and ships with its controls and its nulls.
Prepared for David·August 13, 2026·SIMS synthetic audience & Luna individual twin
The thesis
Park et al. (2024) built agents for 1,052 real Americans from interviews
and reproduced their survey answers at 85% of each person's own two-week
test-retest consistency; demographic-only agents reached 74%. Bisbee et al.
(2024) showed the failure mode of prompted-demographic personas directly:
means correlate with humans while variance collapses and 32% of regression
coefficients flip sign. Both our systems ground in observed behavior, and
both report against that literature's metrics rather than a private
scoreboard.
Twin-2K-500
0.827
of the human retest ceiling, public benchmark (SIMS)
Reply prediction
0.919
ROC-AUC on 5,378 held-out real decisions (Luna)
Corpus
1.25M
real behavioral events, 49 sources (Luna); 2,865 cohorts (SIMS)
Privacy floor
k ≥ 100
aggregation-only surfaces; maps onto the ESOMAR person/persona line
A
SIMS — synthetic audience
METHOD person×creator co-engagement (comments) → inverse-creator-frequency
weighted, log1p, L2-normalized, TruncatedSVD 128d → MiniBatchKMeans, k≥100
floor by merging undersized clusters → 675 YouTube + 2,190 v2 cohorts across
three platforms → per-cohort evidence pack (topic/hashtag/sound distributions
with counts, behavioral aggregates, paraphrased voice) → verbalized-distribution
draws, option shares with Wilson intervals.
Twin-2K-500 (arXiv 2505.17479, CC BY 4.0) — 300 stratified personas, 23,057 items, verbalized elicitation, deepseek-v4-flash, $1.70
Measure
Macro cols
Pooled
Twin accuracy (MAD, 1−|pred−truth|/range)
0.6956
0.6940
Human test-retest ceiling (same sample)
0.8413
0.8441
Random-in-range baseline
0.5658
—
Normalized ratio (twin ÷ retest)
0.827
0.822
Published best ratio (paper, full set)
0.877
—
Coverage 99.8% of items. Leakage guard: the held-out item's own
answer never enters the persona context. The remaining 0.827→0.877 gap is
model tier (deepseek-flash vs frontier) and sample size, not method.
Internal validation
Evidence packs carry real signal
Evidence pack, held-out (disjoint labels)0.747
Demographic persona (arm B)0.504
No-evidence control (arm F)0.497
Best trivial rule0.527
Paired test (opus-5)z≈13, p≈0
Demographic persona ≈ chance
(p=0.425). Attribution holds across models (opus-5 0.783, deepseek 0.733,
controls ≈ chance).
U1 + U2 elicitation upgrade
Verbalized + calibrated
2,589 items, 327 cohorts, disjoint labels
Sampled draws (deepseek-flash)0.555
Verbalized + isotonic0.618
Positive-rate (was 0.10, truth 0.5)0.360
Calibration error (ECE)0.024
Production-model lift+6.2 pts, z=6.0
B
Luna — individual twin
METHOD 1,252,261 real events across 49 consented sources
(iMessage excluded by design) → situation → action → outcome memory: a real
inbound message, the relationship state known then, and whether the subject
replied within seven days, with a hard outcome_known_at = ts + 7d
so retrieval can never use a not-yet-observable result. Retrieval bank
predates every training and test row.
Reply/ignore prediction — 5,378 held-out decisions, 5 rolling folds, pooled, clustered by relationship
Arm
PR-AUC
ROC-AUC
Brier
State (relationship + message)
0.8934
0.9064
0.1274
Semantic control (similar msgs, outcomes hidden)
0.8931
—
—
Outcome memory (adds what he actually did)
0.9024
0.9190
0.1174
Outcome − semantic (gate)
+0.0094
95% CI [+0.0042, +0.0155] · passes
The semantic control isolates outcome learning from mere word
similarity. Gate was preregistered: both outcome−semantic and outcome−state
clustered 95% lower bounds above zero.
Does a language model render that memory into accuracy? — 5,367 held-out items, local qwen3:30b, 0% call failures
Arm
PR-AUC
vs control (95% CI)
None (state only)
0.649
—
Semantic control
0.665
—
Outcome pack rendered to the model
0.771
+0.106 [+0.043,+0.156]
A discriminative model proves the signal exists (0.90); this
proves an LLM converts the rendered pack into a large, held-out gain
(+0.106), not just that similar text was present.
External cross-check
Twin-2K-500, local
Full persona, 25 people0.808
Truncated persona (16k of 94k)0.751
Human retest ceiling (this sample)0.799
Same public benchmark as SIMS,
run on a free local model. Persona length was the handicap, as expected.
Questionnaire (honest ceiling)
At the friend baseline
Best construction (enneagram × vote)0.663
Human friend (Duncan)0.614
Chance0.435
Cross-validated router vs best-single+0.012, CI∋0
62 items cannot resolve more; the
router gain's interval spans zero. An instrument limit, stated as one.
C
Why the numbers hold
Semantic control
Δ isolated
outcome vs same-retrieval-outcomes-hidden separates learning from word overlap
Label circularity
p=0.41
circular vs clean labels differ +1.2 pts, not significant
Option-order bias
debiased
permuted per draw (Dominguez-Olmedo): removes the "A" artifact
Leakage
time-gated
outcome_known_at = ts+7d; bank predates train and test rows
Retest normalization
÷ 0.80–0.84
accuracy stated against human self-consistency, the real ceiling
Intervals
clustered
paired bootstrap over relationships; claim only when CI clears 0
Reported, not buried
Nulls and limits
For a technical reviewer these
are the credibility. Each was measured on our own data and published as a
negative.
Engagement-rate forecasting loses to a popularity rule, 0.56 vs 0.64 balanced. The evidence arm does not yet win this task.not won
The 62-item questionnaire is instrument-capped. A cross-validated router adds +0.012 (CI [−0.066,+0.094]) over best-single; the 0.807 oracle is hindsight that evaporates out of sample.ceiling
Fine-tuning does not beat prompting on judgment. 7B pure-voice collapses to 0.437 (chance); 32B reaches 0.555 and loses 6 points on decisions.dead end
The twin loses on professional decisions, 0.583 vs a 0.625 generic prior (n=60). Personal decisions flip positive but under-powered, +0.014 (n≈40).mixed
Trend adoption is under-powered in our window: +6 to +9 points directional, p=0.054–0.23, on only ~20–40 real trends in 13 weeks.thin
Model tier matters. The cheap model's draw collapses to positive-rate 0.10 where truth is 0.5. Verbalized elicitation plus isotonic fixes it.fixed
D
Where this sits
Axis
The field
Us
Provenance
most rivals prompt demographics, the end the record discredits
observed co-engagement and real message history
Validation
most LLM-ABMs validate on stylized facts (Larooij–Törnberg)
preregistration, controls, paired CIs, published nulls
Honest ceiling
Societies.io: 86% vs a stated 91% human ceiling
Twin-2K 0.827 of retest, stated as a ratio
Privacy
Simile carries per-individual consent; others scrape traces
Reference points: Simile ~$300M raised,
$2B valuation (Feb 2026), CVS 100k patient twins, no public accuracy.
SimBench public ceiling 40.8/100 for prompt-only frontier models. Twin-2K-500
(2,058 people × 500 questions) is the emerging public benchmark.
E
Questions you'll ask
Isn't the cohort signal circular, since the labels come from the same engagement graph?
We ran the label-circularity control directly: circular vs
clean labels differ by +1.2 points, paired p=0.41, not significant. The
evidence-pack result (0.747) holds on a disjoint-label held-out set, and a
batched-leak we found early (arm F 0.699) was killed by cohort-mixed batches,
dropping it back to chance (0.46).
Does the twin just match means while variance collapses, like the ANES critique?
That is the exact failure we test for. We report the
variance ratio (sim/human) and cross-item correlation structure, not
only marginals. The single-model draw does collapse on cheap models
(positive-rate 0.10 vs truth 0.5); verbalized elicitation restores it to
0.36 and isotonic calibration brings ECE to 0.024. On the individual side,
the reply model is scored by PR-AUC and Brier, which penalize a collapsed,
over-confident distribution.
How do you know the benchmark isn't leaked or memorized?
Twin-2K-500 answers are held out per person: the target item's
own answer never enters the persona context, and we reproduce the dataset's
human test-retest ceiling from the data itself (0.80–0.84), which a
memorizing model could not do. Luna's outcome memory is hard-gated by
outcome_known_at = situation + 7 days, and the retrieval bank predates
every training and test row.
Why should I trust a +0.009 discriminative gain?
Two reasons. It clears a preregistered, clustered 95%
interval [+0.0042, +0.0155] on 5,378 real decisions, and the rendered
version of the same memory produces a far larger LLM gain, +0.106 PR-AUC
[+0.043, +0.156], with 0% call failures. The small discriminative number
is the conservative floor; the renderer number is the product-relevant one.
The questionnaire only ties a friend. Isn't that weak?
On that instrument, yes, and we say so. A cross-validated
router adds nothing over best-single (+0.012, CI spans zero), so 62 items
cannot resolve a larger gap. The strength is behavioral: predicting 5,378
real future decisions at 0.90 AUC is something a friend cannot do. The quiz
is a weak instrument, not a weak model.
What about privacy and consent, given this is real people's data?
SIMS never exposes an individual: a k≥100 floor at
construction, no per-individual rows in any output, paraphrased exemplars,
which maps directly onto the ESOMAR 2025 person/persona distinction. Luna is
a single consented subject with iMessage excluded by design, stored owner-only
on an encrypted machine, and it never sends a message on anyone's behalf.
Can this run cheaply, or does it need a frontier model?
The 0.827 Twin-2K result used deepseek-flash for $1.70
across 1,199 calls. Verbalized elicitation matches frontier-model quality at
roughly 1/30 the cost on this task class. Frontier models add accuracy on
high-nuance items and are where the 0.827→0.877 gap likely closes.
F
What's next
External-benchmark expansion — OpinionQA, SimBench, ANES, GSS on the shipped verbalized engine, with distributional metrics staged.queued
Close the Twin-2K gap (0.827→0.877) — full persona plus a frontier model; the truncated and local passes already showed persona length was the handicap.ready
Scale and power — all 2,058 Twin-2K personas, larger question banks so effects clear the noise floor.ready
Rate forecasting and propagation — the one open scientific loss; thresholded Monte-Carlo propagation and adoption calibration are shipped, backtest pending.open