mimi
Technical Review

Behavioral digital twins, measured against the field's own yardsticks

Two systems, one bet: simulate people from observed behavior, not prompted demographics. Every figure below is held out, carries a confidence interval, and ships with its controls and its nulls.

Prepared for David·August 13, 2026 ·SIMS synthetic audience & Luna individual twin

The thesis

Park et al. (2024) built agents for 1,052 real Americans from interviews and reproduced their survey answers at 85% of each person's own two-week test-retest consistency; demographic-only agents reached 74%. Bisbee et al. (2024) showed the failure mode of prompted-demographic personas directly: means correlate with humans while variance collapses and 32% of regression coefficients flip sign. Both our systems ground in observed behavior, and both report against that literature's metrics rather than a private scoreboard.

Twin-2K-500
0.827
of the human retest ceiling, public benchmark (SIMS)
Reply prediction
0.919
ROC-AUC on 5,378 held-out real decisions (Luna)
Corpus
1.25M
real behavioral events, 49 sources (Luna); 2,865 cohorts (SIMS)
Privacy floor
k ≥ 100
aggregation-only surfaces; maps onto the ESOMAR person/persona line
A

SIMS — synthetic audience

METHOD  person×creator co-engagement (comments) → inverse-creator-frequency weighted, log1p, L2-normalized, TruncatedSVD 128d → MiniBatchKMeans, k≥100 floor by merging undersized clusters → 675 YouTube + 2,190 v2 cohorts across three platforms → per-cohort evidence pack (topic/hashtag/sound distributions with counts, behavioral aggregates, paraphrased voice) → verbalized-distribution draws, option shares with Wilson intervals.

Twin-2K-500 (arXiv 2505.17479, CC BY 4.0) — 300 stratified personas, 23,057 items, verbalized elicitation, deepseek-v4-flash, $1.70
MeasureMacro colsPooled
Twin accuracy (MAD, 1−|pred−truth|/range)0.69560.6940
Human test-retest ceiling (same sample)0.84130.8441
Random-in-range baseline0.5658
Normalized ratio (twin ÷ retest)0.8270.822
Published best ratio (paper, full set)0.877

Coverage 99.8% of items. Leakage guard: the held-out item's own answer never enters the persona context. The remaining 0.827→0.877 gap is model tier (deepseek-flash vs frontier) and sample size, not method.

Internal validation

Evidence packs carry real signal

  • Evidence pack, held-out (disjoint labels)0.747
  • Demographic persona (arm B)0.504
  • No-evidence control (arm F)0.497
  • Best trivial rule0.527
  • Paired test (opus-5)z≈13, p≈0

Demographic persona ≈ chance (p=0.425). Attribution holds across models (opus-5 0.783, deepseek 0.733, controls ≈ chance).

U1 + U2 elicitation upgrade

Verbalized + calibrated

2,589 items, 327 cohorts, disjoint labels
  • Sampled draws (deepseek-flash)0.555
  • Verbalized + isotonic0.618
  • Positive-rate (was 0.10, truth 0.5)0.360
  • Calibration error (ECE)0.024
  • Production-model lift+6.2 pts, z=6.0
B

Luna — individual twin

METHOD  1,252,261 real events across 49 consented sources (iMessage excluded by design) → situation → action → outcome memory: a real inbound message, the relationship state known then, and whether the subject replied within seven days, with a hard outcome_known_at = ts + 7d so retrieval can never use a not-yet-observable result. Retrieval bank predates every training and test row.

Reply/ignore prediction — 5,378 held-out decisions, 5 rolling folds, pooled, clustered by relationship
ArmPR-AUCROC-AUCBrier
State (relationship + message)0.89340.90640.1274
Semantic control (similar msgs, outcomes hidden)0.8931
Outcome memory (adds what he actually did)0.90240.91900.1174
Outcome − semantic (gate)+0.009495% CI [+0.0042, +0.0155] · passes

The semantic control isolates outcome learning from mere word similarity. Gate was preregistered: both outcome−semantic and outcome−state clustered 95% lower bounds above zero.

Does a language model render that memory into accuracy? — 5,367 held-out items, local qwen3:30b, 0% call failures
ArmPR-AUCvs control (95% CI)
None (state only)0.649
Semantic control0.665
Outcome pack rendered to the model0.771+0.106 [+0.043,+0.156]

A discriminative model proves the signal exists (0.90); this proves an LLM converts the rendered pack into a large, held-out gain (+0.106), not just that similar text was present.

External cross-check

Twin-2K-500, local

  • Full persona, 25 people0.808
  • Truncated persona (16k of 94k)0.751
  • Human retest ceiling (this sample)0.799

Same public benchmark as SIMS, run on a free local model. Persona length was the handicap, as expected.

Questionnaire (honest ceiling)

At the friend baseline

  • Best construction (enneagram × vote)0.663
  • Human friend (Duncan)0.614
  • Chance0.435
  • Cross-validated router vs best-single+0.012, CI∋0

62 items cannot resolve more; the router gain's interval spans zero. An instrument limit, stated as one.

C

Why the numbers hold

Semantic control
Δ isolated
outcome vs same-retrieval-outcomes-hidden separates learning from word overlap
Label circularity
p=0.41
circular vs clean labels differ +1.2 pts, not significant
Option-order bias
debiased
permuted per draw (Dominguez-Olmedo): removes the "A" artifact
Leakage
time-gated
outcome_known_at = ts+7d; bank predates train and test rows
Retest normalization
÷ 0.80–0.84
accuracy stated against human self-consistency, the real ceiling
Intervals
clustered
paired bootstrap over relationships; claim only when CI clears 0
Reported, not buried

Nulls and limits

For a technical reviewer these are the credibility. Each was measured on our own data and published as a negative.

  • Engagement-rate forecasting loses to a popularity rule, 0.56 vs 0.64 balanced. The evidence arm does not yet win this task.not won
  • The 62-item questionnaire is instrument-capped. A cross-validated router adds +0.012 (CI [−0.066,+0.094]) over best-single; the 0.807 oracle is hindsight that evaporates out of sample.ceiling
  • Fine-tuning does not beat prompting on judgment. 7B pure-voice collapses to 0.437 (chance); 32B reaches 0.555 and loses 6 points on decisions.dead end
  • The twin loses on professional decisions, 0.583 vs a 0.625 generic prior (n=60). Personal decisions flip positive but under-powered, +0.014 (n≈40).mixed
  • Trend adoption is under-powered in our window: +6 to +9 points directional, p=0.054–0.23, on only ~20–40 real trends in 13 weeks.thin
  • Model tier matters. The cheap model's draw collapses to positive-rate 0.10 where truth is 0.5. Verbalized elicitation plus isotonic fixes it.fixed
D

Where this sits

AxisThe fieldUs
Provenancemost rivals prompt demographics, the end the record discreditsobserved co-engagement and real message history
Validationmost LLM-ABMs validate on stylized facts (Larooij–Törnberg)preregistration, controls, paired CIs, published nulls
Honest ceilingSocieties.io: 86% vs a stated 91% human ceilingTwin-2K 0.827 of retest, stated as a ratio
PrivacySimile carries per-individual consent; others scrape tracesk≥100 aggregation, ESOMAR-clean, paraphrased exemplars

Reference points: Simile ~$300M raised, $2B valuation (Feb 2026), CVS 100k patient twins, no public accuracy. SimBench public ceiling 40.8/100 for prompt-only frontier models. Twin-2K-500 (2,058 people × 500 questions) is the emerging public benchmark.

E

Questions you'll ask

Isn't the cohort signal circular, since the labels come from the same engagement graph?
We ran the label-circularity control directly: circular vs clean labels differ by +1.2 points, paired p=0.41, not significant. The evidence-pack result (0.747) holds on a disjoint-label held-out set, and a batched-leak we found early (arm F 0.699) was killed by cohort-mixed batches, dropping it back to chance (0.46).
Does the twin just match means while variance collapses, like the ANES critique?
That is the exact failure we test for. We report the variance ratio (sim/human) and cross-item correlation structure, not only marginals. The single-model draw does collapse on cheap models (positive-rate 0.10 vs truth 0.5); verbalized elicitation restores it to 0.36 and isotonic calibration brings ECE to 0.024. On the individual side, the reply model is scored by PR-AUC and Brier, which penalize a collapsed, over-confident distribution.
How do you know the benchmark isn't leaked or memorized?
Twin-2K-500 answers are held out per person: the target item's own answer never enters the persona context, and we reproduce the dataset's human test-retest ceiling from the data itself (0.80–0.84), which a memorizing model could not do. Luna's outcome memory is hard-gated by outcome_known_at = situation + 7 days, and the retrieval bank predates every training and test row.
Why should I trust a +0.009 discriminative gain?
Two reasons. It clears a preregistered, clustered 95% interval [+0.0042, +0.0155] on 5,378 real decisions, and the rendered version of the same memory produces a far larger LLM gain, +0.106 PR-AUC [+0.043, +0.156], with 0% call failures. The small discriminative number is the conservative floor; the renderer number is the product-relevant one.
The questionnaire only ties a friend. Isn't that weak?
On that instrument, yes, and we say so. A cross-validated router adds nothing over best-single (+0.012, CI spans zero), so 62 items cannot resolve a larger gap. The strength is behavioral: predicting 5,378 real future decisions at 0.90 AUC is something a friend cannot do. The quiz is a weak instrument, not a weak model.
What about privacy and consent, given this is real people's data?
SIMS never exposes an individual: a k≥100 floor at construction, no per-individual rows in any output, paraphrased exemplars, which maps directly onto the ESOMAR 2025 person/persona distinction. Luna is a single consented subject with iMessage excluded by design, stored owner-only on an encrypted machine, and it never sends a message on anyone's behalf.
Can this run cheaply, or does it need a frontier model?
The 0.827 Twin-2K result used deepseek-flash for $1.70 across 1,199 calls. Verbalized elicitation matches frontier-model quality at roughly 1/30 the cost on this task class. Frontier models add accuracy on high-nuance items and are where the 0.827→0.877 gap likely closes.
F

What's next