June 2026
•
10 min read
Proof, Not Promises: A New Bar for Behavioral AI



By Sophia Lanuzo, Melrose Tia, and Jason Albia

Article
Research and Development
This article surveys the competitive landscape in the AI-driven behavioral simulation space. The central finding is straightforward: the industry converges on internal accuracy claims while avoiding the harder question of commercial validity, and no competitor has yet published a blind backtesting exercise against independently sourced, real-world campaign performance data. Predikta's backtesting study with AdSpark is the first to close that gap, combining a strong scientific foundation with an externally validated commercial proof point built on real ad metrics. That is not a marginal differentiator: it is a different category of evidence entirely, and it sets a standard the rest of the industry will now need to respond to.
01 · The Validation Problem
What the Industry Is Not Measuring
The synthetic audience and behavioral simulation market has moved from research curiosity to contested commercial territory in under two years. From well-funded U.S. startups like Aaru to specialized European platforms like Lakmoos, vendors increasingly position validation and predictive accuracy as primary differentiation across the competitive landscape.
Yet in practice, most claims rest on simulation fidelity: how closely does the simulated response match a known human response on a held-out dataset?
That is a meaningful metric, and an important first step. It demonstrates that simulation can reproduce patterns observed in existing human data.
But replication alone is not enough. The real test is whether simulation signals translate into real-world outcomes: does a higher simulation score translate to a better-performing ad in the real world, against real audiences, on real platforms?
No behavioral AI simulation platform in this space, to our knowledge, has moved beyond response replication. Predikta is the first to do so, having conducted a blind backtest against independently sourced historical campaign data from AdSpark, measuring whether higher simulation scores predict actual Click-Through Rates (CTR) and Cost Per Click (CPC) outcomes.
A simulation score that predicts human responses but cannot predict market outcomes has not yet proven its commercial value. This is the gap the Predikta x AdSpark backtesting study was designed to close. Its findings carry implications not only for Predikta, but for the standard of evidence the industry should be required to meet.
02 · Predikta's Evidence Architecture
Two Layers, Two Different Questions
Most platforms in this space have a single layer of evidence: internal benchmarking against their own data. Predikta has deliberately built two distinct and complementary validation layers, each answering a question the other cannot.
L1 — Scientific Foundation: Rigorous Internal Validation
Published as Sentiment Simulation using Generative AI Agents (Tia et al., arXiv:2505.22125, May 2025), the study grounded agents in a nationally representative survey of 2,485 Filipinos. The agents achieved 92% alignment with original survey responses and 81–86% accuracy in predicting ground-truth human sentiment, with high stability across repeated trials (±0.2–0.5% SD) and negligible sensitivity to framing variation (p = 0.9676, Cohen’s d = 0.02). Platform benchmarks: 88% simulation accuracy at national level and 96% behavioral fidelity in psychographic representation. Taken together, these results answer a foundational question: the simulation faithfully represents how real people think and feel.
L2 — Commercial Validation: Blind External Backtesting with AdSpark
Conducted in partnership with AdSpark and supported by 917Ventures and Brave Connective, the study evaluated real historical campaigns across multiple industries. Predikta scored ads in blind simulations, with results benchmarked against independently sourced KPI data across nearly 20,000 ad pairs. Predikta achieved a 20-point advantage over random selection in identifying the stronger-performing ad, while ads in the top 20% of Predikta scores delivered 3.4× higher CTRs and 2.6× lower CPCs than lower-scoring ads. That answers the harder question: simulated audience responses can reliably identify winning campaigns before they go to market.
The significance of the two-layer structure is not just in the numbers each layer produces: it is the logical completeness. Layer 1 establishes that the behavioral simulation model is accurate. Layer 2 establishes that behavioral accuracy has predictive value in the market. Both questions must be answered before a marketing team can have justified confidence in using simulation scores to make pre-launch decisions. Predikta is the only platform in this competitive roundup that has answered both.

03 · The Backtesting Study
Design, Results, and Implications
The backtesting study was designed specifically to avoid the methodological shortcuts that make internal validation exercises easy to dismiss. Three methodical design choices defined its credibility.
Independent data source. Campaign data came directly from AdSpark's own historical records: real campaigns run for real clients, with real performance outcomes already recorded. The dataset covered multiple industries such as telecommunications, IT and Software Services, banking/finance, and others. Predikta had no involvement in selecting which campaigns were included in this study.
Blind simulation. Predikta's behavioral scores were generated without access to the actual campaign performance outcomes.
Real commercial metrics. Performance was evaluated against CTR and CPC, widely used measures of advertising and media effectiveness, rather than proprietary engagement proxies or internal scoring.
PREDIKTA × ADSPARK BACKTESTING STUDY — KEY RESULTS
+20%
Advantage over random chance in head-to-head ad selection
3.4×
Higher CTR for top-20% Predikta scores vs. lower-scoring ads
2.6×
Lower CPC for top-20% Predikta scores vs. lower-scoring ads
Pre-launch behavioral scores predicted real-world campaign performance across nearly 20,000 ad pairs, across industries.
Across nearly 20,000 head-to-head ad comparisons spanning multiple industries, Predikta correctly identified the stronger-performing ad approximately 70% of the time for conversion ads. Ad choices were determined entirely by simulated audience responses, then validated against real campaign data using CTR and CPC. Where random guessing would be correct half the time, Predikta achieved a 20-percentage-point advantage before a single ad was deployed.
That result is not incidental. It indicates that simulated audience behavior is meaningfully connected to real market outcomes, and that what Predikta measures reflects something true about how people actually respond to advertising.
Picking the stronger-performing ad is one thing. Understanding how much better the higher-ranked ads perform is another. Ads in the top 20% of Predikta scores delivered 3.4 times the CTR of lower-scoring ads at a cost of 2.6 times less per click. At typical campaign spending levels, those differences are commercially significant before a single peso is allocated to a live campaign.
The study also identified where the advantage is strongest: Awareness and Conversion campaigns, the two objective types where creative quality has the most direct influence on outcomes. In these campaigns, success typically depends less on platform optimization or algorithm, and more on how audiences think, feel, and respond to the ad message itself. That is precisely the layer Predikta's psychographic simulation is designed to model.
"Testing Predikta against real-world campaign outcomes has been a long-standing priority for our team. Moving beyond replication of survey responses to evaluating Predikta's output against observed campaign performance represents a meaningful development step. Opportunities to conduct this type of commercial validation are rare, and seeing the relationships hold up against observed campaign performance data reinforced confidence in the methodological foundation behind Predikta."
— Sophia Lanuzo, Lead R&D
04 · Research Frontier
Academic Work Shaping the Science
The backtesting study results do not stand alone. The broader academic literature on psychographically grounded agent simulation is converging on findings that reinforce precisely what the AdSpark study demonstrates in commercial terms.
Our Foundational Paper (arXiv:2505.22125). Our arXiv paper established the behavioral architecture that the backtesting study later tested in a real commercial setting. The paper instantiated agents from survey data covering personality, values, beliefs, and socio-political attitudes. Critically, it showed that agents could maintain stable, psychographically anchored outputs regardless of how scenarios are framed — the behavioral property that makes real-world commercial prediction possible. A model that is highly sensitive to stimuli framing cannot produce reliable pre-launch predictions. Predikta's foundational results suggest a more stable behavioral signal. The AdSpark backtesting results validate that stability in practice.
Adjacent Research Worth Knowing
LLM Agents Grounded in Self-Reports (arXiv:2411.10109). Using 1,052 Americans, this study showed that agents built from qualitative interview and structured survey data reached 82–86% agreement on held-out social attitude items, compared with 74% for demographics-only agents. This independently supports Predikta's design choice: grounding agents in rich psychographic data is measurably more effective than demographic profiling alone. While not conducted in an advertising context, the performance gap suggests practical relevance for applied behavioral prediction tasks.
AgentSociety (arXiv:2502.08691). A large-scale social simulation framework that uses LLM-driven agents integrating Maslow's hierarchy of needs, different emotional states, motivations, and cognitive processes to model human-like behavior. This signals where the frontier is heading: motivational agent architectures that model dynamic needs and affective states beyond static profile-based simulation.
Polarization Simulation (Springer, 2025). A 100-agent study grounded in psychometric and demographic data from Serbian social media users demonstrates that psychographically grounded agents can produce coherent ideological clustering and differentiated behavior across staged social media interactions. This provides adjacent support for the stability property Predikta's paper documents and the backtesting study depends on.
Open challenge: A 2025 meta-review notes LLMs "struggle significantly in scenarios that more accurately reflect real-world conditions", particularly on challenges around emotional nuance, cultural context, and group dynamics. Predikta's Philippines-specific, survey-grounded dataset directly responds to the cultural-context critique. The AdSpark backtesting results are evidence that the response is working. That evidence is further contextualized by how Predikta's validation architecture compares to every other platform currently operating in this space.

05 · Product Landscape
Who Is Doing What
The synthetic audience and behavioral simulation space is increasingly converging around the same question: not whether synthetic respondents can be generated, but how their outputs are validated. Well-funded companies, specialized synthetic research platforms, and research firms are all building validation stories through benchmark studies, partner datasets, client case studies, or performance-linked models. Most visible players, however, remain Western or globally generalized in their data orientation. To our knowledge, no other platform has published a comparable validation architecture in the Philippines or broader Southeast Asia — one that integrates local survey grounding, population-scale modeling, and externally verified commercial outcomes.
Aaru
The highest-profile funded competitor. Founded in March 2024 and reported to have raised a Series A at a $1B headline valuation (Redpoint Ventures). Backed by Accenture Ventures with anchor partnerships at IPG and EY. Uses multi-agent AI across public and proprietary data. Best-known validation: correctly predicted the New York Democratic primary. Validation approach is event-prediction based — not tied to ad creative performance metrics. No peer-reviewed paper, no commercial backtest against CTR, CPC, or campaign conversion data published.
$1B valuation Multi-agent Accenture-backed US-primary
Atypica.AI
Positioned as an AI consumer research and product strategy platform. Claims 300,000+ persona library built from in-depth interviews, rapid time-to-insight, and 100× cost efficiency versus traditional research agencies. Validation rests on internal benchmarks and client testimonials. No peer-reviewed accuracy figures or published external backtest against real-world ad performance data.
300K+ personas Global panels Strategy focus
Evidenza
The most direct product-use-case competitor to Predikta's Campaign Simulation Lab. Creates audience-specific synthetic personas to validate brand messaging, advertising creatives, and copy variations across demographic and psychographic audience segments. No published peer-reviewed accuracy metrics or commercial backtesting results against live campaign performance data found.
Ad creative testing Marketing validation Psychographic segments Synthetic audiences
Synthetic Users
General-purpose synthetic research participants for interviews, concept tests, surveys, and usability studies. Multi-agent architecture coordinating multiple LLMs. SOC 2 compliant. Strong in UX and product research; less focused on campaign sentiment simulation or marketing performance forecasting. Validation emphasises qualitative realism over quantitative accuracy against real-world outcomes.
Multi-LLM SOC 2 compliant UX / product focus
Lakmoos
Differentiates on architecture: neuro-symbolic AI (neural networks combined with symbolic reasoning) rather than pure LLMs. Reports 98%+ similarity scores across 20 client benchmark studies in 2025, with a Belkin case study. Benchmark figures are client-sourced, not peer-reviewed or independently verified against real campaign performance metrics.
Neuro-symbolic AI 98% similarity High auditability
Ditto / Deepsona
A cluster of specialized platforms with distinct output formats focused on concept validation, pricing, messaging, and go-to-market research. Ditto emphasizes synthetic persona panels and fast qualitative-to-quantitative studies, while Deepsona positions itself as a predictive marketing simulation platform for campaigns, pricing, products, and ideas. Both operate primarily in Western markets with no published culturally specific models for Southeast Asia.
Narrative outputs Pricing strategy Western market research

06 · Competitive Validation Matrix
A Standard the Field Must Now Meet
The table below summarizes the validation landscape across the platforms evaluated, focusing on five dimensions that matter when deciding whether simulation scores can be trusted for real media budget decisions.
Public Scientific Evidence
Predikta — Yes. arXiv:2505.22125; 92% profile alignment; 88% national-level simulation accuracy.
Others — No or partial; generally no platform-specific published scientific evidence.
External Commercial Backtest
Predikta — Yes. AdSpark study; ~20,000 ad pairs; blind design.
Others — No or partial; some validation claims or use cases, but no published ad-performance backtests.
Independent Data Source
Predikta — Yes. AdSpark historical campaigns, independently sourced.
Others — Often internal benchmarks, partner datasets, or non-independent sources.
Real Ad Metrics (CTR / CPC)
Predikta — Yes. 3.4× CTR uplift; 2.6× CPC reduction; ~70% head-to-head accuracy.
Others — Public CTR/CPC validation metrics generally not disclosed.
Cultural / Market Grounding
Predikta — Philippines-specific; 2,485 real survey respondents; 70M-population model.
Others — Typically global, Western, or European; audience-based models without comparable survey-grounded population modeling.
At a high level, Predikta is the only platform in this comparison with publicly documented evidence across all five dimensions. Other platforms may provide validation claims, partner integrations, or market-specific capabilities, but publicly available evidence is generally partial, unpublished, or not directly tied to campaign-level advertising outcomes.
Sources & References
[1] Tia, M., Lanuzo, J.S., Baltazar, L.R., Lopez-Relente, M.J., Quiñones, D.M., Albia, J. (2025). Sentiment Simulation using Generative AI Agents. arXiv:2505.22125 [cs.MA]. Netopia AI / UP Diliman / UP Los Baños.
[2] Netopia AI and AdSpark (May 2026). Predikta × AdSpark Backtesting Study — Press Release. Key figures: 3.4× CTR, 2.6× CPC reduction, ~70% head-to-head accuracy, ~20,000 ad pairs, an estimated 12M additional clicks and up to ₱50M in media value.
[3] Greenbook GRIT Report (2025). AI adoption in qualitative research — 72% figure, cited via Perspective AI blog, April 2026.
[4] PyMC Labs (2026). Synthetic Consumers: A Practical Guide. pymc-labs.com. Synthetic data 50% projection by 2027.
[5] TechCrunch (Dec 2025). AI synthetic research startup Aaru raised a Series A at a $1B headline valuation.
[6] Accenture Newsroom (Mar 2025). Accenture Invests in and Collaborates with AI-Powered Agentic Prediction Engine Aaru.
[7] Marketing Dive (Aug 2025). IPG partners with Aaru for AI-powered consumer simulations.
[8] AiMultiple (Mar 2026). Synthetic Users Explained: Top 7 AI User Research Tools.
[9] Ditto (Feb 2026). Synthetic Research Platforms: The 2026 Market Map. askditto.io.
[10] arXiv:2411.10109 (Nov 2024). LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals.
[11] arXiv:2502.08691 (Feb 2025). AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents.
[12] Springer (2025). Agent-Based Simulation of Politicized Topics Using Large Language Models (RecSysLLMsP).
[13] Atypica.AI Blog (Dec 2025). 10 Best AI Market Research Tools in 2025.
Ready to know before
you launch?
Find out what works before you launch your next campaign.


