Spyglasses Research

A synthetic panel matched to human sub-intent reproduces the consideration set — but not the share of voice.

Spyglasses Research · August 6, 2026 · spec & code

Key findings

Why we asked

Four days ago we published a bet. Having found that brand-anchored prompt panels measure a brand’s home field rather than its market, we argued the fix was specific and testable:

If a scenario generator’s output is stratified to the human panel’s sub-intent profile — budgets, recipients, use cases, form factors, in human proportions — the data says it should mirror the human panel at both the response level and the share level.

That prediction came from a post-hoc pilot on cells as small as 5 prompts, which is exactly the kind of result that evaporates under a real test. So we pre-registered the real test, froze the panels before collecting a single answer, and committed to publishing whatever came back.

The stakes are not academic. Every AI-visibility product, ours included, reports numbers computed over a panel of prompts that no human wrote. David McSweeney’s 31 methodology questions asks vendors to show that their tracked prompts resemble the population’s (Q3) and that volume is not being substituted for representativeness (Q7). SparkToro’s research makes the same point from the data side. If a synthetic panel can be made to stand in for a human one, that is a concrete answer. If it cannot, the honest move is to say what it can and cannot stand in for.

Three panels ran side by side on the same platform for the same five days:

The stratified panel was generated once, by regenerate-until-valid, and never hand-edited — the same one-shot process a customer would get.

Methods in brief

1,230 runs evaluated in this study: 238 headphone-panel prompts × 5 daily waves plus the 40-prompt control × 1, collected 2–6 August 2026 through DataForSEO’s LLM scraper (en-US, web search forced, model gpt-5-5 throughout — not the logged-in consumer product). From each answer we extracted recommended brands using a frozen alias lexicon, cited source domains, and grounding-query tokens. Total collection cost: $2.95.

Four pre-registered hypotheses, all with 90% prompt-cluster bootstrap intervals and TOST-style equivalence logic, frozen at 80e6c09 before wave one:

We also registered five predictions in public. Three held; the central one did not.

#Registered predictionOutcome
1The stratified panel passes H1Confirmed — −0.001 on brands
2The stratified panel passes H2Falsified — 0.089 against a 0.05 band
3The unstratified panel passes H1 on brands and sourcesSplit — sources yes, brands landed just outside
4The unstratified panel fails H2Confirmed — 0.137
5Matched sub-intents come back equivalentConfirmed — travel −0.009, music −0.038

The positive-control gate passed (same-intent pairs cite far more similar sources than cross-intent pairs, Δ = 0.306 [0.274, 0.338], permutation p = 0.0002) and the placebo split was null (−0.003 [−0.008, 0.005]). The human baseline replicated a third time: between-prompt brand overlap 0.517 here versus 0.528 and 0.537 in the two prior studies, and human brand shares tracked the previous study within about 3 points. A manipulation check confirmed the stratified panel hit all twelve of its target sub-intent cells exactly.

Result 1: the synthetic prompts are indistinguishable from human ones

This is the cleanest exchangeability result in the program, and it beat our own prediction.

ArtifactHuman × humanHuman × stratifiedDifference [90% CI]
Recommended brands0.5170.516−0.001 [−0.038, +0.037]
Cited sources0.3080.306−0.001 [−0.030, +0.026]
Grounding queries0.2880.233−0.054 [−0.082, −0.028]

Pick a synthetic prompt and a human prompt at random and compare the two answers: they agree exactly as much as two differently-worded human prompts agree with each other. The equivalence band was 0.10; the estimate is 0.001.

Restricting to pairs where both prompts carry the same sub-intent tightens it further — travel −0.009 [−0.051, +0.030] across 46 prompts, music −0.038 [−0.088, +0.009] across 41. The pilot that motivated this study replicated at eight times the sample size.

Two honest qualifications. The grounding-query contrast sits inside the band but its interval excludes zero, so the panel searches slightly differently even when it answers the same. And the unstratified panel’s brand contrast came back at −0.055 [−0.106, −0.006] — the point estimate is close to the −0.041 we measured last month, but the interval spills just past our band, so by the rule we froze it counts as a real difference rather than equivalence. It is a borderline interval landing on the other side of a threshold, not a reversal.

Result 2: the share vector still misses

Horizontal bar chart of human-panel brand mention shares with overlaid markers for both synthetic panels. Sony's human bar reaches 89% with the stratified panel's marker at 83%; Bose 81% versus 80%; Apple's human bar reaches 46% while both synthetic markers sit near 25–28%.

Share of answers mentioning each brand: human-panel bars (with 90% intervals) against each synthetic panel’s markers, pooled over the five waves evaluated in this study.

Averaged across the six brands humans mention most, the stratified panel is 8.9 points off [7.1, 12.5] — better than the unstratified panel’s 13.7 and better than the previous study’s neutral panel at 11.2, but well outside the 5-point band we committed to. The prediction failed.

BrandHumanStratifiedDifference
Sony89.0%83.3%−5.7
Bose80.8%80.4%−0.5
Sennheiser76.6%66.5%−10.1
Anker69.8%78.2%+8.4
Apple46.3%25.1%−21.2
JBL23.4%15.6%−7.7

Stratification bought real ground — Bose lands within half a point, and Sennheiser’s 28-point miss on the unstratified panel shrinks to 10 — but Apple alone accounts for 40% of the total deviation, and the panel under-mentions five of six brands while over-mentioning Anker.

The equivalence bound is the claim: we could have detected agreement within 5 share points, and the best panel missed by 8.9.

Result 3: what the miss actually is

Everything from here is post-hoc and labelled as such. It is also the part that changed how we read the pre-registered result.

A share figure pools a yes/no per answer, which means “off by 8.9 points” can describe two very different failures. Either the synthetic panel surfaces different brands, or it surfaces the same brands at different rates. Splitting the statistic answers it.

Two side-by-side horizontal bar charts for the travel sub-intent. The left panel, showing the share of prompts surfacing each brand at least once in five runs, has human and synthetic bars of nearly equal length for all six brands. The right panel, showing the share surfacing each brand in all five runs, shows the bars diverging sharply, with Apple at 9.5% for humans and 0% for the synthetic panel.

The same six brands, split two ways, within the travel sub-intent evaluated in this study. Left: does this prompt surface the brand at all? Right: does it surface the brand every time?

Within the travel sub-intent, the panels agree on which brands belong. Counting the prompts that surface a brand at least once across their five runs: Sony 99% of human prompts and 96% of synthetic ones; Bose 95% and 96%; Sennheiser 94% and 89%; Anker 91% and 100%. Even Apple, the worst case, is 82% versus 72%. Both panels return the same six-brand consideration set, in nearly the same order — the rank correlation between the two share vectors is 0.87.

They disagree on how often. Apple comes back in all five runs for 9.5% of human prompts and 0% of synthetic ones. Sennheiser: 59% versus 22%. Among prompts that surface Apple at all, humans get it in 55% of runs and the synthetic panel in 37%.

That is a meaningfully different result from “synthetic panels see a different market.” They see the same market and report it at different volumes.

Ruling out the two easy explanations

It is not the panel mix. The obvious reading of a share gap is that the panel asks a different distribution of questions. If so, holding the sub-intent fixed should close it. It does not: within travel the gap is 0.085, within music 0.110, against an unconditional 0.089.

Bar chart comparing mean absolute brand-share difference for the unconditional comparison and three sub-intent-restricted comparisons, all sitting at or above the 0.05 equivalence band drawn as a dashed line.

Mean absolute brand-share difference, unconditional and restricted to prompts carrying the same sub-intent flag, with 90% intervals and the pre-registered 0.05 band.

This matters commercially. The natural workaround is to have a brand name the sub-intents it cares about and generate prompts for those, sidestepping the need to estimate a mix. The data says that would produce a realistic consideration set and still not a trustworthy percentage.

It is not noise, and this test was pre-registered. A synthetic prompt re-run five times is exactly as self-consistent as a human prompt re-run five times: −0.038 [−0.086, +0.009] for the stratified panel and −0.018 [−0.060, +0.029] for the unstratified one, both null against a 0.10 band.

That is a non-replication of our own study. Last month we measured all three synthetic panels as materially noisier than humans — −0.104, −0.111 and −0.089, every one a real difference — and concluded synthetic prompts “occupy a less stable region of the response space.” This study’s unstratified panel is an exact replication draw of that generator and lands at −0.018, while the human baseline barely moved (0.724 to 0.704). The two studies’ intervals overlap in the −0.060 to −0.039 range, so they are not statistically distinguishable and this is a non-replication rather than a refutation. What does not survive is the general claim that synthetic prompts are inherently flakier.

With mix and noise both excluded, what remains is a systematic shift in what synthetic phrasing surfaces — consistent in direction in every stratum we can measure: Apple down 17 to 19 points, Sennheiser down 13 to 18, Anker up about 6.

How far the panel drifts depends on where a brand sits in the ranking. Expressed as error relative to each brand’s own human share, rather than in absolute points:

BrandRank in the human panelRelative error
Bose2−0.6%
Sony1−6.4%
Anker4+12.0%
Sennheiser3−13.2%
JBL6−33.1%
Apple5−45.8%

Below this basket it degrades further, and not in one direction. Four brands that human prompts surface in 1.5–3.8% of answers never appear in the stratified panel at all (Bowers & Wilkins, Sonos, Focal, Beyerdynamic), while Shokz appears 2.6× more often than in human answers (1.8% → 4.7%) and Technics somewhat more (2.2% → 2.9%). The accurate statement is not that synthetic panels under-surface the alternatives — it is that outside the top few brands, a synthetic panel’s per-brand number stops being reliable in either direction.

Where the residual lives

The stratification worked precisely where it was applied and bought nothing anywhere else. Measuring how far each panel’s phrasing sits from the human panel’s, across the six dimensions we matched versus the nine we did not:

PanelSix stratified dimensionsNine unstratified dimensions
Stratified0.0390.135
Unstratified0.1970.094

Matching cut error five-fold on its targets and left the rest untouched. One dimension dominates the remainder: the stratified panel raises comfort in 45% of its prompts against the human panel’s 12% — an artifact of the clause template that hits the six targets. Excluding it, the two panels are effectively tied on the unstratified dimensions, so the honest statement is that constraining six dimensions did not improve the other nine, not that it degraded them.

The deeper limit is combinatorial. Human prompts in this study form 79 distinct phrasing profiles; the stratified panel reproduces 13% of them, the unstratified panel 8%. It never asks about watching films on a plane (31% of human prompts), the recipient’s age (10%), how many options to return (9%), or what format to answer in (4%). Matching six marginal rates does not reconstruct the joint distribution — and the answer depends on the joint.

What our own instrument got wrong

This study argues that measurement instruments mislead, so it owes you its own errors. We ran a blind audit: one reviewer labelled 30 answers by hand with the extractor’s output withheld, so the labels were generated independently rather than checked against ours.

On the raw comparison the extractor scored 88.0% precision — below the 95% we had committed to. Re-reading every disagreement against the source text resolved 11 of 14 in the extractor’s favour: they were brands the labelling had passed over, most of them sitting in comparison tables, parentheses, and secondary product headings rather than the main recommendation list. That puts the audited figures at 97.4% precision and 91.9% recall, above the frozen thresholds. Recall is measured against what a human noticed, so it is an upper bound. The three surviving errors, and one gap the audit found, are worth reporting.

Our extractor counts brands used as platforms. “A streaming subscription such as Apple Music” and “what phone does he use — Samsung Galaxy, Google Pixel” both register as brand mentions. This is differential: 5.0% of human answers carry a platform-only Apple reference versus 1.1% of the stratified panel’s. Restricting Apple to product aliases narrows its gap from −21.2 to −17.3 points and moves the headline from 0.089 to 0.083. Roughly 18% of the single largest number in this article is our artifact, and it biases in the direction that flatters our own conclusion.

One brand was missing from the lexicon entirely. JLab appears in 6.7% of human answers, above the 5% threshold for inclusion, so the comparison should have run over seven brands rather than six. Adding it moves the stratified panel to 0.078 and the unstratified panel to 0.119.

Both corrections push the same way, and neither changes a verdict: with both applied the stratified panel sits at 0.073 and the unstratified panel at 0.114, still well outside the 0.05 band. Because the lexicon was frozen before collection, we report these as corrections rather than re-running the analysis with a tuned instrument — adjusting your measuring device after seeing the results is the precise failure that pre-registration exists to prevent.

What we can and cannot claim

What this means if you track a brand in AI answers

Data and reproducibility

Changelog