Spyglasses Research
Pre-registered studies of how AI assistants find, evaluate, and cite sources, from the team behind the Spyglasses AI visibility platform. Every study publishes its methodology, analysis code, and — where possible — an anonymized dataset.
- We tried sourcing prompts from real human conversations. The material is abundant. Using it is the hard part.
A methodology note, not a study. The advice to source AI-visibility prompts from Reddit and similar venues sounds like it dissolves the synthetic-prompt problem. We ran it properly for one product category: volume was never the obstacle, but only about 1 post in 80 is usable as a prompt unedited. Clipping them at sentence boundaries fixes that for 93% and strips the author's sub-intent from most of them. Includes the wrong answer we reached first, and why our own frozen protocol did not catch it.
- A synthetic panel matched to human sub-intent reproduces the consideration set — but not the share of voice.
Experiment 005, pre-registered: a scenario-generated prompt panel stratified to 143 human survey phrasings' sub-intent profile, run beside an unstratified panel and the re-run human panel on ChatGPT for five days. Its answers are statistically indistinguishable from human ones — the closest match we have measured — and its brand-share vector still misses by 8.9 points against a 5-point band. The gap turns out to be about how often each brand appears, not which brands appear. We also failed to replicate our own finding that synthetic prompts are noisier.
- We tested our own prompt generator against 143 humans. It measures your home field, and we found what a market-matching panel needs.
Experiment 003: Four prompt panels, 143 human survey phrasings, two brand-anchored panels from our own production generator, and one neutral scenario panel — run side by side on ChatGPT for five days, pre-registered. Brand-anchored panels fail every market-representativeness test; what they measure instead is your standing on the buying situations your own positioning claims. But a neutral scenario panel's prompts are statistically indistinguishable from human phrasings, and when synthetic and human prompts share a sub-intent, their answers match. This is a concrete, testable path to synthetic panels that mirror human ones, which we are pre-registering next.
- A reader asked if 'rank is unstable' was measured or folklore. So we measured it.
Follow-up to our SparkToro replication: the first public coefficient on AI recommendation rank stability. Across 1,001 ChatGPT runs, the order of recommended brands agrees at Kendall's tau 0.64 when the same prompt is re-run and 0.36 across different phrasings — and nearly half of same-prompt pairs return shared brands in exactly the same order. The 'rank chaos' folklore mostly isn't about rank.
- We pre-registered 'prompt phrasing barely matters.' ChatGPT proved us wrong.
Extending SparkToro's prompt-diversity study: we pre-registered the hypothesis that prompt wording doesn't matter when intent matches, then re-ran their 143 de-identified survey prompts against one platform — ChatGPT — every day for a week. The leading brands held, but rewording moved the shortlist, the cited sources, and the searches underneath, well beyond run-to-run noise.
- Only one AI surface deep-links to the moment inside a YouTube video
Google AI Overviews attaches a timestamp to 44% of the YouTube videos it cites, sending users to the exact second that answers their question. ChatGPT, Gemini, Perplexity, and Claude never do. We looked at what predicts a moment citation — and whether Google is reading the transcript.