Spyglasses Research

A reader asked if 'rank is unstable' was measured or folklore. So we measured it.

Spyglasses Research · July 24, 2026 · spec & code

Key findings

Why we measured this

After we published our prompt-phrasing study, a reader asked a question we couldn’t answer with a number: when we said brand rank in AI answers isn’t stable, was that a computed statistic or an eyeball read?

It was an eyeball read, inherited from the January 2026 SparkToro study our experiment replicated and extended. Their data showed ordering bouncing around freely. But their design ran each prompt once across seven different platforms, which means observed rank variation bundled together three things: genuine rank instability, brands entering and leaving the answer entirely (set churn), and platform-to-platform differences. Our replication data — one platform, 143 phrasings, each run daily for a week — can separate them. The reader was right that nobody had published the number, and right that this dataset could produce it. Everything below is a post-hoc analysis of the frozen experiment 002 data, labeled exploratory throughout.

Methods in brief

“Rank” here means the order in which brands are first mentioned in the answer text — a proxy, since ChatGPT’s answers are prose and only sometimes explicit rankings. For every pair of runs (1,001 headphone runs evaluated in this study), we computed Kendall’s tau over the brands both answers contain, requiring at least three shared brands. That makes the statistic order-only and conditional on membership: the churn of brands in and out of answers — which our main study already quantified with set overlap — is deliberately excluded, so tau answers the reader’s actual question. Pairs are compared within-prompt (same prompt, different days; 82% have 3+ shared brands) and between-prompt (different phrasings, same day; 70% qualify). Confidence intervals come from the same prompt-level cluster bootstrap as the main study.

Kendall’s tau runs from +1 (identical order) through 0 (no association) to −1 (exactly reversed). For context, tau on three shared items can only take four values, so the averages below are coarse-grained per pair and stable in aggregate.

Result 1: order is sticky — very sticky for the same prompt

ECDF curves of pairwise Kendall's tau for two conditions. Same-prompt repeated-run pairs concentrate at high positive tau with 46% at exactly 1. Different-prompt same-intent pairs sit lower but still mostly positive, with 31% at exactly 1.

Distribution of pairwise order agreement (Kendall’s tau over shared brands, 3+ required). A large mass sits at tau = 1 — identical ordering — in both conditions. Exploratory.

Re-run the same prompt on different days and the shared brands come back in the same relative order far more often than not: mean tau +0.644 [+0.615, +0.673], perfect agreement in 46.2% of pairs, order-reversals in 0.9%. Only 12% of pairs show zero or negative association. For a statistic the folklore says is noise, that is a lot of signal.

Across differently-worded prompts the agreement drops but doesn’t collapse: mean tau +0.358 [+0.311, +0.410], with 31.2% of pairs still in perfect agreement. Rewording the question costs about as much ordering agreement as it costs set membership — but in neither case does it approach randomness.

The mechanism is no mystery in hindsight: the order of first mention tracks the same popularity gradient that drives membership. Sony doesn’t just appear in 90% of answers; it opens 60% of them. Order inherits stability from the same skewed mention frequencies that make the leaderboard stable.

Result 2: the top slot is the genuinely unstable part

Bar chart of the share of pairs agreeing on the first-mentioned brand: 68% for same-prompt repeated runs, 40% for different prompts with the same intent, with a dashed chance line at 40%.

Agreement on which brand is mentioned first. The dashed line is the chance rate implied by the overall first-brand mix. Between-prompt agreement lands exactly on it. Exploratory.

Whether two answers open with the same brand is a different story, and it’s where the instability intuition survives. The same prompt re-run keeps its opener 68.3% of the time [65.4%, 71.1%] — far above chance. But two differently-worded prompts agree on the opener 39.7% of the time [36.0%, 43.7%], and the chance baseline computed from the empirical first-brand mix is 39.9%. The match is almost embarrassing: with respect to the lead position, rewording the prompt is statistically indistinguishable from drawing an opener at random in proportion to how often each brand opens answers overall. (Answer share within this study, to be clear — we measured nothing about real-world market share, however correlated the two may be.)

So if your interest is “who gets mentioned first,” phrasing genuinely scrambles it. If your interest is “does the ordering of the list mean anything,” it does — and it means the same thing across phrasings much more often than chance.

What this refines about the folklore

The two statistics are designed to be complements: the parent study’s set overlap measures whether a brand survives from one answer to the next, and tau measures order among the survivors. Put them together and the combined claim is sharper than either alone: the churn lives in which brands make the answer; the ordering of the brands that stay is largely conserved. Roughly a quarter of the set turns over between same-prompt runs, and the surviving order mostly holds. That’s more precise than a blanket verdict for or against rank as a metric, and we haven’t seen it stated with numbers attached before.

The “AI rank is chaos” impression comes from real observations, but our decomposition suggests it was mostly measuring other things:

  1. Set churn, not rank churn. Brands entering and leaving the answer (our main study’s headline effect) destroys apparent rankings without any reordering of what remains. Tau conditions that away and finds order beneath the churn.
  2. Platform mixing. Single runs spread across seven platforms — the original study’s design, appropriate for its question — can’t distinguish rank instability from platforms simply ordering brands differently. On one platform, held constant, order is moderately-to-strongly conserved.
  3. The top slot really is unstable across phrasings — at exactly chance — and the top slot is what most people look at. The most salient rank position is the least stable one, which is exactly the recipe for a folklore of chaos built on accurate glimpses.

To be clear about the relationship to SparkToro’s work: their conclusion was sound at their design’s scope, and their study is the reason this dataset exists. This result doesn’t overturn it; it decomposes it, on one platform, into a stable component (relative order among shared brands) and an unstable one (lead position across phrasings).

Two ways these numbers could flatter us — checked

The same reader flagged two selection effects that could make tau look better than it is. Both are checkable, so we checked (the numbers are in the analysis output).

Are the excluded pairs the unstable ones? Tau requires 3+ shared brands, which drops 18% of within-prompt pairs — plausibly the least stable, since low overlap and disorder could travel together. The largest excluded group, pairs sharing exactly two brands, permits one order comparison each: within-prompt, those pairs are concordant 76.1% of the time against a 50% coin — equivalent to tau ≈ +0.52. Between-prompt, 59.4% (tau ≈ +0.19). So the direction of the concern is right — the dropped tail is less ordered than the included pairs — but the size is small: the excluded pairs are diluted, not chaotic, and no headline changes if you fold them in. Within included pairs, mean tau is also nearly flat across shared-set sizes (+0.60 at three shared brands, +0.65 at four, +0.70 at five), so there’s no gradient suggesting the tail below the cutoff falls off a cliff.

Does the 12% of disordered pairs mean some prompts are hopeless? Only if that disorder concentrated in particular prompts — and it doesn’t. The 295 zero-or-negative within-prompt pairs are scattered across 63 different prompts; the five worst prompts account for just 54 of them. At the prompt level, exactly 2 of 136 prompts have a mean within-prompt tau at or below zero, and the 10th-percentile prompt still averages +0.22. Ordering stability is heterogeneous — the median prompt averages +0.68 and the bottom decile is mediocre rather than inverted — but essentially no prompt lives in the chaotic regime. A negative-tau pair is an event that happens to a stable prompt occasionally, not a property some prompts have.

What we can and cannot claim

What this means if you track a brand

Data and reproducibility

Changelog