You Can't Test Your Way Out of a Black Box - How to Reframe LLM 'Visibility' Measurement
From SMX Advanced Berlin, Amanda King @ FLOQ walks through how to understand what is accurate and inaccurate about what we're being sold about LLM measurement platforms, and how to think about it instead.
how we know what we know. THE SETUP THE PIPELINE THE RESPONSE The pitch, and the play we have run before. The retrieval pipeline, in three tiers. Where the energy goes instead. • 01 · The pitch being sold • • • 02 · Rankings and last-click 04 · Tier one: documented 07 · What the dashboard measures • 05 · Tier two: reverse-engineered • 08 · Redirect: retrievability • 06 · Tier three: genuinely opaque • 09 · Leading indicators • 03 · What AI exposed FLOQ
Share points. Basis points. The room speaks fluent points, and this score speaks it back. It has a denominator: mentions over responses tracked. Soft in three places. • Whose prompt set • How many runs • Which surface. The API is not the app https://www.tryprofound.com/blog/how-to-track-your-visibility-in-ai-search · https://docs.peec.ai/metrics/brand-metrics/share-of-voice FLOQ
a better dashboard. The current advice for AI visibility, at every price point. Build a rigorous enough testing protocol and you will finally know where you stand. FLOQ
Citation rate. Point scores, no intervals. Probabilistic outputs, sold as deterministic numbers your CMO can put in a slide. • A score, to one decimal • No confidence interval in sight • A trendline, week on week https://floq.co/strategy/metrics-breaking-seo-success-different-now/ FLOQ
$100M a year goes into AI visibility, on Rand Fishkin's estimate. He went looking for research showing the answers hold still enough to track. Nothing. The market got built first. • $100M+ a year, estimated • Zero consistency research before 2026 • The question came after the invoice https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-s hould-take-care-when-tracking-ai-visibility/ FLOQ
last-click attribution. Take an opaque system, build a tracker that produces a number, optimise the number, congratulate ourselves when it goes up. FLOQ
a personalised, fluctuating system, reported as a fact. Two people, same query, different page one. We knew. We reported position three anyway. • Personalised per user • Fluctuating per day • Reported as a constant https://www.google.com/intl/en_us/search/howsearchworks/how-search-works/ranking-results/ FLOQ
causation, filed under attribution. Even Google walked away: data-driven attribution replaced last-click as the Ads default in 2021. • A guess with a decimal point • A model everyone agreed not to poke • A number the bonus was tied to https://www.adexchanger.com/online-advertising/goodbye-last-click-attribution-google-ads-changes-default-to-data-modeli ng/ FLOQ
crawl-to-refer. At one 2025 read, Anthropic's platforms made roughly 71,000 page requests per referral sent back. Google: about five to one. • Training drives ~80% of AI crawling • Search-purpose crawling: under 10% • Cloudflare flags its own blind spot https://blog.cloudflare.com/ai-search-crawl-refer-ratio-on-radar/ FLOQ
on. AI summaries eat the informational clicks that padded the graph. Discover and AI Mode land as direct. The visits that arrive convert at 4.4x and get filed as noise. • Pew: clicks fall from 15% to 8% under a summary • Lily Ray: 363K Discover clicks, invisible in GA4 • AI referrals: 4.4x conversion, filed as noise https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-a ppears-in-the-results/ FLOQ
ends 16 February 2026. Gemini 3 publishes January 2025. The world keeps moving. The model's memory of you doesn't, until the next training run. • GPT-5.6: February 2026 • Gemini 3: January 2025 • Retrieval patches answers, not memory https://metehan.ai/articles/llm-knowledge-cutoff-dates/ FLOQ
snapshot of the web into weights. The pages are shredded on the way in. The model keeps the gist, loses the sources. A citation from memory is an invention. • Knowledge without provenance • Nothing to link to • Citation-shaped confabulation https://arxiv.org/abs/2005.11401 FLOQ
system runs real searches, fetches pages, and anchors the answer to them. A citation is a pointer to a document it just held. The only path a brand can win. • Fan-out queries, real fetches • Citations: receipts of retrieval • No lookup, nothing to win https://arxiv.org/abs/2005.11401 · https://developers.googleblog.com/en/gemini-api-and-ai-studio-now-offer-grounding-with-google-search/ FLOQ
enriches answers by searching. Gemini reads Google. ChatGPT launched on Bing, now a blended stack. Grounding, RAG, browsing: same job. Timeliness, rented from an index. • Gemini: Google • ChatGPT: blended, Bing loosening • Perplexity: closest to classic search https://blogs.bing.com/search/may_2023/Bing-at-Microsoft-Build-2023 · https://developers.googleblog.com/en/gemini-api-and-ai-studio-now-offer-grounding-with-google-search/ FLOQ
it. Every prompt gets a prediction score, 0 to 1, for whether grounding would help. Developers set the threshold. Default 0.3. Under the line: no retrieval, no citations. • Prediction score: 0 to 1 • Default threshold: 0.3 • Under the line, citations do not exist https://developers.googleblog.com/en/gemini-api-and-ai-studio-now-offer-grounding-with-google-search/ FLOQ
it, with no mechanics. OpenAI documents the API tool. Anthropic publishes the app's search rules. Microsoft documents query derivation and shows the queries. • Google: the name, no mechanics • OpenAI: the API, not the app • Anthropic and Microsoft: the app, with rules https://developers.google.com/search/docs/appearance/ai-features · https://developers.openai.com/api/docs/guides/tools-web-search FLOQ
app. Two modes: non-reasoning search passes your query straight to the tool; agentic search lets a reasoning model plan. The output lists each action: search, open page, find in page. • Queries usually visible in the output • Every search action billed • Nothing on what ChatGPT itself does https://developers.openai.com/api/docs/guides/tools-web-search ; https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/web-search FLOQ
prompt with every release. Twenty-nine snapshots, including the search rule: stable information, answer; time-sensitive, search immediately. • The app's instructions, published • When to search: a written rule • Query construction: leaked, not published https://docs.claude.com/en/release-notes/system-prompts · https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool FLOQ
how queries derive from prompts, with examples, and lists the exact Bing queries in each response. Consumer Copilot exposes them after the fact, in Webmaster Tools. • M365 Copilot: queries shown in the answer • Copilot Studio: named selection checks • Bing WMT: grounding queries per site, Feb 2026 https://learn.microsoft.com/copilot/microsoft-365/manage-public-web-access · https://searchengineland.com/bing-webmaster-tools-ai-performance-report-468751 FLOQ
$35 per 1,000 grounded prompts, and on Gemini 3, per search query the model fires. One prompt, several queries, each one billed. Grounding budgets are not a metaphor. • $35 per 1,000 grounded prompts • Gemini 3: billed per fan-out query • Every citation, someone paid to find https://ai.google.dev/gemini-api/docs/pricing FLOQ
ages. Independent researchers filling the gap the platforms leave. The good ones publish their error bars. The findings are real. So are the expiry dates. FLOQ
OVERLAP 46% 87% of ChatGPT prompts fire a live search of citations matched Bing's top 20 80 million queries reviewed. Semrush. In 2024. Controlled study, Seer Interactive. https://www.semrush.com/blog/chatgpt-search-insights/ · https://www.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-results FLOQ
only 12% of AI-cited URLs rank in the top 10 for the original prompt. They rank for the fan-outs. Profound: citations spread almost evenly across positions one to ten. • 12% overlap with the prompt • Perplexity the outlier: 1 in 3 • Flattening: position 10 nearly equals 1 https://ahrefs.com/blog/ai-search-overlap/ · https://www.tryprofound.com/blog/ai-search-shift FLOQ
CONTESTED Retrieval is search-grounded Which index, and when Top-20 visibility feeds the shortlist The grounding budget numbers Your patterns propagate Any number older than a quarter Every engine rents timeliness from an index. The pool retrieval draws from stays shallow. Fan-out returns the description you codified. Blended stacks shift with each partnership. Publicly challenged. Petrovic engages with it. Treat published percentages as dated. https://www.tryprofound.com/blog/ai-search-shift · https://www.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-re FLOQ
the Bing API in 2025. OpenAI built its own crawler and index. Profound measured Bing alignment collapsing from 26% to 8% while Google alignment roughly tripled. • 87% overlap in 2024 • API retired in 2025 • 8% alignment by 2026 https://learn.microsoft.com/en-us/lifecycle/announcements/bing-search-api-retirement · https://www.tryprofound.com/blog/ai-search-shift FLOQ
index. A September config leak enumerated 69 retrieval engines, 47 of them Labrador's: web, PDF, news by recency, arXiv, Reddit, legal, medical, finance, local. • Sworn testimony: building since 2023 • 69 engines in the registry, 47 Labrador • Fan-out routes to vertical engines https://www.linkedin.com/pulse/complete-list-every-search-engine-hidden-inside-leak-david-konitzny-uwofe/ · https://www.justice.gov/atr/media/1402141/dl FLOQ
Data and Oxylabs scraping Google in, Yelp and TripAdvisor licensed, Bing mainly in Deep Research, and Microsoft's Web IQ, which Microsoft confirms powers ChatGPT. • Google: scraped in, confirmed by GSC spikes • What looks like Bing may be Web IQ • Lockdown mode serves cached pages https://peec.ai/blog/chatgpt-built-its-own-search-index · https://www.microsoft.com/en-us/webiq FLOQ
Gemini's grounding: what gets pulled, how much gets read, which sentences survive. Roughly 2,000 words of budget, a third of a page, front-loaded. • A ~2,000-word grounding budget • Roughly 32% of a page used • Write for extraction, or lose the read https://dejan.ai/blog/hacking-gemini/ FLOQ
weeks. Early August: 5.6 Luna becomes the free default. Single-query fan-outs fall from 94% of prompts to 43.5% overnight. Retrieved sources double. Every prior study describes a different machine. • Single fan-outs: 94% to 43.5% • Sources per prompt: roughly doubled • site: operator: 0.3% to 23% https://www.linkedin.com/pulse/chatgpts-new-default-model-gpt-56-more-retrieval-content-konitzny-8htce/ · https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/ FLOQ
real searches: clicks fell 15% to 8% under a summary. Ahrefs put the position-one hit at 58%. Semrush found zero-click rates didn't rise. All three are careful. • Pew: clicks nearly halve • Ahrefs: 58% down at position one • Semrush: no change on matched keywords https://ahrefs.com/blog/ai-overviews-reduce-clicks-update/ · https://www.semrush.com/blog/semrush-ai-overviews-study/ FLOQ
and citation are moving in opposite directions. More pages considered per answer, fewer credited. Half the fights between studies come down to which layer got counted. • Retrieved URLs: rising • Cited domains: falling • Check which one a study measured https://lilyraynyc.substack.com/p/what-we-can-learn-from-evolving-chatgpt FLOQ
THE EXPERIMENT THE RESULT 2,961 <1 in 100 runs across three platforms runs returned the same list of brands 600 volunteers, 12 prompts. ChatGPT, Claude, Google AI. Same list in the same order: fewer than 1 in 1,000. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-s hould-take-care-when-tracking-ai-visibility/ FLOQ
prompts for two weeks. The engines don't search for what you typed. They fan out, sometimes a dozen queries, sometimes departing from your intent entirely. • ChatGPT: wide-net researcher • Perplexity: nearly 1:1 with your prompt • Copilot: compresses to a few strings https://www.tryprofound.com/blog/what-ai-engines-actually-search-for FLOQ
appends best-of to a quarter of advice fan-outs. It fetches reviews you never asked for. Fusion scoring picks winners. The question answered is not quite the question asked. • Best-of: added to 24.3% of advice fan-outs • Reviews: fetched unasked • Fusion scoring picks the winners https://peec.ai/blog/patterns-we-see-in-chatgpt-query-fanouts FLOQ
prompt, a thousand runs, randomness off: dozens of different answers. Output depends on batch size. Batch size depends on server load. • Non-determinism is the default, at kernel level • Load changes the maths path • The fix: perfect, open-sourced, 61.5% slower. https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/ FLOQ
research: the citations for a stable prompt set can change by up to 60% in a single month. Not your content. Not your rankings. The pipe itself, moving. • Up to 60% churn, monthly • The vendor's own finding • Stability is the anomaly https://www.tryprofound.com/blog/ai-search-volatility FLOQ
fell 86-94% across markets February to April 2026, then mostly rebounded in May. One model swap cut cited domains by a fifth overnight. None of it tracked anything brands did. • US zero-citation rate: 28% to 48% in March • Cited domains: 19 to 15 on one model swap • May: mostly rebounded https://www.seoclarity.net/chatgpt-citation-decline-analysis · https://searchengineland.com/inside-chatgpt-search-web-run-fan-out-queries-ai-visibility-477339 FLOQ
March 2026: ChatGPT's default moved to GPT-5.3 Instant. Over 14 weeks of 400 daily prompts, domains per answer fell from 19 to 15, URLs from 24 to 19. The crawl thinned in step. • Domains per answer: 19 to 15 • URLs per answer: 24 to 19 • Fewer fetches, not just fewer citations https://think.resoneo.com/chatgpt/5.3-5.4/ FLOQ
tracked five markets. Wave one, 8 March. Wave two, 19 April, worse. By end of April, citation volume was down 86 to 94% everywhere. In May, it rebounded toward pre-March levels. • US zero-citation answers: 28% to 48% • Down 86 to 94% by end of April • May: back toward February https://www.seoclarity.net/chatgpt-citation-decline-analysis FLOQ
Surfer injected a leaked system prompt into clean API calls: still not the interface. Luxeo found the same across web, API and mobile. Vendors concede it. Nobody documents it. • Web, app, API: different answers • System prompt injected: still different • Even the tracking vendors concede it https://surferseo.com/blog/llm-scraped-ai-answers-vs-api-results · https://luxeo.team/chatgpt-web-vs-api-experiment/ FLOQ
caught ChatGPT running site: searches against parked domains the brands don't own. Netcraft: roughly a third of AI login links point somewhere the brand doesn't control. • site:census.com. Census isn't there. • A third of login links: wrong • Parked domains are buyable https://www.techradar.com/pro/security/chatgpt-and-other-ai-tools-could-be-putting-users-at-risk-by-getting-company-we b-addresses-wrong FLOQ
can't inspect the data or identify the raters. 12,000 people rated Google's AI, 500 words in 15 minutes. Bias analysis covers American English. Their preferences are its ethics. • Training data: uninspectable • Raters: unidentifiable • Governance: shaped by IPO incentives https://floq.co/strategy/what-we-dont-know-about-ai-training/ FLOQ
physical availability, mental availability, distinctiveness, reputation, commercial performance. Strengthen them, you surface. Weaken them, nothing fixes it. • The signals are the cause • The score is the shadow • Most of the six aren't SEO's to drive https://www.jonoalderson.com/conjecture/clicks-dont-count/ FLOQ
DRIFT CALL Presence frequency 55-77% stable KEEP Mention rate, with CIs 5 runs bounded KEEP Fan-out presence 5 runs binary KEEP Rank position in AI <1/100 1/1000 LET FADE n = 1 ±60%/mo LET FADE Single-snapshot scores Rose marks the negative call. If a metric can't survive repetition, it doesn't survive the quarter. https://searchengineland.com/make-prompt-tracking-more-accurate-479708 · https://www.tryprofound.com/blog/research FLOQ
an open tool that predicts whether a query triggers grounding. Run your prompt set through it before you track anything. Prompts under the threshold have no citation layer. • Predict grounding, per prompt • Split the set: grounded vs memory • Only the grounded half has citations to win https://dejan.ai/blog/how-google-decides-when-to-use-gemini-grounding-for-user-queries/ ;https://huggingface.co/dejanseo/query-grounding FLOQ
Kevin Indig's framework: repeated runs, fixed sampling rules, confidence intervals on every rate you report. A distribution you can measure, not a draw you can chase. • Five runs per prompt, per platform • Mention and citation rate, with intervals • Trend, not position https://searchengineland.com/make-prompt-tracking-more-accurate-479708 FLOQ
own docs: Semrush's shared prompt index produced results inconsistent with independent data, with big month-to-month swings driven by methodology changes, not visibility. • Vendor on vendor • Methodology moves the number • Ask what changed: the market, or the index https://peec.ai/ai-instructions FLOQ
same way. Suganthan Mohanadasan watched the network traffic: ChatGPT names brands in its first search, before reading a single page. The shortlist forms before retrieval, in the model's prior. • Named in the fan-out: cited 68.9% • Merely fetched: cited 2.1% • Rank feeds the pool. Recognition picks it. https://suganthan.com/blog/chatgpt-decides-before-it-searches/ FLOQ
to understand who you are. Jes Scholz's framing, and she's right. Inconsistency makes it 40, 60, 80, and one will be Joe Blow's Reddit thread from ten years ago. • Same name, everywhere • Same description, everywhere • One pattern for the machines to return https://searchengineland.com/establish-brand-entity-for-seo-guide-427717 FLOQ
with 600 press articles. Zero AI visibility. Volume without entity presence or corroboration doesn't register. The machines can tell known from placed. • 600 articles • Zero visibility • Corroboration beats volume https://searchengineland.com/ai-recommendations-inconsistent-fix-469250 FLOQ
All of it is what retrieval rewards. Technical Brand Content Legible to the engines that feed the models. One entity, impossible to misread. Fewer pages, aimed at your expertise. • Per-bot robots.txt policy • • • Sitemaps to spec, JS fallbacks One description, everywhere Consolidate: ~35% uplift, done well • • Tech debt, actually addressed About page as evidence locker Kill invalid informational topics • Knowledge panel, press, Wikipedia • Answer what only you can • https://searchengineland.com/geo-and-seo-how-to-invest-your-time-and-efforts-wisely-461424 · https://floq.co/strategy/content-consolidation-meta-study/ FLOQ
and tokenise at retrieval. Schema is not a lever on the model. It's a vehicle for the engines that feed the model. Do it anyway. • Stripped in training • Read by the engines underneath • Worth it for that reason alone https://aws.amazon.com/blogs/machine-learning/an-introduction-to-preparing-your-own-dataset-for-llm-training/ FLOQ
2024: citations, quotations and statistics lifted visibility 30 to 40% on their benchmark. Keyword stuffing went backwards. Evidence density is the lever. • Cite, quote, quantify: +30 to 40% • Keyword stuffing: minus 8 to 10% • A benchmark, not production. Still. https://arxiv.org/abs/2311.09735 FLOQ
88.1% of AI Overview triggers are informational. Chasing head terms padded the graph without building a good business and a good product underneath. AI took the padding. • 88.1% of triggers: informational • Padding, not pipeline • Let it go https://searchengineland.com/google-ai-overviews-13-searches-455057 FLOQ
stays in traditional SEO and brand. Do the 1 to 5% technical changes deferred for years. Pilot the AI layer. Sell it as non-negative and compounding. • 80: foundations and brand • 20: the systemic schema layer • Baby steps beat big bets https://searchengineland.com/geo-and-seo-how-to-invest-your-time-and-efforts-wisely-461424 FLOQ
discards the Reddit pages it pulls 99% of the time. Reddit still tops citation charts on sheer volume. And it now flags ~25,000 GEO-spam posts a day. • Retrieved relentlessly • ~1% survival, still #1 cited • Don't spam it. They're watching. https://dejan.ai/blog/reddit-ai/ · https://www.emarketer.com/content/reddit-geo-crackdown-could-raise-stakes-ai-visibility-strategies FLOQ
of voice above share of market predicts growth. The gap, not the level. Share of search via Trends, or your top 20 SERPs. Pick one. Stick to it. • Track the gap, not the number • Trends or SERP share • Consistency over precision https://www.contagious.com/news-and-views/share-of-search-the-new-most-important-metric-for-brands-google FLOQ
anything, branded search grows. Slowly, lumpily. Free in Search Console, legible to the board. Flat after a year means something isn't landing. Worth knowing. • Free • Legible to stakeholders • Honest when it's flat https://floq.co/strategy/metrics-breaking-seo-success-different-now/ FLOQ
Pick two or three proxies that match what the business is trying to do and watch them quarterly. Know what's under the number when it moves. • Direct demand, next to branded • Reviews: volume, recency, coherence • Close rates, retention, conversion quality https://www.jonoalderson.com/conjecture/clicks-dont-count/ FLOQ
market down 15% is outperformance. Up 3% while a competitor grows 12% is losing. Sales, product and finance all report against the market. We report snapshots. Stop. • Year on year • Against named competitors • Trajectory, not totals https://floq.co/strategy/metrics-breaking-seo-success-different-now/ FLOQ
is tied to the old number. Change mid-year and they'll fight you, however right you are. Run parallel a quarter, cut over at the financial year. You carry the meaning. • Parallel run first • Cut over at year end • Translate, constantly https://floq.co/strategy/metrics-breaking-seo-success-different-now/ FLOQ
trendline and a sentence. Directors get the proxies for their function. Your team gets the messy guts. One dashboard for everyone serves no one. • CxO: excess share of voice, one line • VPs: proxies by function • Team: everything underneath https://portent.com/blog/analytics/marketing-reports.htm FLOQ
Tools now reports AI citations and the grounding queries Copilot generates. The site: fan-outs surface in Search Console: 197,000 impressions, one click. • Bing: AI Performance Report, Feb 2026 • GSC: site: queries, near-zero CTR • Free, first party, already yours https://otterly.ai/blog/bing-webmaster-tools-ai-performance-report/ · https://lilyraynyc.substack.com/p/what-we-can-learn-from-evolving-chatgpt FLOQ
actually know. Make yourself worth retrieving. Movement, against the market. • • Per-bot robots.txt policy • • One entity description, everywhere GSC site: filter + Bing AI report • Share-of-search baseline • Change parallel run to year end Classify prompts by grounding • Five-run fan-out check • Presence frequency, with intervals • Front-load the extractable facts FLOQ
presence frequency at sample size. A poll, not a fact. Move some of the testing budget upstream: entity consistency, tech debt, consolidation. Report movement against the market. The gap, not the snapshot. FLOQ