Upgrade to Pro — share decks privately, control downloads, hide ads and more …

You Can't Test Your Way Out of a Black Box - Ho...

You Can't Test Your Way Out of a Black Box - How to Reframe LLM 'Visibility' Measurement

From SMX Advanced Berlin, Amanda King @ FLOQ walks through how to understand what is accurate and inaccurate about what we're being sold about LLM measurement platforms, and how to think about it instead.

Avatar for Amanda King

Amanda King

September 30, 2026

More Decks by Amanda King

Other Decks in Marketing & SEO

Transcript

  1. SMX ADVANCED BERLIN You can't test your way out of

    a black box. Amanda King floq · 2026
  2. THE AGENDA Forty-five minutes, three tiers. The pipeline, sorted by

    how we know what we know. THE SETUP THE PIPELINE THE RESPONSE The pitch, and the play we have run before. The retrieval pipeline, in three tiers. Where the energy goes instead. • 01 · The pitch being sold • • • 02 · Rankings and last-click 04 · Tier one: documented 07 · What the dashboard measures • 05 · Tier two: reverse-engineered • 08 · Redirect: retrievability • 06 · Tier three: genuinely opaque • 09 · Leading indicators • 03 · What AI exposed FLOQ
  3. THE SCENE These points have a soft denominator. NPS points.

    Share points. Basis points. The room speaks fluent points, and this score speaks it back. It has a denominator: mentions over responses tracked. Soft in three places. • Whose prompt set • How many runs • Which surface. The API is not the app https://www.tryprofound.com/blog/how-to-track-your-visibility-in-ai-search · https://docs.peec.ai/metrics/brand-metrics/share-of-voice FLOQ
  4. MEANWHILE We've all been in that meeting. Me, in the

    car park, wondering if the metric reported was actually real FLOQ
  5. THE PITCH Run more prompts. Run them more times. Buy

    a better dashboard. The current advice for AI visibility, at every price point. Build a rigorous enough testing protocol and you will finally know where you stand. FLOQ
  6. THE PITCH What the tools promise. Visibility scores. Mention frequency.

    Citation rate. Point scores, no intervals. Probabilistic outputs, sold as deterministic numbers your CMO can put in a slide. • A score, to one decimal • No confidence interval in sight • A trendline, week on week https://floq.co/strategy/metrics-breaking-seo-success-different-now/ FLOQ
  7. THE PROMISE Run it rigorously enough and you'll finally know

    where you stand. That's the promise. FLOQ
  8. THE SPEND Nobody checked whether the thing holds still. Over

    $100M a year goes into AI visibility, on Rand Fishkin's estimate. He went looking for research showing the answers hold still enough to track. Nothing. The market got built first. • $100M+ a year, estimated • Zero consistency research before 2026 • The question came after the invoice https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-s hould-take-care-when-tracking-ai-visibility/ FLOQ
  9. THE PATTERN We've run this play before. Twice. Rankings, then

    last-click attribution. Take an opaque system, build a tracker that produces a number, optimise the number, congratulate ourselves when it goes up. FLOQ
  10. PLAY ONE What a ranking ever was. A snapshot of

    a personalised, fluctuating system, reported as a fact. Two people, same query, different page one. We knew. We reported position three anyway. • Personalised per user • Fluctuating per day • Reported as a constant https://www.google.com/intl/en_us/search/howsearchworks/how-search-works/ranking-results/ FLOQ
  11. PLAY ONE 'We rank #3' sounds precise. It tells you

    nothing about whether anything is improving for the business. FLOQ
  12. PLAY TWO Last-click ran the same trick. Correlation dressed as

    causation, filed under attribution. Even Google walked away: data-driven attribution replaced last-click as the Ads default in 2021. • A guess with a decimal point • A model everyone agreed not to poke • A number the bonus was tied to https://www.adexchanger.com/online-advertising/goodbye-last-click-attribution-google-ads-changes-default-to-data-modeli ng/ FLOQ
  13. IN THEIR WORDS Last-click attribution was already a polite lie.

    — Jono Alderson, Friends of Search. https://ddma.nl/kennisbank/jono-alderson-measuring-what-matters-clicks-dont-count/ FLOQ
  14. THE EXPOSURE The exchange rate, measured. Cloudflare built the metric:

    crawl-to-refer. At one 2025 read, Anthropic's platforms made roughly 71,000 page requests per referral sent back. Google: about five to one. • Training drives ~80% of AI crawling • Search-purpose crawling: under 10% • Cloudflare flags its own blind spot https://blog.cloudflare.com/ai-search-crawl-refer-ratio-on-radar/ FLOQ
  15. THE EXPOSURE Then AI ate the clicks the graph ran

    on. AI summaries eat the informational clicks that padded the graph. Discover and AI Mode land as direct. The visits that arrive convert at 4.4x and get filed as noise. • Pew: clicks fall from 15% to 8% under a summary • Lily Ray: 363K Discover clicks, invisible in GA4 • AI referrals: 4.4x conversion, filed as noise https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-a ppears-in-the-results/ FLOQ
  16. TIER ONE What the platforms actually document. This tier is

    short. That is the point. Hold onto how little of it there is. FLOQ
  17. TIER ONE The cutoff moved. The mechanism didn't. GPT-5.6's memory

    ends 16 February 2026. Gemini 3 publishes January 2025. The world keeps moving. The model's memory of you doesn't, until the next training run. • GPT-5.6: February 2026 • Gemini 3: January 2025 • Retrieval patches answers, not memory https://metehan.ai/articles/llm-knowledge-cutoff-dates/ FLOQ
  18. TIER ONE Every answer takes one of two paths: memory,

    or lookup. Only the lookup produces something to cite. Document* grounding is the lookup. FLOQ
  19. TIER ONE Closed book: the memory path. Training compresses a

    snapshot of the web into weights. The pages are shredded on the way in. The model keeps the gist, loses the sources. A citation from memory is an invention. • Knowledge without provenance • Nothing to link to • Citation-shaped confabulation https://arxiv.org/abs/2005.11401 FLOQ
  20. TIER ONE Open book: the lookup path. Before answering, the

    system runs real searches, fetches pages, and anchors the answer to them. A citation is a pointer to a document it just held. The only path a brand can win. • Fan-out queries, real fetches • Citations: receipts of retrieval • No lookup, nothing to win https://arxiv.org/abs/2005.11401 · https://developers.googleblog.com/en/gemini-api-and-ai-studio-now-offer-grounding-with-google-search/ FLOQ
  21. TIER ONE Retrieval is how they stay current. The model

    enriches answers by searching. Gemini reads Google. ChatGPT launched on Bing, now a blended stack. Grounding, RAG, browsing: same job. Timeliness, rented from an index. • Gemini: Google • ChatGPT: blended, Bing loosening • Perplexity: closest to classic search https://blogs.bing.com/search/may_2023/Bing-at-Microsoft-Build-2023 · https://developers.googleblog.com/en/gemini-api-and-ai-studio-now-offer-grounding-with-google-search/ FLOQ
  22. TIER ONE Document* grounding is a scored decision. Google documents

    it. Every prompt gets a prediction score, 0 to 1, for whether grounding would help. Developers set the threshold. Default 0.3. Under the line: no retrieval, no citations. • Prediction score: 0 to 1 • Default threshold: 0.3 • Under the line, citations do not exist https://developers.googleblog.com/en/gemini-api-and-ai-studio-now-offer-grounding-with-google-search/ FLOQ
  23. TIER ONE Fan-out: named once, documented in pieces. Google named

    it, with no mechanics. OpenAI documents the API tool. Anthropic publishes the app's search rules. Microsoft documents query derivation and shows the queries. • Google: the name, no mechanics • OpenAI: the API, not the app • Anthropic and Microsoft: the app, with rules https://developers.google.com/search/docs/appearance/ai-features · https://developers.openai.com/api/docs/guides/tools-web-search FLOQ
  24. TIER ONE OpenAI documents it. In the API, not the

    app. Two modes: non-reasoning search passes your query straight to the tool; agentic search lets a reasoning model plan. The output lists each action: search, open page, find in page. • Queries usually visible in the output • Every search action billed • Nothing on what ChatGPT itself does https://developers.openai.com/api/docs/guides/tools-web-search ; https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/web-search FLOQ
  25. TIER ONE Anthropic publishes the rulebook. Anthropic publishes claude.ai's system

    prompt with every release. Twenty-nine snapshots, including the search rule: stable information, answer; time-sensitive, search immediately. • The app's instructions, published • When to search: a written rule • Query construction: leaked, not published https://docs.claude.com/en/release-notes/system-prompts · https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool FLOQ
  26. TIER ONE Microsoft shows the queries. Twice. Enterprise Copilot documents

    how queries derive from prompts, with examples, and lists the exact Bing queries in each response. Consumer Copilot exposes them after the fact, in Webmaster Tools. • M365 Copilot: queries shown in the answer • Copilot Studio: named selection checks • Bing WMT: grounding queries per site, Feb 2026 https://learn.microsoft.com/copilot/microsoft-365/manage-public-web-access · https://searchengineland.com/bing-webmaster-tools-ai-performance-report-468751 FLOQ
  27. TIER ONE Named, exposed, billed, capped, even explained. What nobody

    publishes is the weighting: why those results won. FLOQ
  28. TIER ONE Grounding has a price tag. Google bills it:

    $35 per 1,000 grounded prompts, and on Gemini 3, per search query the model fires. One prompt, several queries, each one billed. Grounding budgets are not a metaphor. • $35 per 1,000 grounded prompts • Gemini 3: billed per fan-out query • Every citation, someone paid to find https://ai.google.dev/gemini-api/docs/pricing FLOQ
  29. TIER TWO What we've worked out. And how fast it

    ages. Independent researchers filling the gap the platforms leave. The good ones publish their error bars. The findings are real. So are the expiry dates. FLOQ
  30. TIER TWO The pipeline runs through search. THE TRIGGER THE

    OVERLAP 46% 87% of ChatGPT prompts fire a live search of citations matched Bing's top 20 80 million queries reviewed. Semrush. In 2024. Controlled study, Seer Interactive. https://www.semrush.com/blog/chatgpt-search-insights/ · https://www.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-results FLOQ
  31. TIER TWO Cited for queries you never saw. Ahrefs, 2026:

    only 12% of AI-cited URLs rank in the top 10 for the original prompt. They rank for the fan-outs. Profound: citations spread almost evenly across positions one to ten. • 12% overlap with the prompt • Perplexity the outlier: 1 in 3 • Flattening: position 10 nearly equals 1 https://ahrefs.com/blog/ai-search-overlap/ · https://www.tryprofound.com/blog/ai-search-shift FLOQ
  32. TIER TWO What holds, and what's contested. WHAT HOLDS WHAT'S

    CONTESTED Retrieval is search-grounded Which index, and when Top-20 visibility feeds the shortlist The grounding budget numbers Your patterns propagate Any number older than a quarter Every engine rents timeliness from an index. The pool retrieval draws from stays shallow. Fan-out returns the description you codified. Blended stacks shift with each partnership. Publicly challenged. Petrovic engages with it. Treat published percentages as dated. https://www.tryprofound.com/blog/ai-search-shift · https://www.seerinteractive.com/insights/87-percent-of-searchgpt-citations-match-bings-top-re FLOQ
  33. TIER TWO Then the Bing number aged out. Microsoft retired

    the Bing API in 2025. OpenAI built its own crawler and index. Profound measured Bing alignment collapsing from 26% to 8% while Google alignment roughly tripled. • 87% overlap in 2024 • API retired in 2025 • 8% alignment by 2026 https://learn.microsoft.com/en-us/lifecycle/announcements/bing-search-api-retirement · https://www.tryprofound.com/blog/ai-search-shift FLOQ
  34. TIER TWO “Rank in Bing, show up in ChatGPT.” The

    surest fact of 2024 didn't survive to 2026. FLOQ
  35. TIER TWO ChatGPT's index has a name. Labrador: OpenAI's own

    index. A September config leak enumerated 69 retrieval engines, 47 of them Labrador's: web, PDF, news by recency, arXiv, Reddit, legal, medical, finance, local. • Sworn testimony: building since 2023 • 69 engines in the registry, 47 Labrador • Fan-out routes to vertical engines https://www.linkedin.com/pulse/complete-list-every-search-engine-hidden-inside-leak-david-konitzny-uwofe/ · https://www.justice.gov/atr/media/1402141/dl FLOQ
  36. TIER TWO Eight providers and a cache. Around it: Bright

    Data and Oxylabs scraping Google in, Yelp and TripAdvisor licensed, Bing mainly in Deep Research, and Microsoft's Web IQ, which Microsoft confirms powers ChatGPT. • Google: scraped in, confirmed by GSC spikes • What looks like Bing may be Web IQ • Lockdown mode serves cached pages https://peec.ai/blog/chatgpt-built-its-own-search-index · https://www.microsoft.com/en-us/webiq FLOQ
  37. TIER TWO For two months, ChatGPT said exactly where its

    results came from. Then the field disappeared. FLOQ
  38. TIER TWO The reverse-engineering is real work. Dan Petrovic reverse-engineered

    Gemini's grounding: what gets pulled, how much gets read, which sentences survive. Roughly 2,000 words of budget, a third of a page, front-loaded. • A ~2,000-word grounding budget • Roughly 32% of a page used • Write for extraction, or lose the read https://dejan.ai/blog/hacking-gemini/ FLOQ
  39. TIER TWO And it has a shelf life measured in

    weeks. Early August: 5.6 Luna becomes the free default. Single-query fan-outs fall from 94% of prompts to 43.5% overnight. Retrieved sources double. Every prior study describes a different machine. • Single fan-outs: 94% to 43.5% • Sources per prompt: roughly doubled • site: operator: 0.3% to 23% https://www.linkedin.com/pulse/chatgpts-new-default-model-gpt-56-more-retrieval-content-konitzny-8htce/ · https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/ FLOQ
  40. TIER TWO Even the best research disagrees. Pew tracked 69,000

    real searches: clicks fell 15% to 8% under a summary. Ahrefs put the position-one hit at 58%. Semrush found zero-click rates didn't rise. All three are careful. • Pew: clicks nearly halve • Ahrefs: 58% down at position one • Semrush: no change on matched keywords https://ahrefs.com/blog/ai-overviews-reduce-clicks-update/ · https://www.semrush.com/blog/semrush-ai-overviews-study/ FLOQ
  41. TIER TWO Some of it is measurement, not contradiction. Retrieval

    and citation are moving in opposite directions. More pages considered per answer, fewer credited. Half the fights between studies come down to which layer got counted. • Retrieved URLs: rising • Cited domains: falling • Check which one a study measured https://lilyraynyc.substack.com/p/what-we-can-learn-from-evolving-chatgpt FLOQ
  42. IN THEIR WORDS The Descartes reflex: the practiced discipline to

    stop, take apart, and verify before accepting something as true just because it looks true. — Joost de Valk. https://joost.blog/healthy-doubt/ FLOQ
  43. TIER TWO That's not sloppy research. That's what studying a

    moving, personalised, probabilistic system produces. FLOQ
  44. TIER THREE What nobody can tell you. The tier the

    dashboards price as if it does not exist. Independent testing keeps confirming it. No platform documents it. FLOQ
  45. TIER THREE Ask the same question. Count the answers. Probabilistic.

    THE EXPERIMENT THE RESULT 2,961 <1 in 100 runs across three platforms runs returned the same list of brands 600 volunteers, 12 prompts. ChatGPT, Claude, Google AI. Same list in the same order: fewer than 1 in 1,000. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-s hould-take-care-when-tracking-ai-visibility/ FLOQ
  46. TIER THREE The retrieval layer confirms it. Profound tracked 10,000

    prompts for two weeks. The engines don't search for what you typed. They fan out, sometimes a dozen queries, sometimes departing from your intent entirely. • ChatGPT: wide-net researcher • Perplexity: nearly 1:1 with your prompt • Copilot: compresses to a few strings https://www.tryprofound.com/blog/what-ai-engines-actually-search-for FLOQ
  47. TIER THREE It edits your question before answering it. ChatGPT

    appends best-of to a quarter of advice fan-outs. It fetches reviews you never asked for. Fusion scoring picks winners. The question answered is not quite the question asked. • Best-of: added to 24.3% of advice fan-outs • Reviews: fetched unasked • Fusion scoring picks the winners https://peec.ai/blog/patterns-we-see-in-chatgpt-query-fanouts FLOQ
  48. TIER THREE Temperature zero doesn't mean what you think. Same

    prompt, a thousand runs, randomness off: dozens of different answers. Output depends on batch size. Batch size depends on server load. • Non-determinism is the default, at kernel level • Load changes the maths path • The fix: perfect, open-sourced, 61.5% slower. https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/ FLOQ
  49. TIER THREE Citations drift 60% in a month. Profound's own

    research: the citations for a stable prompt set can change by up to 60% in a single month. Not your content. Not your rankings. The pipe itself, moving. • Up to 60% churn, monthly • The vendor's own finding • Stability is the anomaly https://www.tryprofound.com/blog/ai-search-volatility FLOQ
  50. TIER THREE The instability shows up at platform scale. Citations

    fell 86-94% across markets February to April 2026, then mostly rebounded in May. One model swap cut cited domains by a fifth overnight. None of it tracked anything brands did. • US zero-citation rate: 28% to 48% in March • Cited domains: 19 to 15 on one model swap • May: mostly rebounded https://www.seoclarity.net/chatgpt-citation-decline-analysis · https://searchengineland.com/inside-chatgpt-search-web-run-fan-out-queries-ai-visibility-477339 FLOQ
  51. TIER THREE One default swap, a fifth fewer citations. 4

    March 2026: ChatGPT's default moved to GPT-5.3 Instant. Over 14 weeks of 400 daily prompts, domains per answer fell from 19 to 15, URLs from 24 to 19. The crawl thinned in step. • Domains per answer: 19 to 15 • URLs per answer: 24 to 19 • Fewer fetches, not just fewer citations https://think.resoneo.com/chatgpt/5.3-5.4/ FLOQ
  52. TIER THREE Two waves, five markets, nothing brands did. seoClarity

    tracked five markets. Wave one, 8 March. Wave two, 19 April, worse. By end of April, citation volume was down 86 to 94% everywhere. In May, it rebounded toward pre-March levels. • US zero-citation answers: 28% to 48% • Down 86 to 94% by end of April • May: back toward February https://www.seoclarity.net/chatgpt-citation-decline-analysis FLOQ
  53. TIER THREE Two waves, one model swap, zero brand actions.

    Your dashboard filed all of it as your performance. FLOQ
  54. TIER THREE Ask through the API, get a different answer.

    Surfer injected a leaked system prompt into clean API calls: still not the interface. Luxeo found the same across web, API and mobile. Vendors concede it. Nobody documents it. • Web, app, API: different answers • System prompt injected: still different • Even the tracking vendors concede it https://surferseo.com/blog/llm-scraped-ai-answers-vs-api-results · https://luxeo.team/chatgpt-web-vs-api-experiment/ FLOQ
  55. TIER THREE Sometimes it guesses your domain. Wrong. Malte Landwehr

    caught ChatGPT running site: searches against parked domains the brands don't own. Netcraft: roughly a third of AI login links point somewhere the brand doesn't control. • site:census.com. Census isn't there. • A third of login links: wrong • Parked domains are buyable https://www.techradar.com/pro/security/chatgpt-and-other-ai-tools-could-be-putting-users-at-risk-by-getting-company-we b-addresses-wrong FLOQ
  56. TIER THREE And underneath: a training set nobody audits. You

    can't inspect the data or identify the raters. 12,000 people rated Google's AI, 500 words in 15 minutes. Bias analysis covers American English. Their preferences are its ethics. • Training data: uninspectable • Raters: unidentifiable • Governance: shaped by IPO incentives https://floq.co/strategy/what-we-dont-know-about-ai-training/ FLOQ
  57. MEANWHILE The vendor methodology page. Pay no attention to the

    retrieval layer behind the curtain. FLOQ
  58. TIER THREE This is the tier your testing protocol lives

    in. Rigour doesn't change what tier it is. FLOQ
  59. THE REFLECTION So what is the dashboard measuring? Jono Alderson's

    frame is the cleanest: visibility is a reflection of underlying signals across the web. Strengthen the signals, become easier to find. FLOQ
  60. THE REFLECTION Visibility reflects. It doesn't originate. Six capabilities: experience,

    physical availability, mental availability, distinctiveness, reputation, commercial performance. Strengthen them, you surface. Weaken them, nothing fixes it. • The signals are the cause • The score is the shadow • Most of the six aren't SEO's to drive https://www.jonoalderson.com/conjecture/clicks-dont-count/ FLOQ
  61. THE REFLECTION A stable score on an unstable system is

    measuring the reflection, not the thing casting it. FLOQ
  62. THE CALL What survives scrutiny, and what doesn't. METRIC REPEAT

    DRIFT CALL Presence frequency 55-77% stable KEEP Mention rate, with CIs 5 runs bounded KEEP Fan-out presence 5 runs binary KEEP Rank position in AI <1/100 1/1000 LET FADE n = 1 ±60%/mo LET FADE Single-snapshot scores Rose marks the negative call. If a metric can't survive repetition, it doesn't survive the quarter. https://searchengineland.com/make-prompt-tracking-more-accurate-479708 · https://www.tryprofound.com/blog/research FLOQ
  63. THE HONEST VERSION Step zero: classify your prompts. Petrovic built

    an open tool that predicts whether a query triggers grounding. Run your prompt set through it before you track anything. Prompts under the threshold have no citation layer. • Predict grounding, per prompt • Split the set: grounded vs memory • Only the grounded half has citations to win https://dejan.ai/blog/how-google-decides-when-to-use-gemini-grounding-for-user-queries/ ;https://huggingface.co/dejanseo/query-grounding FLOQ
  64. THE HONEST VERSION If you track, track like a pollster.

    Kevin Indig's framework: repeated runs, fixed sampling rules, confidence intervals on every rate you report. A distribution you can measure, not a draw you can chase. • Five runs per prompt, per platform • Mention and citation rate, with intervals • Trend, not position https://searchengineland.com/make-prompt-tracking-more-accurate-479708 FLOQ
  65. THE HONEST VERSION The vendors don't believe each other. Peec's

    own docs: Semrush's shared prompt index produced results inconsistent with independent data, with big month-to-month swings driven by methodology changes, not visibility. • Vendor on vendor • Methodology moves the number • Ask what changed: the market, or the index https://peec.ai/ai-instructions FLOQ
  66. THE HONEST VERSION If the vendor can't walk you through

    the methodology, it's not the tool you need. FLOQ
  67. MEANWHILE Asking the vendor about methodology. Me, asking how the

    score is calculated. The methodology, mid-answer. FLOQ
  68. THE REDIRECT You can't test your way out. You can

    make yourself worth retrieving. FLOQ
  69. THE REDIRECT Everything in tiers one and two points the

    same way. Suganthan Mohanadasan watched the network traffic: ChatGPT names brands in its first search, before reading a single page. The shortlist forms before retrieval, in the model's prior. • Named in the fan-out: cited 68.9% • Merely fetched: cited 2.1% • Rank feeds the pool. Recognition picks it. https://suganthan.com/blog/chatgpt-decides-before-it-searches/ FLOQ
  70. THE REDIRECT Thirty encounters. Roughly what it takes for Google

    to understand who you are. Jes Scholz's framing, and she's right. Inconsistency makes it 40, 60, 80, and one will be Joe Blow's Reddit thread from ten years ago. • Same name, everywhere • Same description, everywhere • One pattern for the machines to return https://searchengineland.com/establish-brand-entity-for-seo-guide-427717 FLOQ
  71. THE REDIRECT You can't fake the encounters. Authoritas manufactured experts

    with 600 press articles. Zero AI visibility. Volume without entity presence or corroboration doesn't register. The machines can tell known from placed. • 600 articles • Zero visibility • Corroboration beats volume https://searchengineland.com/ai-recommendations-inconsistent-fix-469250 FLOQ
  72. THE REDIRECT The structural work. None of it is new.

    All of it is what retrieval rewards. Technical Brand Content Legible to the engines that feed the models. One entity, impossible to misread. Fewer pages, aimed at your expertise. • Per-bot robots.txt policy • • • Sitemaps to spec, JS fallbacks One description, everywhere Consolidate: ~35% uplift, done well • • Tech debt, actually addressed About page as evidence locker Kill invalid informational topics • Knowledge panel, press, Wikipedia • Answer what only you can • https://searchengineland.com/geo-and-seo-how-to-invest-your-time-and-efforts-wisely-461424 · https://floq.co/strategy/content-consolidation-meta-study/ FLOQ
  73. THE REDIRECT Schema, honestly. LLMs strip your markup in training

    and tokenise at retrieval. Schema is not a lever on the model. It's a vehicle for the engines that feed the model. Do it anyway. • Stripped in training • Read by the engines underneath • Worth it for that reason alone https://aws.amazon.com/blogs/machine-learning/an-introduction-to-preparing-your-own-dataset-for-llm-training/ FLOQ
  74. THE REDIRECT The peer-reviewed bit. Princeton and IIT Delhi, KDD

    2024: citations, quotations and statistics lifted visibility 30 to 40% on their benchmark. Keyword stuffing went backwards. Evidence density is the lever. • Cite, quote, quantify: +30 to 40% • Keyword stuffing: minus 8 to 10% • A benchmark, not production. Still. https://arxiv.org/abs/2311.09735 FLOQ
  75. THE REDIRECT The informational padding was never yours to keep.

    88.1% of AI Overview triggers are informational. Chasing head terms padded the graph without building a good business and a good product underneath. AI took the padding. • 88.1% of triggers: informational • Padding, not pipeline • Let it go https://searchengineland.com/google-ai-overviews-13-searches-455057 FLOQ
  76. THE REDIRECT GEO stays 20% work. Most time and money

    stays in traditional SEO and brand. Do the 1 to 5% technical changes deferred for years. Pilot the AI layer. Sell it as non-negative and compounding. • 80: foundations and brand • 20: the systemic schema layer • Baby steps beat big bets https://searchengineland.com/geo-and-seo-how-to-invest-your-time-and-efforts-wisely-461424 FLOQ
  77. THE REDIRECT Reddit: retrieved always, cited almost never. Petrovic: ChatGPT

    discards the Reddit pages it pulls 99% of the time. Reddit still tops citation charts on sheer volume. And it now flags ~25,000 GEO-spam posts a day. • Retrieved relentlessly • ~1% survival, still #1 cited • Don't spam it. They're watching. https://dejan.ai/blog/reddit-ai/ · https://www.emarketer.com/content/reddit-geo-crackdown-could-raise-stakes-ai-visibility-strategies FLOQ
  78. THE INDICATORS Measure what actually moves. Borrowed from disciplines that

    never had clean attribution and invested anyway. Go talk to finance about how they measure TV. FLOQ
  79. ANCHOR ONE Excess share of voice. Binet and Field: share

    of voice above share of market predicts growth. The gap, not the level. Share of search via Trends, or your top 20 SERPs. Pick one. Stick to it. • Track the gap, not the number • Trends or SERP share • Consistency over precision https://www.contagious.com/news-and-views/share-of-search-the-new-most-important-metric-for-brands-google FLOQ
  80. ANCHOR TWO Branded search trend. If the work is building

    anything, branded search grows. Slowly, lumpily. Free in Search Console, legible to the board. Flat after a year means something isn't landing. Worth knowing. • Free • Legible to stakeholders • Honest when it's flat https://floq.co/strategy/metrics-breaking-seo-success-different-now/ FLOQ
  81. ANCHOR THREE Everything underneath the surface. Don't build six dashboards.

    Pick two or three proxies that match what the business is trying to do and watch them quarterly. Know what's under the number when it moves. • Direct demand, next to branded • Reviews: volume, recency, coherence • Close rates, retention, conversion quality https://www.jonoalderson.com/conjecture/clicks-dont-count/ FLOQ
  82. THE FRAME Incremental, against the market. Down 5% in a

    market down 15% is outperformance. Up 3% while a competitor grows 12% is losing. Sales, product and finance all report against the market. We report snapshots. Stop. • Year on year • Against named competitors • Trajectory, not totals https://floq.co/strategy/metrics-breaking-seo-success-different-now/ FLOQ
  83. THE POLITICS Time the change with the money. Someone's bonus

    is tied to the old number. Change mid-year and they'll fight you, however right you are. Run parallel a quarter, cut over at the financial year. You carry the meaning. • Parallel run first • Cut over at year end • Translate, constantly https://floq.co/strategy/metrics-breaking-seo-success-different-now/ FLOQ
  84. THE DASHBOARD Three layers, one story. The board gets one

    trendline and a sentence. Directors get the proxies for their function. Your team gets the messy guts. One dashboard for everyone serves no one. • CxO: excess share of voice, one line • VPs: proxies by function • Team: everything underneath https://portent.com/blog/analytics/marketing-reports.htm FLOQ
  85. THE INDICATORS The platforms leak their own telemetry. Bing Webmaster

    Tools now reports AI citations and the grounding queries Copilot generates. The site: fan-outs surface in Search Console: 197,000 impressions, one click. • Bing: AI Performance Report, Feb 2026 • GSC: site: queries, near-zero CTR • Free, first party, already yours https://otterly.ai/blog/bing-webmaster-tools-ai-performance-report/ · https://lilyraynyc.substack.com/p/what-we-can-learn-from-evolving-chatgpt FLOQ
  86. THE DASHBOARD We track presence like a poll. We invest

    in what makes us retrievable. We report the gap against the market. FLOQ
  87. MONDAY The take-home. MEASURE FIX REPORT Know what you can

    actually know. Make yourself worth retrieving. Movement, against the market. • • Per-bot robots.txt policy • • One entity description, everywhere GSC site: filter + Bing AI report • Share-of-search baseline • Change parallel run to year end Classify prompts by grounding • Five-run fan-out check • Presence frequency, with intervals • Front-load the extractable facts FLOQ
  88. TAKEAWAYS Three things to do next. Swap rank tracking for

    presence frequency at sample size. A poll, not a fact. Move some of the testing budget upstream: entity consistency, tech debt, consolidation. Report movement against the market. The gap, not the snapshot. FLOQ
  89. WHO'S BEHIND FLOQ Amanda King is human. 15+ years in

    SEO. Business and product focused, AI forward. 40+ countries, lived in three. Always learning. Slightly obsessed with tea.