Responses, not links: how we measure AI visibility
A search page used to deliver six links, and the person chose which to open. An assistant reads the sources itself, decides what it believes, and answers in a paragraph. We measure what lands in that paragraph: whether your brand appears, whether it gets recommended, whether it is named the best, whether the answer tells the buyer where to purchase it, and which sources make all of that happen.
This page is the method. The studies run on answer.show, our own measurement platform, and the full technical version of this document lives there: https://answer.show/es/geo-measurement/method
Responses, not links
In 2015, a results page delivered six links and position mattered, because there was a list. A model reads the sources, keeps what it considers credible, and writes one paragraph. “Entraron seis fuentes. Salieron dos marcas. Nadie hizo clic en un solo link.” Six sources went in. Two brands came out. Nobody clicked a single link.
Your brand is inside that paragraph or it is not. And the brands left out have no way of even finding out they were left out. That is the problem measurement solves first.
What gets counted
Every response in every run is read and coded the same way. Four outcome rates, measured identically each time:
Mentioned. Does the brand appear in an answer to a question that never names it. The floor: presence in the category conversation. In our invented sample: Mentioned 26% (103 of 396 responses).
Recommended. Does the assistant actively suggest it, or merely list it alongside others. Being listed among eight is not being told to buy it. 11% (44 of 396).
Flagged as the best. Is it distinguished instead of grouped with the rest. The strongest and rarest form of presence. 3% (12 of 396).
With purchase route. Does the answer tell the person where to buy it. The closest honest proxy to purchase intent, with no analytics required. 7% (28 of 396).
These figures come from the sample brand we use to show the report structure, Solaria, a sunscreen sold in Mexican pharmacies:
Marca inventada, competidores inventados, cifras inventadas. La estructura es real; los números no son de nadie.
The percentages change from study to study; the format, with the denominator printed next to every rate, never does.
Every study also records:
- Competitor mentions, which makes share of category conversation possible.
- Sources: every cited site, deduplicated, with its citation rate, ordered by who can actually influence it.
- Negatives and evasive answers, recorded as their own status. “Que un asistente se niegue a recomendar es información, no un dato faltante.” An assistant refusing to recommend is information, not a missing data point. Models that show no sources are recorded the same way, because counting them as zero silently corrupts every comparison.
Optional counters a brand can add: sentiment, alerts when an assistant repeats a specific critique or a stale study, retailers named as purchase channels, whether a legally approved claim is being made on the brand’s behalf, and subbrand attribution.
Why it is not measured from your website
Almost every tool in this category inherited its frame from SEO: it crawls the client’s site and grades its pages. Coherent for companies that sell through their own site. Wrong for anyone who sells through retailers, pharmacies, distributors or advisors.
Follow a real purchase: the buyer asks an assistant, opens a retailer app, reads two reviews, and picks the product off the shelf. “Veces que el comprador visitó el sitio web de la marca: 0.” Times the buyer visited the brand’s website: zero.
So the measurement is built around the category conversation, not the domain. And one rule follows from it: “para saber si tu marca aparece por sí sola, la pregunta no puede mencionarla.” To know whether your brand appears on its own, the question cannot mention it.
The three layers of the prompt bank
The bank holds three layers, and each one answers a different question:
Unbranded. Category questions that never name you. The only honest way to measure spontaneous presence.
Brand-named. What happens once someone names you: what the model believes, where it sends buyers, which critique arrives attached, whether your subbrands get credited correctly.
Competitive. Who owns the category conversation: head-to-head comparisons, and the comparisons happening without you in them.
Each layer gets its own numbers, and they are never averaged. “Mézclalas y cada cifra se mueve por razones que nadie puede explicar.” Mix them and every figure moves for reasons nobody can explain. Bank size lands between 50 and 100 questions, so every layer has cuts big enough to report.
How the questions are written
Start from the category, never the brand. Nobody asks for a category; they ask about a situation inside it. First, list every theme a real buyer brings. For a sunscreen: protecting children, sensitive skin, daily facial use, beach and sport, ingredient safety, price and value, where to buy.
Then give each theme several genuinely distinct angles:
- Who is asking: “bloqueador para bebé de menos de seis meses”
- The situation: “cómo evito que mis hijos se quemen en la playa”
- How far along the decision is: “cuál bloqueador para niños es el mejor”
- The comparison being weighed: “bloqueador mineral o químico para niños”
An assistant answers those four differently. They are not one question reworded.
The themes are not guessed. They come from documented demand: Search Console queries, Google autocomplete, People Also Ask, forums, review platforms, call center FAQs. Every question carries its origin in the deliverable. Questions with documented demand, not questions a consultant invented at a whiteboard.
The three types of assistant
The same question meets three different machines, and they are measured and reported separately, never blended:
Memory-only call. What the model absorbed in training. The closest thing to spontaneous brand awareness, and the hardest to move.
Search-then-answer. The model searches, then answers. What most consumers see today, and the dirtiest to measure, because whether it searched varies between identical questions.
The search engine that answers. Google AI Mode and AI Overviews. The bridge to a search budget someone in the room already owns.
The split is often the most useful table in the report. Same brand, mention rate by assistant type, from the invented Solaria sample: memory-only 9% (12 of 132), search-then-answer 22% (29 of 132), the answering search engine 31% (41 of 132). Same disclaimer as before:
Marca inventada, competidores inventados, cifras inventadas. La estructura es real; los números no son de nadie.
The reading is the point: “el awareness en la memoria del modelo es débil; la web abierta es la que está cargando la marca.” Awareness inside the model’s memory is weak; the open web is what carries the brand. That is why the lever is PR, retail content and reviews, which is exactly the work this measurement aims.
How the study runs
Six phases:
- Levantamiento. One working session to set the scope: does the money arrive through your website or through a shelf, what decision would you make differently based on the result, and who you compete with, including the brands the models treat as interchangeable with yours even if your sales team does not.
- Bank written and reviewed. You review the bank: do the questions sound like your customers, and is any theme missing. Review does not mean rewriting the questions toward the brand, “que es el único cambio que invalidaría el resultado”.
- Freeze and run. The approved bank freezes and becomes the instrument. It runs on the chosen models, with repetitions, and the number of observations and the cost ceiling are known before the first prompt is sent.
- Reading and interpretation. Every response is read and coded identically: mention or recommendation, subbrand attribution, sources ordered by who can actually move them. This is the expensive, judgment-heavy part.
- Report and working session. What moved, what it implies, which team owns each action, so your people can defend the numbers in a meeting where we are not present.
- Second wave. The same frozen bank, asked again after the work had time to publish and index. Same instrument, so the runs compare directly.
Study size is a range, not a package
A study has three knobs: questions x models x repetitions. There is no standard configuration; the right size depends on the decision the study must inform.
The range runs from a small directional sample (a few dozen prompts, fewer models, fewer repetitions, enough to see where you stand) to a deep study (around 100 prompts, 4 or more models, 5 repetitions, enough to compare segments and models with confidence).
One illustrative configuration, labeled as such: 60 questions x 3 assistants x 3 repetitions = 540 responses read. It is an example of the arithmetic, not a default.
Size and price are agreed in the levantamiento, before the first prompt is sent. The scope conversation happens before the money is spent, not after.
Rates, not scores
Every figure in the report is a rate with its denominator printed next to it: Mentioned 26% (103 of 396 responses). “Un porcentaje sin población atrás es un adorno.” A percentage without a population behind it is a decoration.
There are no composite scores. A score hides its inputs; a rate shows its evidence. If another provider hands you a “visibility score” for your brand, ask what population stands behind the number. If there is no answer, it is a decoration too.
Measure twice
Run one is a baseline plus a list of what is citing the competitors. Then the work happens. It is mostly editorial, retail content and PR rather than a website migration. Run two, against the same frozen bank, is what tells you whether it worked.
The honest caveats, kept visible:
- The comparison does not prove causation. A model update or new third-party content can move the same number.
- Questions added later start their own baseline. You never splice banks into a trend.
- If nothing moved, the report says nothing moved.
Five things to unlearn before you decide
Most of the confusion in this market comes from importing habits that made sense for search rankings. Five of them need to go:
- There is no position one. An assistant does not rank brands; it writes an answer. Whoever offers you position one invented it.
- Most of the response was written by someone else. The assistant assembles from sources it trusts, and most of those sources are not yours.
- The same question answers differently each time. One screenshot is an anecdote, and anecdotes fail in both directions: false comfort and false panic.
- A topic can matter before anyone searches it. The questions that trigger your category’s answers are asked every day, with or without search volume behind them.
- The first run is a baseline, not a fix. A first study tells you where you stand and what to do about it. The proof of movement arrives with the second run.
What this does not do
The limits, stated because they convert skepticism into trust:
- No position. “No existe la posición uno; lo que te dé una, se la inventó.” There is no position one; whoever gives you one invented it.
- No sales attribution. The purchase route rate is the honest proxy, not a verified sale.
- No promised movement. Third parties write most of what gets cited, so nobody controls the outcome, including us.
- No fact-checking of assistant claims. If an assistant repeats a wrong claim about your brand, the study records it and shows where it lives. Correcting it is a separate job.
- A first run sets priorities. It does not deliver a proven fix. Whoever promises a verified fix from a single baseline is selling something else.
Start with the question, not the tool
The method only matters if it answers a decision you actually need to make. Tell us your category and the decision, and we will tell you whether a study is the right instrument, and what it should contain either way.