What sources does ChatGPT actually cite, and how to find out for your industry

“¿Qué fuentes usa ChatGPT?” The useful version of this question is not global. The sources that carry answers about sunscreen are not the sources that carry answers about ERP software. What you need is the list for your category, in your country, with an owner attached to each source so you know what to do about it. Here is how to get there, and what the list looks like.

How to find out (and the trap in doing it casually)

For questions that search first, ChatGPT shows its sources. So the naive method is: ask a question, read the citations, done. The trap is variability. The same question asked again returns a different answer with different sources. One ask is an anecdote; the source list from one ask is an anecdote too.

The honest method: a bank of real buyer questions for the category, asked several times each across engines, with every cited source recorded and deduplicated. After that, you are not guessing which media outlet to pitch. You know.

The three ownership classes

Once you have the list, every source falls into one of three classes, and the class decides the move:

What the list looks like

From our invented sample, Solaria, a sunscreen brand sold in Mexican pharmacies (competitors Bloqsol and Dermalux):

SourceOwner classMove
profeco.gob.mxCannot edit (regulator)“Aceptar y rodear”: be precise and present around it
pharmacy retailer pageCan rewrite“Este trimestre, sin pedir permiso”: rewrite the listing copy
dermatology reference siteEarned“Editorial y relaciones públicas”: feed the topic with substance
Marca inventada, competidores inventados, cifras inventadas. La estructura es real; los números no son de nadie.

Three sources, three different plays, one afternoon of decisions. That is what the list is for. In a real report, each row also carries its citation share with the denominator, so you know which sources carry weight and which are noise.

What the research says about who writes the answers

The short version: mostly, not the brand. There is a widely quoted analysis in this space reporting that earned media contributes 48% of what assistants cite, third-party commercial content 30%, and the brand’s own site 23%. We have not been able to identify and verify the original study behind those figures, so treat them as directional until we can name the source. The structural conclusion is safe regardless: the brand’s own site is a minority of what gets cited. That is why “fix our website” is not a strategy for AI visibility, and why the source list for your category is the actual map.

The instrument for getting that list: a study with a frozen prompt bank, repetitions, and every response saved and readable. It is the first thing we run: /auditoria-visibilidad-ia/.

Next step: The AI visibility study