New reverse-engineering research published August 17 by the Resoneo team gives the clearest look yet at what actually happens between ChatGPT finding a page and citing it and the gap between those two steps is bigger than most people assume.
The funnel
Across the research team’s corpus, 61,332 URLs were surfaced into ChatGPT’s sources sidebar pages the system found and considered relevant. Of those, 5,032 became the lead source behind an actual citation. But only 759 pages were fully opened and read by the system, and all of those happened specifically in ChatGPT’s “thinking mode.” That’s roughly 1% of everything surfaced.
Why that 1% matters disproportionately
Here’s the number worth remembering: a page that gets fully opened and read ends up cited 74% of the time. A page that’s merely retrieved but never actually opened gets cited only about 7% of the time. Being read, not just being found, is what actually predicts citation by a factor of roughly 10.
The pool of winners is shrinking, not growing
A separate piece of analysis from the same period found something that reinforces this: the number of unique domains cited per ChatGPT response dropped from 19 to 15 after a recent model update, even as the number of pages retrieved during research grew. More pages get surfaced, but the citations concentrate into fewer domains not a wider spread. Separately, the most consistently cited domains overall remain a familiar, small list: Reddit, Wikipedia, Forbes, Merriam-Webster, Consumer Reports, Healthline, and Walmart.
The honest caveat
The researchers themselves are upfront that they don’t have this fully solved. Exactly what makes ChatGPT decide to open and read one page over another is still partly a hypothesis, not a confirmed rule, OpenAI doesn’t publish that logic, and the researchers note the system keeps getting harder to observe over time as routing and provider details shift. Treat this as the clearest evidence available right now, not a finished formula.
What this actually means practically
The takeaway isn’t “write differently for AI” it’s the same fundamentals this section keeps coming back to, now with a clearer number attached. A page that’s specific, clearly structured, and directly answers a real question gives a retrieval system an actual reason to open it instead of skimming a snippet and moving on. Generic, broad, thin content is exactly the kind of page that gets surfaced into that 99% that never gets read.
Practical takeaway: if your website content answers questions in vague, broad strokes rather than specific, well structured detail, this is a concrete reason to fix that, not because of some new AI trick, but because it’s still true that specific and well-structured content wins, and now there’s a number showing exactly how much it wins by.
Want a plain, honest look at whether your own site’s content is built to be read, not just skimmed?