The page types I used to ignore now win in LLMs: features, use cases, case studies, comparisons, reviews, integrations, and persona-based 'who we help' pages. They give the LLM something to cite and say 'built for you.' Which are you prioritizing?
Strong opinions, loosely held, and everyone deserves to get paid. The acronym fight isn't about being right - it's that good people work hard to get brands visibility and should make a living. AI search is awesome. So is getting paid for it.
Full episode with @garrettsussman
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
So the implied AEO advice here is to make individual sections of your page useful when read on their own:
> Answer directly within each section. Only selected passages reaches the answering model. Name the subject and include relevant conditions and limitations nearby so replace vague references like “this approach” or “the results above” with explicit descriptions.
> Make important information understandable as plain text. Pages become simplified text during retrieval. Give charts written takeaways and tables clear column labels. Don’t rely on color, icons, or placement to convey meaning. This was a hard learned lesson for me with benchmarks.
> Cover the subquestions behind a decision (e.g. use cases, industries, etc). One question can trigger multiple related searches. Address capabilities, difficult cases, costs, requirements, and alternatives, with a clear answer to each substantive question. This is pretty aligned with the fanout best practice.
> Keep claims beside their evidence. Individual passages are scored, so a claim and its supporting context may become separated. Put the metric, source link, methodology, and caveat together.
> Make metadata specific and consistent. Titles and descriptions should accurately describe the page’s subject and match its content. Metadata surviving retrieval doesn’t establish a ranking benefit.
> Distinguish “retrieved” from “cited.” Many retrieved pages never become citations. Inspect the passages that earn citations and identify which questions or evidence your page leaves unanswered. Super interesting, I'm going to be digging deeper here because my tracking does not reach that level of granularity...yet.
A decent heuristic is: if a section were lifted out of the page, would its meaning, evidence, and limitations still be clear?
What do you want the internet to know about this business?' is a real exercise to run with local clients. Business provides service. Business located in. Business serves these areas. Build the triples, then place them everywhere. Doing this yet?
Full episode with @DarrenShaw_
Everyone has an AI-search theory.
@otterlyAI is actually testing them.
One wild result: publishers with OpenAI licensing deals were cited by ChatGPT ~48% more often.
Plus: fake brands can earn recommendations, schema may not help ChatGPT, and YouTube citations are rising.
ChatGPT retrieval system LEAK alert 💥
I found "something huge" in ChatGPT's server-sent events over the weekend.
Inside the stream is a detailed debug view of ChatGPT's web retrieval system. Which queries it wrote, which engines it called, what came back, scoring objects,what it fetched, how it split each page, and what finally made it into the answer.
I have been pulling this data for a few days. On Saturday I started getting rate limits, so I was probably the most active user of that stream this weekend :)
A few things I can share today:
One question is never one search. Yes, we know there are fanouts. But actually more fanouts behind the scenes!
For a single "best AI visibility tools" prompt, ChatGPT ran 5 search rounds, wrote 18 different queries(hidden queries), made 50 engine calls, pulled 228 results, fetched 223 URLs with selected chunks, and cited 16. You see 16 links.
✍ Let's start today with renderer. How actually ChatGPT uses your page for retrieval.
Your meta tags travel with every result.
There is a separate og_data object in the payload and it is empty on every single result. The raw meta_tags list is what is actually kept.
The page body is not HTML. It is a markdown-like text render. I compared its fingerprints against the common HTML to text parsers. Two-space "* bullet", "* * " for horizontal rules, "# heading" and "--- | ---" table separators with no outer pipes all match the Python html2text library. Turndown, markdownify and Trafilatura each match only one or two of those. So html2text is the strongest candidate, but it is not a stock build.
That render is what gets cut into blocks of roughly 170 words and scored. The model does not see your full page. It sees one to three of those blocks per source.
The images below is not a mockup. Every code block is copied from the stream as is, including our own Peec AI product page.
What I showed above is a small slice. One prompt produces a 150,000-line JSON dump, and the two fields in the image are maybe 2 percent of it.
The same stream also carries:
Every rewritten query, and which of roughly ten internal engines each one was sent to (web, news, Wikipedia, Reddit, arXiv, YouTube, PDF and more)
📍 A per-result score, plus a score object that breaks that score into its components
📍 A per-chunk score for every block of every fetched page, and which blocks were kept for the prompt
📍 A should_fetch decision on each result, with crawl date and publication date
📍 A second ranking pass done in the model's reasoning, where domains are re-ordered before the answer is written
📍 The exact prompt the model receives, with the word budget it is given per source
📍Separate result types for shopping and local queries, with their own fields
If you work on GEO or AI visibility, this is the closest look at ChatGPT's retrieval pipeline I have seen. MORE TO COME.
Follow @DavidKonitzny, @TomekRudzki, @MalteLandwehr, there is a lot more coming THIS WEEK.
@tompeham 's team ran a real schema markup test. Google organic and AI Overviews visibility went up.
ChatGPT and the other AI engines: zero impact, and some hallucinated the exact data sitting in the schema.
New episode out now with Thomas Peham of https://t.co/v3JrjM6usb.
Something I've noticed across our clients is the layers of content approval isn't always tied to company size. The variable is more about how many people get to say "yes, but."
So, how many approval layers does a piece of content pass through before it publishes at your company?
I'm not a huge user of @perplexity_ai, but I just noticed this for the first time ... a "Trusted" badge added to select citations.
Is this new?
cc: @awebranking , @gfiorelli1
AI is making it easier to create the average answer.
Which might make having an actual opinion more valuable than ever.
@LidiaInfanteM on why going against consensus could become a real advantage.
🎧 Listen to the new @Page2Podcast + subscribe.
https://t.co/2XIc3EcBiJ
#SEO
SEO teams keep trying to manufacture content quality. @LidiaInfanteM says it can't be done, only highlighted and enhanced. What's the tactic you've seen actually work for proving expertise on a page?
Our only successful LLM tests so far? Giving more information. The LLM doesn't get bored reading specs... a human does. Doesn't even need to be UGC; manufacturer data, product detail, more exposed info all help. What's working in your LLM tests?
@SEO_Humorist runs grounding fanout queries through the Gemini API and says this behavior barely existed six months ago. What's the strangest thing you've seen an AI system check unprompted?
Also in there: fan-out queries running site: commands on domains that don't exist, and how his team turned that into an expired-domain finder.
Full episode: https://t.co/ILTgLiuoXL
Which of the three above are you still doing anyway?
@MalteLandwehr has tested most of the AI search checklist people are selling right now.
Three of the most popular items have never produced a measurable visibility lift for him.
The thing that does work is boring enough that nobody builds a product around it.
The boring thing that works: describe yourself identically everywhere.
Website: "SEO agency for lawyers."
Founder's LinkedIn: "digital visibility in the jurisdictional industry."
Press release: "search visibility for law offices."
To an LLM, that's three companies.