Getting cited by AI search.
How AI Overviews, ChatGPT, and Perplexity retrieve sources, and what structured data really does.
No statistics on this page, deliberately.
There are no sourced numbers for AI citation rates in roofing, so none appear here. Any figure claiming that some percentage of AI answers cite a given source type would be invented, and this site does not publish invented figures. What follows is mechanism, which is checkable, and it is the only honest way to write about this pillar today.
How the retrieval actually works.
Every answer engine has to point at something. The routes differ: one reads its own index, others fetch pages at query time through their own crawlers and, for some queries, a third-party index. The requirement does not differ. The answer has to be sitting in readable markup under a heading that matches the question, rather than assembling by script once the page has loaded.
What that changes about writing a page.
What gets ranked is no longer the page, it is the passage. One question per page, the answer in the first sentence beneath a heading that matches it, and no ambiguity about which company is answering. What that costs is the standard roofing page, where the answer sits three paragraphs below the positioning.
What makes a company nameable.
No property on the list guarantees a mention, and a site missing any of them is hard to mention at all. What the company does and where it works has to be legible in the first heading. The organisation has to resolve to one consistent entity across the site, the schema, and every external citation. Pages have to answer questions rather than list capabilities. And the whole thing has to render without JavaScript for crawlers that are not Googlebot, because several of these systems fetch pages themselves.
What structured data does and does not do.
Schema removes guesswork about who is answering, what they do, and where they do it. It does not force a citation and it is not a ranking lever on its own. The types that matter for roofing are an organisation block, Service with an Offer per service, FAQPage for real question sections, Article on guide pages, and BreadcrumbList for hierarchy. Every claim in the schema should be visible on the page itself, because schema that contradicts the page undermines the attribution it was meant to establish.
llms.txt, honestly.
A root-level text file describing a site's important content for language-model consumers. It is a proposal, not a standard, support varies by system, and Google is not among the systems it moves. Ship it, because it takes an afternoon and some retrieval systems do read it. Judge anyone selling it as an AI-visibility strategy accordingly.
What carries over from ordinary SEO.
Almost all of the technical foundation: crawlability, speed, clear information architecture, genuine topical depth, and consistent local signals. Anyone claiming AI search makes SEO obsolete is describing a different internet from the one your enquiries come from. Three things are genuinely different: format, attribution, and crawler coverage. That is the whole list, and it is enough to change how a site is built.
Spoken queries are the practical version of this.
Typed queries compress to roof repair cost. Spoken ones expand to how much does it cost to repair a small leak on a tile roof, which contains the roof type, the severity, and the intent. The page that wins that answers a narrower question than any general service page does, which is the same discipline that makes a site extractable. The two efforts are one effort.
What we can and cannot promise.
Nobody can guarantee AI citations, and anyone who does is guessing. What can be done is making the site extractable, correctly attributed, fast without scripts, and genuinely the best available answer to specific questions your market asks. A Paid Discovery audit reports where your site stands on each of those today, and which questions in your market have no adequate answer published by anyone yet.
Why extraction favours plain markup.
A retrieval system reads a page and looks for a passage that answers the query. A page whose answer sits in the first sentence under a heading matching the question is trivial to extract. A page whose content assembles after a framework boots, or whose answer arrives in paragraph four after three paragraphs of positioning, is not. That is the entire practical difference, and it is a writing and build decision rather than a technical trick.
Entity clarity, in concrete terms.
Five things that make it unambiguous who is answering.
Crawler coverage is now part of the job.
Google's AI Overviews draw on Google's index, but ChatGPT browsing and Perplexity fetch pages themselves, and some systems lean on third-party indexes. A site reachable only by Googlebot is invisible to part of that landscape. Check your robots rules, make sure nothing important sits behind JavaScript-only navigation, and treat crawlability as a plural problem rather than one about Google alone.
What answer-shaped content looks like.
One question per page, stated in the heading roughly as a person would ask it, answered in the first two sentences, then explained underneath for anyone who wants the detail. Ranges rather than refusals on price questions. Process rather than legal advice on claims questions. This is not a format trick. It is what a page looks like when it was written to be useful rather than to rank.
Where roofing companies have an unusual advantage.
The questions are specific, local, and largely unanswered. Nobody has published a good answer to what a hail claim looks like in a particular state, or how a tile roof repair differs from a shingle one on a two-storey home, or what a TPO specification should contain for a warehouse of a given age. Volume-led content strategies skip these because the search volume looks small. Answer engines do not care about volume. They need an answer to attribute.
How to tell whether any of this is working.
Ask the systems directly and record what they say. Run the queries your buyers use in AI Overviews, ChatGPT, and Perplexity, note which companies get named and which pages get cited, and repeat monthly. It is manual and imperfect, and it is considerably more honest than a dashboard metric nobody can define. It is also exactly what the citation checks in an audit and a post-launch crawl do.
What would make us wrong.
A decisive move toward licensed data partnerships and away from open crawling. In that world, being licensed matters more than being extractable, and a contractor's own site loses leverage. It is a genuine possibility and worth watching. The hedge is that clear answers, clean markup and consistent entities are also what make a site rank and convert, so being early costs close to nothing.
Every guide in this pillar.
Two acronyms for the same shift: being the source an answer engine quotes rather than a link it lists.
The retrieval and attribution mechanism, described without invented statistics.
Nine checks, each verifiable before launch, and the two that fail most often.
A young convention, cheap to ship, and not a ranking lever. Here is what it is actually for.
An honest comparison for a roofer deciding where to spend attention.
Conversational queries are longer and more specific. That changes what a page has to answer.
Questions about this pillar.
Can you guarantee AI citations?+
No, and anyone who does is guessing. What can be done is making the site extractable, correctly attributed, and genuinely the best answer to specific questions.
Is this different from SEO?+
It overlaps heavily. The real differences are format, answers stated plainly. Attribution, entity clarity about who is answering. And crawler coverage, since several systems retrieve pages themselves.
Does llms.txt matter?+
It is cheap and harmless, and some retrieval systems use it. It has limited effect on Google specifically and is not a substitute for pages that answer questions.
Why no statistics in this cluster?+
Because no sourced figures exist for AI citation rates in this trade. Publishing an invented percentage would undermine every sourced number elsewhere on this site.
Should we block AI crawlers instead?+
That is a legitimate choice for a publisher selling content. For a roofing contractor whose goal is to be recommended, blocking the systems that recommend you is the wrong end of the trade.
Find out if your site is citable.
$2,500, and you see every answer a competitor is cited for and you are not.