A growing share of commercial questions now get answered without a click. Someone asks an assistant, gets a synthesised answer with a handful of citations, and either follows one or does not. If your pages are not the ones being cited, you are invisible in that entire interaction.
This is what people mean by generative engine optimisation. Most of it is not new, and none of it is a trick.
How assistants pick sources
Broadly, three things decide it:
- Whether you rank at all. Most assistants retrieve from conventional search results before summarising. Classic SEO is the entry ticket, not a separate discipline.
- Whether your page answers the question directly. A model extracting an answer needs a passage that is the answer, not a page that eventually arrives at one.
- Whether the claim is specific enough to attribute. “We deliver excellent results” cannot be cited. “Merging thin pages usually improves rankings” can.
Write answer-first
The single highest-impact change. Under every heading that poses a question, answer it in the first sentence, then explain.
Compare: “There are many factors that influence WordPress performance, and understanding them requires context…” against “Slow WordPress sites are most often caused by the database, not the front end.” The second can be lifted and attributed. The first cannot.
Be specific, and be willing to be wrong
Vague copy is unciteable copy. Replace adjectives with facts:
- “Fast” → the actual metric, and what it was before
- “Experienced” → years, project count, what kind of work
- “Affordable” → a range, or the model you price on
- “We work with many platforms” → name them
Specific claims are commitments. That is exactly why they carry weight — and why so few pages make them.
Structure that survives extraction
Assume any section could be pulled out on its own, without the rest of the page.
- Descriptive headings that state the topic, not clever ones
- Short paragraphs — one idea each
- Lists and steps where the content is genuinely a list or a sequence
- A summary near the top for anything long
- Tables for comparisons, which models parse very reliably
Structured data that matches the page
Schema helps machines understand what a page is about. Organization, Service, Article, FAQPage and BreadcrumbList all do useful work.
The rule people break: your FAQ schema answers must match the visible text on the page exactly. Schema that says something the page does not say is a violation, and it is the single most common structured-data mistake we find.
Be clear about who you are
Models need to know what entity your site represents before they will attribute anything to it. That means a consistent organisation name everywhere, an About page with real people and real credentials, consistent details across your site and your third-party profiles, and named authors on articles.
Anonymous content from an unclear entity is exactly what a system trying to attribute a claim will skip.
What gets cited most
In our experience, roughly in this order:
- Direct answers to specific questions — the FAQ format works because it matches how people ask
- How-to content with real steps, especially with an order and a reason for it
- Comparisons — “X vs Y”, where the page is honest about both
- Definitions of terms in your field
- Original data, which is the hardest to produce and the most valuable
Sales pages are cited least. Guides are cited most — which is why a services page and a blog do different jobs and you need both.
Measuring it
Attribution is genuinely immature here. What you can do today: watch for referral traffic from assistant domains in analytics, periodically ask the major assistants questions you should be the answer to and see who they cite, and track branded search volume, which tends to rise when you are being mentioned without a link.
A closing caution
None of this is a hack, and anything marketed as one is likely to age badly. The pages that get cited are clear, specific, honest, well-structured and worth citing. That has been good writing advice for a long time — the difference is that now there is a second audience reading it.
