For the complete documentation index, see llms.txt. Every page on this site is also served as Markdown: append `.md` to any URL, or send `Accept: text/markdown`.
Attensira Logo
Attensira
GEO Glossary

Content clusters

A group of pages covering one subject at different depths, linked to a central page, so that a system can see coverage of a topic rather than one page about it.

Karl-Gustav KallasmaaKarl-Gustav Kallasmaa, Founder & CEOLast updated

A content cluster is a group of pages that cover one subject at different depths — a broad central page plus narrower pages answering specific questions about it — connected by internal links so that a system encountering any one of them can see the rest.

The idea is structural, not editorial. Nothing about a cluster changes what a single page says. What it changes is what is visible about a site: whether the subject is represented by one page that mentions many things, or by a set of pages each of which answers one thing completely.

The mechanism

Three components, all of which Google documents as ordinary site-organisation advice rather than as a named technique.

Division of subject. Each page owns one question. This is the part people skip, and it is the part that decides whether the cluster works: pages that overlap compete rather than compound.

Internal links with descriptive anchors. Google describes link text as "the text part of a link that you can see" that "tells users and Google something about the page you're linking to," and notes that links help demonstrate topical knowledge. The anchor is what states the relationship between two pages; "click here" states nothing.

A structure that survives scale. Google's own threshold is explicit: "if you have more than a few thousand URLs on your site, how you organize your content may have effects on how Google crawls and indexes your site." Below that, structure is mostly for readers. Above it, it is also for crawlers — grouping similar content in directories, Google notes, helps it understand how often URLs in a given directory change. See crawling and indexing.

Why it matters when the reader is an agent

A retrieval-based assistant does not read a site. It fetches or selects passages and writes from them, which means the cluster is never quoted as a unit. This is worth stating plainly because it changes the argument for building one.

The benefit is not that the cluster is seen. It is that a cluster forces the property a retrieval system rewards: one page per question, answered in full, on the page. A page covering five subtopics has a diluted meaning, whether that meaning is expressed as a vector for semantic matching or as a human judgement about what the page is for. Splitting it produces pages whose entire content is on-topic for one query.

The second benefit is corroboration. When a question about your subject has an obvious page, the same claim appears in a consistent form across your site rather than in five slightly different phrasings. Inconsistent restatements of your own facts are a common source of an assistant getting your details subtly wrong.

Failure modes

  • Cannibalisation. Two pages written for near-identical questions. Both are weaker than the merged version would have been, and Google's crawl guidance is direct about the cost: "eliminate duplicate content to focus crawling on unique content rather than unique URLs."
  • The hub with nothing under it. A pillar page linking to four thin pages is one page and four liabilities.
  • Cluster as a publishing quota. Deciding the shape first and finding twelve subtopics to fill it produces pages that fail the substitution test — swap the subject noun and the page still reads fine, which means it is about nothing in particular.
  • Links only downward. Pillar-to-spoke links without spoke-to-spoke and spoke-to-pillar links leave the narrow pages orphaned from each other, which is where readers and crawlers actually travel.
  • Structure without depth. A tidy hierarchy of shallow pages is still a set of shallow pages. Coverage is the claim a cluster makes; thin pages make it falsely.

How to act on it

  1. Write down the distinct questions first. If two of them have the same answer, they are one page.
  2. Give each page a title that states its question, and a first sentence that answers it.
  3. Link between siblings, not just from the hub, and make every anchor describe the destination.
  4. Use descriptive directory paths once the site is large enough for it to matter; Google recommends words in URLs over random identifiers.
  5. Audit for overlap on a cadence. Clusters decay by accretion — the sixth page about the same thing is always added for a good local reason. See content authority.

Frequently asked questions

What makes a group of pages a cluster?

Descriptive internal links plus a clean division of subject — one question per page.

At what size does site structure start to matter to crawlers?

Google names "more than a few thousand URLs" as the point where organisation affects crawling and indexing.

Do clusters help with AI assistants?

Indirectly. The cluster is not retrieved as a unit, but it forces narrow, self-contained pages, which is what passage retrieval selects.

What is the main risk?

Overlap. Duplicate pages waste crawl and compete with each other.

Frequently Asked Questions about Content clusters

Links with descriptive anchor text between them, and a division of subject so that each page answers a different question. Google describes link text as "the text part of a link that you can see" that "tells users and Google something about the page you're linking to" — the connective tissue is the anchor, not the directory name.

Both, at different scales. Google states that "if you have more than a few thousand URLs on your site, how you organize your content may have effects on how Google crawls and indexes your site," and recommends grouping similar content in directories with descriptive words in the URL rather than random identifiers. Below that scale the links carry most of the weight.

As many as there are genuinely distinct questions, and no more. The failure test is substitution: if swapping the subject noun in one page produces a valid page about something else, that page has no reason to exist separately and should be a section of another.

Indirectly, and by a different mechanism than in search. A retrieval system usually selects a passage, not a site, so the cluster does not get quoted as a unit. What it buys is that whichever page is retrieved is narrow, self-contained, and answers exactly the question asked, instead of burying that answer inside a page about five things.

Yes, when the pages overlap. Google's crawl budget guidance says to "eliminate duplicate content to focus crawling on unique content rather than unique URLs," and warns that time spent crawling URLs it should not crawl means crawlers "might not explore the rest of your site." Six near-identical pages are worse than one good one.
Share this term

Track how your brand shows up in ChatGPT, Claude, and Google AI

Attensira monitors your visibility across AI search platforms so you know exactly when and how you're being recommended.