Content clusters
A group of pages covering one subject at different depths, linked to a central page, so that a system can see coverage of a topic rather than one page about it.
Karl-Gustav Kallasmaa, Founder & CEOLast updated A content cluster is a group of pages that cover one subject at different depths — a broad central page plus narrower pages answering specific questions about it — connected by internal links so that a system encountering any one of them can see the rest.
The idea is structural, not editorial. Nothing about a cluster changes what a single page says. What it changes is what is visible about a site: whether the subject is represented by one page that mentions many things, or by a set of pages each of which answers one thing completely.
The mechanism
Three components, all of which Google documents as ordinary site-organisation advice rather than as a named technique.
Division of subject. Each page owns one question. This is the part people skip, and it is the part that decides whether the cluster works: pages that overlap compete rather than compound.
Internal links with descriptive anchors. Google describes link text as "the text part of a link that you can see" that "tells users and Google something about the page you're linking to," and notes that links help demonstrate topical knowledge. The anchor is what states the relationship between two pages; "click here" states nothing.
A structure that survives scale. Google's own threshold is explicit: "if you have more than a few thousand URLs on your site, how you organize your content may have effects on how Google crawls and indexes your site." Below that, structure is mostly for readers. Above it, it is also for crawlers — grouping similar content in directories, Google notes, helps it understand how often URLs in a given directory change. See crawling and indexing.
Why it matters when the reader is an agent
A retrieval-based assistant does not read a site. It fetches or selects passages and writes from them, which means the cluster is never quoted as a unit. This is worth stating plainly because it changes the argument for building one.
The benefit is not that the cluster is seen. It is that a cluster forces the property a retrieval system rewards: one page per question, answered in full, on the page. A page covering five subtopics has a diluted meaning, whether that meaning is expressed as a vector for semantic matching or as a human judgement about what the page is for. Splitting it produces pages whose entire content is on-topic for one query.
The second benefit is corroboration. When a question about your subject has an obvious page, the same claim appears in a consistent form across your site rather than in five slightly different phrasings. Inconsistent restatements of your own facts are a common source of an assistant getting your details subtly wrong.
Failure modes
- Cannibalisation. Two pages written for near-identical questions. Both are weaker than the merged version would have been, and Google's crawl guidance is direct about the cost: "eliminate duplicate content to focus crawling on unique content rather than unique URLs."
- The hub with nothing under it. A pillar page linking to four thin pages is one page and four liabilities.
- Cluster as a publishing quota. Deciding the shape first and finding twelve subtopics to fill it produces pages that fail the substitution test — swap the subject noun and the page still reads fine, which means it is about nothing in particular.
- Links only downward. Pillar-to-spoke links without spoke-to-spoke and spoke-to-pillar links leave the narrow pages orphaned from each other, which is where readers and crawlers actually travel.
- Structure without depth. A tidy hierarchy of shallow pages is still a set of shallow pages. Coverage is the claim a cluster makes; thin pages make it falsely.
How to act on it
- Write down the distinct questions first. If two of them have the same answer, they are one page.
- Give each page a title that states its question, and a first sentence that answers it.
- Link between siblings, not just from the hub, and make every anchor describe the destination.
- Use descriptive directory paths once the site is large enough for it to matter; Google recommends words in URLs over random identifiers.
- Audit for overlap on a cadence. Clusters decay by accretion — the sixth page about the same thing is always added for a good local reason. See content authority.
Frequently asked questions
What makes a group of pages a cluster?
Descriptive internal links plus a clean division of subject — one question per page.
At what size does site structure start to matter to crawlers?
Google names "more than a few thousand URLs" as the point where organisation affects crawling and indexing.
Do clusters help with AI assistants?
Indirectly. The cluster is not retrieved as a unit, but it forces narrow, self-contained pages, which is what passage retrieval selects.
What is the main risk?
Overlap. Duplicate pages waste crawl and compete with each other.
Terms related to Content clusters
The credibility a specific piece of content and its named creator carry on a specific topic — how raters are told to assess it, and why it is a page-level property rather than a site-wide score.
The two separate stages that decide whether a page can be retrieved at all, and the reason a serving rule on a blocked page is never read.
Retrieval by meaning rather than by matching strings, what it is genuinely better at, and the class of query where it reliably fails.