Short answer
Keyword clustering is splitting a collected keyword list into groups so that each group can be covered by a single page. You group by intent overlap, not by word similarity: if the person typing query A and the person typing query B want the same thing, those queries belong on one page.
The practical test is the SERP. Enter both queries in Google and compare the first ten results: a noticeable overlap means Google treats them as one topic, so one page can rank for both. Different results mean different clusters and different pages.
Why group at all instead of writing a page per keyword
The "one page per keyword" idea sounds logical but falls apart in practice: any service has dozens of close phrasings, and a site of a hundred nearly identical pages gains nothing. The opposite happens instead — the pages start competing with each other, and the search engine picks one of them, rarely the one you would have picked.
The other extreme is more common: every keyword in the topic is piled onto a single service page in the hope that it will collect everything. That does not work either — the page becomes vague, answers none of the questions precisely, and loses to competitors who have a dedicated answer for each intent.
Clustering is the search for the middle ground between these two. This is the step where keyword research stops being a spreadsheet and becomes a site structure: you can see how many pages you need, what each of them is about and what is missing.
What to group keywords by
There are three approaches and they produce different results. I use all three, but in a specific order.
| Approach | How it works | Where it fails |
|---|---|---|
| By words | Group phrases that share words | Separates synonyms ("apartment renovation" and "apartment refurbishment") and merges different intents |
| By SERP | Compare the top 10 for each query; overlapping URLs = one cluster | Needs SERP data; the top is unstable for rare queries |
| By intent, manually | Look at what the person wants to get | Slow, depends on understanding the niche |
The working sequence is this. First a draft grouping by SERP overlap: it is objective because it relies not on my opinion but on how Google has already divided the topics. Then a manual check of the boundaries: automation does not know that two different terms mean the same service in your niche, or that the same term means different things for you and for a competitor. Grouping purely by shared words is only good for the initial sort of a raw list.
The overlap threshold is a setting, not a truth. The stricter it is, the more small clusters and pages you get. On small sites I deliberately use a loose threshold: one strong page beats three thin ones.
One cluster = one page = one intent
The rule is simple and broken constantly. All keywords inside a cluster must share one intent. A classic case of mixing:
- "buy kitchen table" — the person is ready to buy;
- "how to choose a kitchen table" — the person is still comparing;
- "standard kitchen table dimensions" — the person needs one factual answer.
These are three clusters and three pages: a category or product page, a guide article and a reference piece. Merging them into one page produces text that is too long for a buyer and too promotional for someone still deciding.
Commercial modifiers deserve a separate check. "SEO services" and "SEO services pricing" are almost always one cluster, while "SEO services" and "what is SEO" are not, even though they share words. That is why I never rely on shared words where intent decides.
How many keywords one page can target
The answer clients dislike: as many as the meaning allows. A cluster is defined by the intent boundary, not by a count.
In practice the distribution looks roughly like this:
- A narrow service — 3–10 phrasings: the query itself, synonyms, the city variant, the "price" and "order" variants.
- A service section or category — dozens of phrasings with modifiers.
- An e-commerce category with filters — hundreds, including "brand + type + attribute" combinations. A separate task starts here: which filter combinations become indexable pages and which stay closed.
- An informational article — one question and its rephrasings, plus long-tail for the FAQ.
What is worth controlling is not the size of a cluster but its consistency. If the group contains phrases the page would have to answer in different ways, size no longer matters: split it.
When to split a cluster and when to merge
The signals I use on live projects:
Split when:
- commercial and informational intent are mixed in the group;
- the SERPs for queries inside the group barely overlap;
- the page is already written but stalls on page three or four for part of the queries while ranking fine for the rest;
- Search Console shows queries with clearly different expectations hitting the same page.
Merge when:
- two pages fight over the same query and both rank unstably;
- the pages differ only by a synonym in the title;
- the site is small and splitting leaves you with two thin pages instead of one solid one.
The second case is cannibalisation. It is expensive: impressions get divided between the pages, internal links get diluted, and both versions stay weaker than one page would have been. Often the problem dates back to a time when the structure was built without keyword research and pages were added as ideas came up.
What I see on projects
An honest limit first: there is no universal figure like "correct clustering gives +N%" — too much depends on the niche and the state of the site. But the recurring patterns are worth naming.
First: on sites where the structure was built from inside the company, clustering almost always reduces the number of pages rather than increasing it. Half the sections turn out to be named after internal terms nobody searches for, while half of the real demand is not covered at all.
Second: the most expensive mistakes happen not in small clusters but at the top level — when two major landing pages are built for the same intent. It is not obvious at first, because both rank somehow, and the problem surfaces months later.
Third: clustering is not a result yet. Even a perfectly grouped keyword set gives nothing if the pages are not indexed or the site is technically broken. That is why right after grouping I check why Google is not indexing the pages, and only then move on to content.
Clustering on a multilingual site
In Estonia this is separate work, not a translation. Estonian and Russian search for the same topic means different competitors, different phrasings and a different number of queries. Clusters overlap partially: some groups match, some break apart differently, and a few exist in one language only.
Transferring the grouping from one language to another mechanically is a typical mistake. The result is a structure built for demand that does not exist in that language. The right way is to collect keywords for each language separately and redraw the cluster boundaries against that language's SERP. The final section tree usually ends up similar: the languages differ, the business does not. I covered the choice of the priority language in more detail in the article on which language to target for SEO in Estonia.
How to cluster keywords yourself: the sequence
- Collect the raw keyword set and clean out competitor brands and irrelevant phrases.
- Make a draft grouping — by shared words or with an automated tool. This is only a blank.
- Check the doubtful pairs against the SERP. Unsure whether it is one cluster or two — compare the top 10 for both queries.
- Label the intent of each group: commercial, informational, navigational.
- Map clusters to pages — existing and missing ones. One group, one page, no exceptions.
- Find the overlaps — two pages for one cluster. Decide: merge with a redirect, or separate by meaning.
- Prioritise — start with the commercial clusters closest to an enquiry.
After this you no longer have a keyword table but a task list: which pages to create, which to rewrite, which to merge. If the site is still being designed, build this map into the section structure during web development — changing URLs later costs more. And if the structure already exists and you suspect it was not built around demand, it makes sense to start SEO with rebuilding the keyword set.
