The short answer
Open a filter combination for indexing only when it has a search query of its own, at least 10 products in the listing and its own copy on the page. In practice that means one filter, sometimes two — "leather sofas", "corner sofas". Everything else gets closed off from crawling, because a category with six filters of five values each produces 46,655 combinations of the same product listing, and that is exactly where the crawl budget goes.
How many addresses a catalogue actually creates
The arithmetic matters more than the theory here. Each filter is either unset or holds one of its values, which gives 6 states for 5 values. Six such filters make 6⁶ = 46,656 states, of which 46,655 differ from the clean category. Across 40 categories that is 1.8 million addresses on a 3,000-product store.
Then it multiplies again. ?color=black&size=42 and ?size=42&color=black return the same listing, but to a crawler they are two URLs. With three active filters, permutations give 6 variants of the page; with four, 24. Sorting and list-view switches add their own multipliers.
This is why faceted navigation is the first thing I look at when the Crawl Stats report shows tens of thousands of requests a day while new products take three weeks to reach the index. How that connects to crawling in general is in my article on crawl budget.
What Google itself recommends
This is where the familiar advice parts ways with the documentation. In its guide to crawling faceted navigation (checked 5 October 2026) Google puts robots.txt first: "Oftentimes there's no good reason to allow crawling of filtered items, as it consumes server resources for no or negligible benefit". Filtering via a URL fragment after # is listed as an option with no effect on crawling at all. Canonical and rel="nofollow" are grouped in the same text as methods that are "generally less effective in the long term".
| Method | What it does | What it does not do | When I use it |
|---|---|---|---|
| Disallow in robots.txt | Stops the parameter string from being fetched | Does not remove addresses already in the index | A new catalogue, or facets not yet indexed |
Filtering via #fragment | Creates no separate addresses for the crawler | Gives no landing pages for queries | Utility filters: sorting, view mode, stock |
noindex meta tag | Drops the address from the index on the next crawl | Does not save crawl — the page is still fetched | Cleaning up what is already indexed |
| Canonical to the clean category | Passes signals to the parent page | Is not binding and does not stop crawling | Sorting parameters and UTM, where the URL must work |
That table implies an order which is the one most often broken: noindex first, disallow second. Disallow a page that is already indexed and the crawler will never fetch it, never see the noindex, and the address stays in search as "Indexed, though blocked by robots.txt". I covered that mechanism, and how it differs from consolidation, in my piece on duplicate content.
Google also asks for two technical details that most guides skip: separate parameters with & only, because commas, semicolons and brackets are hard for crawlers to detect, and return a 404 for any combination that produced no products.
The rule of three conditions
I open a filter combination only when all three conditions hold at once: two out of three is not a reason. Out of the same 46,655 combinations, 10 to 20 usually pass.
- It has a query of its own that people actually type — verifiable in Search Console, the Google Ads search terms report and autocomplete. On an Estonian catalogue this is a separate job, because demand is split across languages: how to collect it is in my article on keyword research in Estonian.
- The combination holds at least 10 products, and that number does not drop to zero on normal stock fluctuations. A two-product page will not hold a visitor and will look thin.
- The page has at least two or three paragraphs of its own copy, an H1 matching the query and its own title. If there is nothing to write beyond an auto-generated "Leather sofas — 14 products", it is not a landing page but a copy of the category.
Such pages stop being filters and become part of the structure: they get a link from the menu or the subcategory block, and they go into the sitemap and the internal linking. Anything in the sitemap has to be indexable: a contradiction between sitemap and disallow is more common than it sounds, and against the 50,000-URL limit per sitemap file those blocked addresses also eat space that real products need. How to set it up is in the article on sitemap.xml and robots.txt.
The order I work through a live catalogue
When facets have already multiplied, this is the sequence I follow:
- Export every row from the Page Indexing report with duplicate statuses and with "Indexed, though blocked by robots.txt" — that is the factual picture, not an assumption about what is indexed.
- Group addresses by parameter set rather than by category: typically 80% of the junk comes from two or three parameters such as sorting and stock.
- Pick the combinations that pass the rule of three. On a 3,000-product catalogue that has come out at 15–40 pages for the whole site, not per category.
- Turn the chosen ones into static sections with their own parameter-free addresses and point internal links at them.
- Give everything else a noindex and wait for it to leave the report. On a catalogue of a few tens of thousands of addresses that takes 6 to 10 weeks, and it cannot be rushed — only nudged by submitting small batches for recrawl.
- Only once they are gone from the index, disallow the parameter set in robots.txt so it does not come back.
Step 5 is where projects stall. The client expects filters to be "closed by tomorrow", while two months pass between the fix and a clean report. That expectation is cheaper to discuss before the work starts than to explain in week five.
What I would not do
Two mistakes cost the most, and both look like sensible optimisation.
The first is opening two- and three-filter combinations "just in case", betting on the long tail. On a furniture catalogue I have run since 2017, single-filter pages brought traffic, while two-filter combinations collected a handful of visits a month and took impressions away from the parent category. I closed them, and wrote about that project in more detail in the furniture store case study.
The second is putting a canonical from a facet to the clean category and considering the matter closed. A canonical does not stop crawling: the bot keeps fetching all 46,655 addresses, and Google is free to disagree with the hint. For sorting and UTM parameters it is a working tool; for the facet space it is not.
And an honest boundary: everything above applies to catalogues between 500 and 50,000 products. On a store with 80 items, faceted navigation does not create a problem worth solving — crawling never hits a limit there, and the time is better spent on product pages. Above 100,000 items the opposite task appears: deliberately building tens of thousands of templated landing pages, which is a different approach I wrote about in the article on programmatic SEO.
