Groups up to 10,000 queries by similarity right in your browser — nothing is uploaded to a server. You can paste in search volume and export the result to CSV.
keyword or keyword; volume (separator —
comma, semicolon or tab). 0 keywords
The tool splits a list of search queries into groups based on shared words. This helps you distribute keywords across pages: one cluster usually maps to one page or section. It also shows the main topics in a large keyword export and the queries that don't group with anything.
The tool takes 2 to 10,000 unique queries, one per line. You can add search volume after a semicolon, comma or tab: iphone 15 price; 8100. Write volume without thousands separators: in query; 5,400 the comma splits the columns and the volume becomes 5. For the same reason, a query itself must not contain a comma.
Clustering takes four steps:
Case-only duplicates are dropped, keeping the first. A new “Min. cluster size” takes effect the next time you click “Cluster”.
Grouping has two stages. First, each query becomes a set of words, and then the sets are compared. This runs in a background browser thread (a Web Worker), so the page does not freeze.
The tool removes punctuation, one-letter words and stop words. Stop words include function words in Ukrainian, Russian and English, plus commercial words like buy and price. That is why “buy iphone 15” and “iphone 15 price” look identical to the tool.
The “Ignore word endings” option trims typical endings from words of 6 or more letters. English words lose only a final s, so “laptops” becomes “laptop”.
Similarity is the share of common words among all words of both queries. This is the Jaccard coefficient: for “iphone 15 pro” and “iphone 15” it is 2/3, or about 0.67. If the similarity reaches the threshold, the queries join one group.
Groups merge in a chain: when A is similar to B and B to C, all three end up together, even if A and C are barely alike. This is single-linkage clustering, and at a low threshold it makes clusters grow fast. For speed, each query is compared with about 400 candidates that share its rarest word.
A summary above the clusters shows the number of clusters, keywords and unclustered queries, plus total volume if you provided it. Clusters are sorted by total volume, and by the number of queries when volume is equal. A cluster's name is its one or two most frequent words. These words are processed, so they may be lowercase or shortened.
Groups smaller than “Min. cluster size” are collected in the collapsed “No cluster (singletons)” block. The keyword-clusters.csv file has the columns keyword, cluster, cluster_size and volume. For unclustered queries, the cluster column says “(no cluster)”.
Your queries never leave the browser. The tool ignores Google results and does not recognise synonyms or transliteration: “sneakers” and “trainers” are different words to it. Words found in more than 35% of the queries, and in at least five, are not used to find candidates. A query made only of such words, like “iphone” in a list about iPhones, stays unclustered.
The candidate limit may also miss some similar pairs in large lists. Review the result by hand before you distribute queries across pages.
Encyclopedia articles describe the similarity formula and the chaining effect. MDN documents background processing.