ShipworkSite inspection
Consolidation

Crawlability

PaginationAre paginated pages set up right?Internal searchIs site search eating crawl budget?URL parametersAre parameters creating duplicates?Log analyserWhat does Googlebot actually crawl?Robots builderNeed a robots.txt file?Soft 404sAre pages “not found” but returning 200?JavaScriptCan crawlers see it without running JS?FreshnessIs my site quietly going stale?Robots.txtDoes robots.txt say what I think?

Indexing

Sitemap lastmodAre my sitemap dates valid?SitemapIs my sitemap actually fetchable?IndexingIs Google allowed to index my pages?CanonicalsIs Google indexing the wrong URL?Index signalsDo sitemap and index signals agree?Redirect builderNeed the redirect rules?RedirectsAre my URLs answering directly, over HTTPS?

On-page

Keyword ideasWhat are people searching for?SERP previewHow does my page look in Google?Content qualityAre your pages too thin or too alike?Image weightAre images slowing the page?CannibalizationAre pages competing with each other?On-page checkDoes the page use its target phrase?ImagesAre my images accessible and loading?DuplicatesDo my pages compete for one query?

Links

Orphan pagesWhich pages can no link reach?Link graphHow deep do your pages sit?Anchor textDo links say what they point at?Outbound linksAre external links still alive?Broken linksAre internal links sending visitors nowhere?

Structured data

Rich resultsIs my markup eligible for a rich result?Structured dataIs my product schema valid?Schema coverageDo my key pages carry structured data?Social previewHow does my page look when shared?Schema vs pageDoes markup match the page price?Schema builderNeed valid JSON-LD?

International

Hreflang sitemapDo page and sitemap hreflang agree?HreflangDo my language versions link back?Hreflang builderNeed the hreflang tags?
Every error, explainedGuides →
Pricing Learn Guides
Sign in

Signing in is optional: it keeps your account, watches and connections on any device. Every free check, the audit and the score work with no account at all.

Sign in with GoogleOpens your account, or creates a free one PricingPlans and credit packs for the paid jobs

Which near-duplicate pages should become one?

The near-duplicate content check reports that pages are about the same thing. That is a symptom a shop owner cannot act on. This check turns the same pages into a decision: which one wins, what to merge, what to redirect, what unique content to save first, and the order to do it in.

Crawls a set of pages on your site and groups the near-duplicates from their text. Nothing is stored.

What this check inspects

Shipwork crawls a set of pages, embeds the text of each, and groups the ones that read as near-duplicates using a strict similarity threshold, deliberately higher than the near-duplicate content check so a merge is only suggested for pages that genuinely repeat each other. From each group it picks a keeper: the page with the most inbound links from the crawl, and where that ties, the most complete content by words, headings, images and structured data. Every other page is given an action, merge or redirect, a list of the unique content to move across first, and a numbered plan ending in a 301 and a sitemap removal. When a search property is connected it corroborates each group with the one fact similarity cannot see: whether the pages both rank for the same query. It also reports cannibalisation, a query two of your own pages both get impressions for, which embeddings cannot detect because the pages read differently.

What a failure means

This check needs a set of comparable pages and a working text model, so with too few readable pages or no model it reports a reason rather than a plan. Redirected and noindex pages are skipped, because a redirected page is a copy of its target and a noindex page is not competing. A group of near-duplicate pages is a warning with a keep, merge and redirect plan. A query two of your own pages both rank for is a separate warning, since one is splitting the signals of the other. Without a connected property the plan is exactly the same; the only loss is the corroborating search evidence and the cannibalisation list.

How to fix it

  1. Keep the page the plan names as the keeper and make sure it carries a self-referencing canonical tag.
  2. Before removing a page, move the unique content the plan lists into the keeper.
  3. 301 the other pages to the keeper and remove them from the sitemap, rather than leaving redirects and dead entries.
  4. Update internal links so they point at the keeper instead of the pages that were merged away.
  5. For a contended query, pick the page you want to rank and give each remaining page a clearly different intent.

A typical failure, worked through

The setupA store has three collection pages that say almost the same thing under slightly different names. All three are in the same crawl, and two of them both appear in search data for one query.

What the check reportsThe check groups the three pages, names the best-linked one as the keeper, lists the unique sections to move across before the others are redirected, and adds that two of the pages both rank for the same query.

The pointA similarity score tells you pages are alike. This turns that into which page to keep, what to rescue first, and the exact sequence of edits, which is what a person needs in order to act.

Questions

How is this different from the near-duplicate content check?
That check reports that pages are about the same thing. This one decides which page should win and gives a merge, redirect and content plan, at a stricter threshold so a merge is only proposed for genuine near-duplicates.
Does it need a search connection?
No. The plan is built from the pages alone. A connection only adds corroboration that two pages really rank for one query, plus the cannibalisation list.
Why is a redirected page skipped?
It lands on the page it points at, so it would always look like a copy of its target and would pollute the group.
Is a merge safe?
The plan tells you which sections are unique and must be moved to the keeper first, and says when a page had too little text to judge, so nothing is deleted blind.

Related checks and guides