WEB150M+ pages · canonical-aware

Online duplicate content checker.

Find duplicate content across the live web, your own folders, and your team's content library. Canonical-aware — distinguishes legitimate syndication from scraped copies. Free up to 2,500 words; folder-level scans on Pro.

0 / 1,500 words
Web + Common Crawl indexedNear-duplicate + spun-content detectionFolder-level library scan
DUPLICATE TYPES

Four flavours of duplicate. Different treatment, different fix.

Most “dupe checkers” only catch exact matches. We classify each match into one of four categories so you know which to ignore (legitimate quote), which to canonicalize (syndication), and which to rewrite (scraped or spun).

01 · EXACT

Word-for-word match

Identical strings of 30+ characters. Includes copy-pasted boilerplate, syndicated articles republished verbatim, scraped pages.

Fix: canonicalize, rewrite, or accept (if it's your own syndicated piece).
02 · NEAR-DUP

Lightly reworded

Sentence shape preserved, swapped synonyms or shuffled adjectives. The classic “spin” output. Catches what Copyscape misses by design.

Fix: full rewrite — the underlying argument needs to be yours.
03 · SPUN

Aggressively paraphrased

Sentence-level paraphrase with reordered clauses, swapped tenses, inserted filler. Common from AI-rewriting tools and outsourced content mills.

Fix: verify intent. Is this the writer's own paraphrase or a hidden copy?
04 · BOILERPLATE

Cross-page recurring

The same disclaimer, author bio, or product description appearing on hundreds of your own pages — Google's actual duplicate-content concern.

Fix: extract to a shared template; don't have Google index it 800 times.
WHAT GOOGLE ACTUALLY COUNTS AS DUPLICATE

It's not just “copy-paste from another site.”

Google's actual rules (from Search Central docs): three patterns trigger duplicate-content treatment, and most of them are about your own pages, not about scraped competitors.

01

Cross-domain copies without canonical

Same article appearing on multiple sites without a rel=canonical pointing back to the original. Includes syndicated content distributed wrong, scraped competitors not declared, and accidental cross-site duplication when one writer publishes to two outlets.

02

Same content under multiple URLs on your own site

/product/foo and /product/foo?utm_source=email and /product/foo/ trailing-slash. Faceted nav (filter combinations producing identical pages), printer-friendly versions without canonical, mobile/desktop separate URLs. Google groups them, picks one, drops the rest.

03

Boilerplate masquerading as content

An e-commerce category page with five paragraphs of identical category-description copy across 1,200 SKUs. A locations page with the same “Our team has 20 years of experience” copy across 47 city pages. Thin-content duplicate — Google's Helpful Content Update flags this aggressively.

SEO USE CASES

Six teams that run this check weekly.

Same tool, six different workflows — from a solo blogger pre-publish check to an agency QA pass on outsourced articles.

Blog post pre-publish

Run the draft against the live web before hitting publish. Catches accidental reuse of phrases from your own past posts (self-cannibalization), competitor articles you read while researching, and any AI-rewriter slip-ups.

Competitor scraping detection

Paste your article, see who else is publishing it. Useful for DMCA documentation — every match returns the source URL, the matched intervals, and a timestamped snapshot you can save to your folder as evidence.

Content library hygiene

Pro-tier folder-level scan. Drop your whole content team's published library into a folder; the tool flags cross-document duplication — same paragraph appearing in three different articles, boilerplate that should be templated.

Syndication audit

You syndicated an article to Medium, LinkedIn, and Substack. Are they all pointing rel=canonical back to your origin? The tool fetches each, checks the canonical tag, and tells you which republished versions are leaking SEO juice.

Outsourced writer QA

Your freelance writer's deliverable in one pane. Run it before paying. Catches the unfortunately-common scrape-then-spin pattern, the AI-rewriter signature, and quotes from sources without the citation that should have been there.

AI-content provenance

AI assistants pull text from real web sources during their context lookup. The tool surfaces those reused passages so you can decide: cite the source, rewrite, or kill the paragraph. Cleans up Helpful-Content-Update risk.

FOLDER-LEVEL DUPE SCAN

Drop your library in. See what overlaps with what.

Pro-tier feature. Pick a folder of published articles; the engine builds a similarity matrix across every pair. Sortable by overlap, jump-clickable to the offending intervals.

FOLDER / blog-2026-q16 articles · pairwise similarity matrix
3 cross-overlaps > 10%Export matrix
ARTICLEA1A2A3A4A5A6
A1 — Five SaaS pricing models·4%3%38%6%2%
A2 — How to write SaaS pricing4%·5%12%7%3%
A3 — Product-led pricing guide3%5%·8%6%4%
A4 — SaaS pricing playbook38%12%8%·9%5%
A5 — Choosing your SaaS price6%7%6%9%·4%
A6 — SaaS pricing for founders2%3%4%5%4%·
Click any cell to see the overlapping intervals between the two articles
  • 150M+web pages indexed
  • 30dcorpus refresh cadence
  • 94%canonical detection rate
  • 4dupe-classification tiers
  • 2014indexing Common Crawl
VS BASIC DUPE TOOLS

Three things browser dupe-checkers and SEO suites won't catch.

COPYSCAPE-STYLE TOOLS

Exact-match only (Copyscape standard). A reworded sentence reads as “unique” — and “spun” output passes clean. Useful for clipboard scrapes; useless for content QA.

Noplag: 4-tier classification (Exact / Near-dup / Spun / Boilerplate) with intent hints — actionable, not just a Boolean match flag.

SEO SUITES

Bundle a content-uniqueness tool with your $99-499/mo seat. Limited to surface-level n-gram matching; no Common Crawl coverage; no folder-level cross-document scan.

Noplag: standalone Free tier (2,500 words). Pro ($23/mo) adds folder-level matrix and full 150M+ web index — without locking the dupe tool to a $400 SEO bundle.

FREE ONLINE CHECKERS

Black box. No way to see WHICH source matched, just a similarity number. Often capped at 1,000 words. Most don't refresh their index — checking against a 2021 web snapshot.

Noplag: every match shows source URL + canonical + matched intervals + corpus snapshot date. Apache 2.0 — read the matching algorithm yourself.

CANONICAL-AWARE DETECTION

A reposted article with rel=canonical isn't “duplicate content.” We treat it that way.

When an article appears on multiple URLs, the SEO question isn't “are they identical?” — it's “do they all declare the same canonical?” If your guest post on Medium points rel=canonical back to your origin URL, Google groups them, attributes ranking to the origin, and treats the rest as legitimate syndication. If it doesn't, you're competing with your own re-published content. The tool fetches every match, parses its <link rel="canonical"> tag, and labels each result accordingly: ORIGIN, SYNDICATED-OK, SYNDICATED-LEAK, or SCRAPED.

Read the canonical logic
FAQ

The questions a content lead asks before integrating this.

What's the web coverage like vs Copyscape or SEMrush?

150M+ pages indexed from Common Crawl + curated SEO-relevant high-traffic domains, refreshed every 30 days. Wider than Copyscape's exact-match-only index (~120M, focused on plagiarism boilerplate), narrower than Google's full crawl but with documented coverage and a published refresh cadence. The Pro tier adds your own folder as a private corpus.

How does the canonical detection actually work?

For every matched page, the engine fetches the source URL (cached or live), parses the <head>, and extracts rel=canonical. Matches are then labelled ORIGIN (canonical points to itself), SYND-OK (canonical points to a different known page you control), LEAK (canonical missing/elsewhere on your distribution), or SCRAPED (no canonical + non-distribution domain). The classification is shown next to every match, so your dupe number reads honestly.

Does this catch AI-generated content reused from sources?

Yes — AI writing assistants pull text from real web pages during context lookup, and those reused passages show up as Near-dup or Spun matches. We surface them with the source URL so you can decide: cite, rewrite, or kill. AI-content detection (a separate signal) is a different tool that's coming soon — on the roadmap, not live yet — and only the dupe scan is the focus here.

What's the largest document I can scan?

2,500 words on Free, 50,000 on Premium, 250,000+ on Enterprise (or no cap on self-host). For Pro folder-level scans, a single folder can hold up to 2,000 documents; pairwise matrix computed across all pairs in under 5 minutes for that size.

Will this find competitors who scraped my content?

If the scraped copy is in the web index (i.e. on a public, crawlable site), yes — the URL surfaces in the match list with a timestamped snapshot. The snapshot is DMCA-grade evidence: dated, contents-hashed, with the matched intervals highlighted. We don't auto-file DMCAs (that's a legal call), but we give you the paperwork.

Can I exclude my own syndication network from the dupe scan?

Pro tier — add an allowlist of domains you control (medium.com/@you, linkedin.com/in/you, substack.com/p/you). Matches against allowlisted domains tag as SYND-OK and don't count toward your dupe score. The score reflects “unexpected duplication” rather than “any duplication.”

How does folder-level pairwise scan handle a 500-article library?

The retrieval cascade is O(n log n) per pair in practice — winnowing fingerprints filter pair-candidates by 60-bit hashes, then full alignment only runs on candidates. 500 articles → 124,750 pairs scanned in ~4 minutes on a 4-core box. The matrix exports as CSV for whatever downstream pipeline you want to feed it into.

Is the engine actually open-source?

Yes — github.com/NoplagLabs/noplag-engine, Apache 2.0. The chunker, fingerprinter, retrieval cascade, alignment, and dupe-classification logic are all in the repo. The cloud tier adds the corpus + the Common Crawl + Wikipedia mirror; self-host brings your own. Verify the algorithm before you trust the score.

What's the API for content-pipeline integration?

REST endpoint at /v1/checks (Python + Node SDKs). Submit a doc, get a dupe report back synchronously up to 8 seconds, webhook callback above that. Useful for pre-publish CI checks — fail the merge if dupe% exceeds your team's threshold. Full reference at /docs/developers/api.

What about content behind paywalls or auth?

We only crawl what's publicly accessible — Common Crawl, sitemap-declared pages, robots-allowed domains. Paywalled content (NYT, FT, academic journals behind walls) isn't in the index; the engine won't surface those as matches even if your draft reuses a phrase from them. For academic papers, the OpenAlex / CORE indexes (separate corpus) cover the open-access universe.

Scan a draft. Or your whole content library.

Free tier covers 2,500 words against 150M+ web pages. Pro ($23/mo) unlocks folder-level pairwise scans for your published library. Apache 2.0 reference if you want to keep the corpus on your own server.

Duplicate Content Checker — Free Online Tool | Noplag