Online duplicate content checker.
Find duplicate content across the live web, your own folders, and your team's content library. Canonical-aware — distinguishes legitimate syndication from scraped copies. Free up to 2,500 words; folder-level scans on Pro.
Four flavours of duplicate. Different treatment, different fix.
Most “dupe checkers” only catch exact matches. We classify each match into one of four categories so you know which to ignore (legitimate quote), which to canonicalize (syndication), and which to rewrite (scraped or spun).
Word-for-word match
Identical strings of 30+ characters. Includes copy-pasted boilerplate, syndicated articles republished verbatim, scraped pages.
Lightly reworded
Sentence shape preserved, swapped synonyms or shuffled adjectives. The classic “spin” output. Catches what Copyscape misses by design.
Aggressively paraphrased
Sentence-level paraphrase with reordered clauses, swapped tenses, inserted filler. Common from AI-rewriting tools and outsourced content mills.
Cross-page recurring
The same disclaimer, author bio, or product description appearing on hundreds of your own pages — Google's actual duplicate-content concern.
It's not just “copy-paste from another site.”
Google's actual rules (from Search Central docs): three patterns trigger duplicate-content treatment, and most of them are about your own pages, not about scraped competitors.
Cross-domain copies without canonical
Same article appearing on multiple sites without a rel=canonical pointing back to the original. Includes syndicated content distributed wrong, scraped competitors not declared, and accidental cross-site duplication when one writer publishes to two outlets.
Same content under multiple URLs on your own site
/product/foo and /product/foo?utm_source=email and /product/foo/ trailing-slash. Faceted nav (filter combinations producing identical pages), printer-friendly versions without canonical, mobile/desktop separate URLs. Google groups them, picks one, drops the rest.
Boilerplate masquerading as content
An e-commerce category page with five paragraphs of identical category-description copy across 1,200 SKUs. A locations page with the same “Our team has 20 years of experience” copy across 47 city pages. Thin-content duplicate — Google's Helpful Content Update flags this aggressively.
Six teams that run this check weekly.
Same tool, six different workflows — from a solo blogger pre-publish check to an agency QA pass on outsourced articles.
Blog post pre-publish
Run the draft against the live web before hitting publish. Catches accidental reuse of phrases from your own past posts (self-cannibalization), competitor articles you read while researching, and any AI-rewriter slip-ups.
Competitor scraping detection
Paste your article, see who else is publishing it. Useful for DMCA documentation — every match returns the source URL, the matched intervals, and a timestamped snapshot you can save to your folder as evidence.
Content library hygiene
Pro-tier folder-level scan. Drop your whole content team's published library into a folder; the tool flags cross-document duplication — same paragraph appearing in three different articles, boilerplate that should be templated.
Syndication audit
You syndicated an article to Medium, LinkedIn, and Substack. Are they all pointing rel=canonical back to your origin? The tool fetches each, checks the canonical tag, and tells you which republished versions are leaking SEO juice.
Outsourced writer QA
Your freelance writer's deliverable in one pane. Run it before paying. Catches the unfortunately-common scrape-then-spin pattern, the AI-rewriter signature, and quotes from sources without the citation that should have been there.
AI-content provenance
AI assistants pull text from real web sources during their context lookup. The tool surfaces those reused passages so you can decide: cite the source, rewrite, or kill the paragraph. Cleans up Helpful-Content-Update risk.
Drop your library in. See what overlaps with what.
Pro-tier feature. Pick a folder of published articles; the engine builds a similarity matrix across every pair. Sortable by overlap, jump-clickable to the offending intervals.
- 150M+web pages indexed
- 30dcorpus refresh cadence
- 94%canonical detection rate
- 4dupe-classification tiers
- 2014indexing Common Crawl
Three things browser dupe-checkers and SEO suites won't catch.
Exact-match only (Copyscape standard). A reworded sentence reads as “unique” — and “spun” output passes clean. Useful for clipboard scrapes; useless for content QA.
Noplag: 4-tier classification (Exact / Near-dup / Spun / Boilerplate) with intent hints — actionable, not just a Boolean match flag.
Bundle a content-uniqueness tool with your $99-499/mo seat. Limited to surface-level n-gram matching; no Common Crawl coverage; no folder-level cross-document scan.
Noplag: standalone Free tier (2,500 words). Pro ($23/mo) adds folder-level matrix and full 150M+ web index — without locking the dupe tool to a $400 SEO bundle.
Black box. No way to see WHICH source matched, just a similarity number. Often capped at 1,000 words. Most don't refresh their index — checking against a 2021 web snapshot.
Noplag: every match shows source URL + canonical + matched intervals + corpus snapshot date. Apache 2.0 — read the matching algorithm yourself.
A reposted article with rel=canonical isn't “duplicate content.” We treat it that way.
When an article appears on multiple URLs, the SEO question isn't “are they identical?” — it's “do they all declare the same canonical?” If your guest post on Medium points rel=canonical back to your origin URL, Google groups them, attributes ranking to the origin, and treats the rest as legitimate syndication. If it doesn't, you're competing with your own re-published content. The tool fetches every match, parses its <link rel="canonical"> tag, and labels each result accordingly: ORIGIN, SYNDICATED-OK, SYNDICATED-LEAK, or SCRAPED.
The questions a content lead asks before integrating this.
What's the web coverage like vs Copyscape or SEMrush?
150M+ pages indexed from Common Crawl + curated SEO-relevant high-traffic domains, refreshed every 30 days. Wider than Copyscape's exact-match-only index (~120M, focused on plagiarism boilerplate), narrower than Google's full crawl but with documented coverage and a published refresh cadence. The Pro tier adds your own folder as a private corpus.
How does the canonical detection actually work?
For every matched page, the engine fetches the source URL (cached or live), parses the <head>, and extracts rel=canonical. Matches are then labelled ORIGIN (canonical points to itself), SYND-OK (canonical points to a different known page you control), LEAK (canonical missing/elsewhere on your distribution), or SCRAPED (no canonical + non-distribution domain). The classification is shown next to every match, so your dupe number reads honestly.
Does this catch AI-generated content reused from sources?
Yes — AI writing assistants pull text from real web pages during context lookup, and those reused passages show up as Near-dup or Spun matches. We surface them with the source URL so you can decide: cite, rewrite, or kill. AI-content detection (a separate signal) is a different tool that's coming soon — on the roadmap, not live yet — and only the dupe scan is the focus here.
What's the largest document I can scan?
2,500 words on Free, 50,000 on Premium, 250,000+ on Enterprise (or no cap on self-host). For Pro folder-level scans, a single folder can hold up to 2,000 documents; pairwise matrix computed across all pairs in under 5 minutes for that size.
Will this find competitors who scraped my content?
If the scraped copy is in the web index (i.e. on a public, crawlable site), yes — the URL surfaces in the match list with a timestamped snapshot. The snapshot is DMCA-grade evidence: dated, contents-hashed, with the matched intervals highlighted. We don't auto-file DMCAs (that's a legal call), but we give you the paperwork.
Can I exclude my own syndication network from the dupe scan?
Pro tier — add an allowlist of domains you control (medium.com/@you, linkedin.com/in/you, substack.com/p/you). Matches against allowlisted domains tag as SYND-OK and don't count toward your dupe score. The score reflects “unexpected duplication” rather than “any duplication.”
How does folder-level pairwise scan handle a 500-article library?
The retrieval cascade is O(n log n) per pair in practice — winnowing fingerprints filter pair-candidates by 60-bit hashes, then full alignment only runs on candidates. 500 articles → 124,750 pairs scanned in ~4 minutes on a 4-core box. The matrix exports as CSV for whatever downstream pipeline you want to feed it into.
Is the engine actually open-source?
Yes — github.com/NoplagLabs/noplag-engine, Apache 2.0. The chunker, fingerprinter, retrieval cascade, alignment, and dupe-classification logic are all in the repo. The cloud tier adds the corpus + the Common Crawl + Wikipedia mirror; self-host brings your own. Verify the algorithm before you trust the score.
What's the API for content-pipeline integration?
REST endpoint at /v1/checks (Python + Node SDKs). Submit a doc, get a dupe report back synchronously up to 8 seconds, webhook callback above that. Useful for pre-publish CI checks — fail the merge if dupe% exceeds your team's threshold. Full reference at /docs/developers/api.
What about content behind paywalls or auth?
We only crawl what's publicly accessible — Common Crawl, sitemap-declared pages, robots-allowed domains. Paywalled content (NYT, FT, academic journals behind walls) isn't in the index; the engine won't surface those as matches even if your draft reuses a phrase from them. For academic papers, the OpenAlex / CORE indexes (separate corpus) cover the open-access universe.
Scan a draft. Or your whole content library.
Free tier covers 2,500 words against 150M+ web pages. Pro ($23/mo) unlocks folder-level pairwise scans for your published library. Apache 2.0 reference if you want to keep the corpus on your own server.