Plagiarism checker for publishers.
Pre-publish checks for freelancer submissions, syndication audits for republished content, and a REST API your engineering team can drop into your CMS. The detection engine is Apache 2.0 — readable before adoption, not after.
Four desks. One workflow.
Vet a freelance pitch before assignment
10 minutes per submission, before the editing budget commits.
- Paste a draft pitch / clip, see prior published versions
- AI-likelihood per paragraph (coming soon), ESL-aware
- Skip the back-and-forth on “is this original?”
Run the day's queue before press
Bulk-scan today's articles before they hit the CMS publish queue.
- Queue view sorted by similarity + AI score
- Per-section thresholds (op-ed vs news vs review)
- Audit trail by reporter for the standards desk
Investigate a contested submission
Read the algorithm, then read the matched intervals.
- X-Engine-Commit reproducibility
- Override verdict with documented reason
- Apache 2.0 — defensible at corrections-page review
Audit the partners republishing you
Track canonical-tag health across the syndication network.
- Canonical-aware detection (ORIGIN / SYND OK / LEAK)
- Weekly diff report by partner
- DMCA-grade evidence pack on request
Three steps before the article hits the CMS publish queue.
Embedded in your editorial flow via API, run as a bulk pass before press, or one-off in the dashboard. Same engine, same report shape, fits where it fits.
01 · Intake
Submission lands in your CMS queue — drag-drop the .docx / .gdoc / paste, or the CMS pushes via webhook. Author + section + assigned editor metadata travels with the document for the audit log.
02 · Scan + classify
The cascade runs against open web + Common Crawl + (optionally) your own back-catalogue corpus. Citation classification separates direct quotes (with attribution) from uncited paraphrase. AI-likelihood per paragraph is coming soon, ESL-aware. Median 42s per article.
03 · Decide + publish
Editor sees: headline similarity, breakdown by category, sortable source list (AI-score with ESL-adjusted flag is coming soon). Three actions: publish, send back with notes, escalate to standards desk. Decision logged with engine commit + corpus snapshot.
Four CMS surfaces. One REST contract.
Plug Noplag into the CMS your reporters already use. The API surface is the same — what differs is the integration layer (WordPress hooks, Ghost actions, Drupal events, or your own webhooks).
Official plugin from WordPress.org Plugins Directory. Gutenberg-compatible.
- Pre-publish meta-box on every post type
- Block-level highlighting in the editor
- Audit log via WP REST API
Webhook integration + Admin API consumer. Member-content compatible.
- Pre-publish webhook fires on save_draft
- Tag-level scan-policy per content type
- Self-host friendly (Noplag self-host + Ghost self-host)
Contrib module from drupal.org. Field-API-based attach.
- Per-content-type plagiarism field
- Workflow integration with Content Moderation
- Multi-site config supported
Webhooks + OpenAPI 3.0 spec. For headless CMS, in-house systems, or anything else.
- POST /v1/checks from any CMS hook
- HMAC-signed webhook callbacks
- Python + Node SDKs ship from the OpenAPI spec
Today's publish queue. Sortable, audit-logged.
Per-article rows with reporter, section, similarity, AI score, and a click-through to the highlighted report. Sortable. Bulk-approve safe articles, escalate the rest.
See who's republishing you — and whether the canonical tag points home.
Track your distribution network
Add the domains you syndicate to (Medium, LinkedIn, Substack, your network partners) as an allowlist. Weekly crawl reports back which partners have your rel=canonical pointing home (ORIGIN credit), which have it missing (SEO leak), and which are reposting without permission (SCRAPED).
Whose blog is your byline missing from?
Open-web monthly diff scans surface non-allowlisted domains publishing matched intervals of your articles. The match comes with a timestamped snapshot — DMCA-grade evidence dated to the day the scraped copy appeared. We don't file DMCAs (that's a legal call); we give you the paperwork.
Internal duplication across your own pages
Pro folder-level scan: drop your published article corpus into a folder; the engine builds a similarity matrix across every pair. Catches the 47 city-pages with the same boilerplate paragraph — Helpful-Content-Update risk you should know about before Google does.
Three policies your standards desk will want before signing the procurement contract.
First — submitted articles never train a model. Not ours, not anyone else's. Submitted text is fingerprinted, indexed in your tenant, and purged after 30 days unless explicitly retained. Second — the AI verdict (coming soon, on the roadmap) will be per-paragraph, not per-article, and labelled with the engine commit + corpus snapshot timestamp. A contested AI-flag will reproduce months later against the same commit. No silent retraining. Third — Apache 2.0 means your engineering team can clone github.com/NoplagLabs/noplag-engine, run it on your own infrastructure, audit the data flow, and verify the right-to-erasure path end-to-end. Most institutional procurement reviews of plagiarism software stall on those three questions; we tried to answer them upfront.
Read the standards-desk guideWhat the standards desk actually asks before procurement.
- Will our articles be used to train your AI model?
- No. Submitted articles are fingerprinted (winnowing hashes only) and indexed in your tenant-isolated database. The text itself is purged after 30 days unless you opt into retention. The AI detection model — Binoculars-based (Hans et al. 2024), coming soon on the roadmap — will be pre-trained on public corpora; we don't add customer content to it, ever. Explicit in the DPA.
- How will the per-paragraph AI verdict survive a reporter dispute?
- AI detection is coming soon — on the v1.2 roadmap, not live at launch. When it ships, every flagged paragraph will stamp the engine commit and corpus snapshot. Re-run the same paragraph against that commit and the verdict reproduces exactly. Open the algorithm at github.com/NoplagLabs/noplag-engine (Apache 2.0), point to the perplexity threshold, document the ESL-adjusted calibration. We don't have a “trust the score” mode.
- What about our freelancers submitting from low-bandwidth markets?
- Detection is calibrated for English, Spanish, Portuguese, Polish, and Ukrainian at launch, with more languages rolling out. Detection quality is benchmarked on PAN-PC-11; per-language evaluation is on the roadmap. ESL-adjusted AI verdict flags non-native-English writing as ESL ADJ instead of LIKELY-AI when the lexical signature suggests it.
- Can we self-host inside our newsroom infrastructure?
- Yes — Apache 2.0 engine + Docker. Runs entirely inside your network. Pair with your existing Postgres + Redis on your own infra; no submission traffic leaves your VPC. Useful for newsrooms with strict source-protection requirements or pre-publication confidentiality.
- Does it integrate with our CMS?
- WordPress (official plugin), Ghost (Admin API + webhooks), Drupal (contrib module). For anything else — custom CMS, headless, in-house — the REST + OpenAPI 3.0 spec gives your engineering team everything to build a pre-publish hook in under a day. Python and Node SDKs ship from the spec.
- What's the syndication-audit feature actually do?
- Add your distribution domains as an allowlist (medium.com/@you, your partners' domains). Weekly crawl reports back the canonical-tag health: ORIGIN (points to you), SYND OK (points to you from a partner), LEAK (canonical missing or pointing elsewhere — SEO ranking split), or SCRAPED (verbatim copy, no canonical, non-allowlisted domain). The scraped category comes with a timestamped snapshot for DMCA evidence.
- How does pricing work for a 30-reporter newsroom?
- Premium at $79/mo gets 5,000 words/check, 5,000 articles/month, full API + CMS integrations. For high-volume operations (10,000+ articles/mo), Enterprise contracts are volume-priced on the corpus + dedicated infrastructure dimension. Per-reporter seats not part of the model; the volume is what scales.
- What about op-eds, syndication, and book excerpts where some 'reuse' is expected?
- Per-section thresholds on Pro+ — set different sensitivity for news vs op-ed vs review vs syndicated excerpt. An op-ed citing a Senate floor speech verbatim should not trigger the same threshold as a news article reusing a competitor's phrasing. The citation-classification step separates legitimate quoted material from accidental reuse.
- Will it slow down our publish pipeline?
- Median 47 seconds per article on a 1,500–4,000-word check. The CMS-integration hook is async — fire on save_draft, the editor sees the verdict by the time they're done proofreading. Long-form features (8,000+ words) take ~90 seconds and can run as a pre-publish job rather than a synchronous hook.
- Can we add our back catalogue as a private corpus?
- Yes — Enterprise tier. Bulk-ingest your archived content into a tenant-private corpus, so the engine checks each new submission against your own back catalogue. Catches reporter self-recycling (sometimes legitimate, sometimes not), and intentional reuse of older content under a new byline.
Run a draft through the queue. See the per-article view.
Drop in a Word doc, a Google Doc, or a CMS-exported article. Get a per-paragraph report in 47 seconds. Free up to 2,500 words; Pro plan for volume; Enterprise for corpus + dedicated infrastructure.