Plagiarism checker for journalists.
Pre-file self-check before the editor sees it. Detects paraphrased lifts and press-release boilerplate, with AI-generated-submission detection coming soon. Plus: track who's republishing your byline. Free up to 2,500 words; REST API for newsroom integration if you're filing into a CMS.
Four reporter workflows. One self-check.
Daily filing on deadline
Your editor sees it next — make sure it's clean first.
- 30-second self-check on a 900-word filing
- Press-release boilerplate auto-flagged
- Wire-service overlap surfaced before publish
Long-form with reused research
5,000-word features pull from interview transcripts + prior reporting.
- Self-citation: your own past features tagged SELF-CITE
- Quoted-material classification (attribution vs paraphrase)
- Embargo-aware: check against unembargoed materials only
Months of source documents in one piece
Detect accidental quote-without-attribution drift across the draft.
- Per-paragraph source attribution check
- Document-cache scan against your folder of sources
- DMCA-grade snapshot when your scoop gets republished
Multi-outlet filings, no in-house QA
You're the editor; the self-check is the only check before the editor sees it.
- Mobile-ready check on a phone (you're in the field)
- Cross-publication self-recycling flagged with allowlist
- Free tier covers a typical 2,500-word filing
Three steps before the file button.
Embedded in your filing pipeline via API, or one-off in the dashboard. Same engine. Same report. Same 47 seconds.
01 · Self-check the draft
Paste or upload the filed copy. Add your byline-history allowlist (your past pieces, your wire-service syndication list). The engine catches accidental paraphrase of your own prior work (legitimate) separately from accidental paraphrase of someone else's (not).
02 · Verify the sources
The report surfaces every matched interval with the source URL, canonical-tag status, and snapshot date. Press-release boilerplate is auto-flagged as REPRINT (some news outlets allow press-release quoting verbatim, with attribution; others don't). Per-paragraph AI-generated-text classification with an ESL-adjusted band is coming soon.
03 · Decide before filing
Look at the classification: properly-attributed quotes (excluded from headline score), paraphrase-with-citation (low weight), uncited-paraphrase (the one you fix). Once AI detection rolls out, AI-likely paragraphs will need a disclosure or a rewrite per your outlet's policy. Once it's clean, file. The report attaches to the filing as an audit-trail artifact.
Six things every paragraph needs checked before filing.
Every category is shown next to its matched interval. Auto-flag where attribution is required by your outlet's policy. The bar is “would a standards desk accept this?”
01 · DIRECT QUOTE
Quotation marks + speaker attribution in the same sentence. Counted as quoted, excluded from your similarity score. Detected against your transcript folder if you've uploaded interview recordings as primary source material.
02 · PRESS RELEASE
Verbatim text matched against a public press release. Tagged REPRINT. Some outlets allow press-release quoting with attribution; others (most US dailies) require rewriting. Configure per-outlet policy in Settings.
03 · WIRE COPY
Match against AP, Reuters, AFP, Bloomberg, dpa, Kyodo, PA wire feeds (Premium tier with wire-service corpus subscription). Wire copy is typically rewriteable under your outlet's license; the engine confirms the wire source so you know what's been licensed vs lifted.
04 · EMBARGO HOLDS
Add embargoed press materials to a temporary scope. Embargo-aware: matches against embargoed content tag as EMBARGO-OK until the embargo date, then revert to standard source classification. Useful for science journalism where embargoes are routine.
05 · AI-GENERATED (COMING SOON)
Per-paragraph AI-likelihood with ESL-adjusted verdict, on the roadmap. Built for vetting freelancer submissions — AI-generated pieces are common enough now that pre-publish checks will catch them before reviewer 2 (your standards desk) does. Open methodology (Binoculars, Hans et al. 2024).
06 · YOUR OWN PRIOR
Byline-history allowlist: your past published pieces, your portfolio site URL, your Substack handle, your Medium archive. Self-recycling tagged SELF-CITE — common in feature writing when you build on a prior story. Doesn't inflate your similarity score; you'll still see it flagged for review.
What you see after 47 seconds on a 1,500-word filing.
Per-paragraph classification with source URLs, attribution status, and click-through to the matched intervals. Defensible at a standards-desk meeting; reproducible against the engine commit stamped on the report.
Three things to do when your scoop shows up under someone else's byline.
Set your byline as a watch query
Paste your published piece (or list of pieces) into the watch list. Daily web-scan surfaces any non-allowlisted domain publishing matched intervals of your reporting — verbatim copies, light rewrites, AI-spun versions. Notification includes the offending URL, the matched paragraphs, and the timestamp of the scrape.
DMCA-grade snapshot, dated and hashed
Every flagged scrape ships with a HTML snapshot of the scraper's page, dated to the day it was crawled, hashed for tamper-evidence. Pair with the timestamp on your original byline (your CMS publish date, your tweet, the Wayback Machine capture) and you have the priority chain.
Your byline on legitimate syndication
If you syndicate to Medium, LinkedIn, Substack, or partner outlets, the engine tracks whether their canonical tag points back to your original. ORIGIN means SEO credit flows to you; LEAK means the syndication is splitting your ranking. Common LinkedIn issue: they strip canonical tags by default.
Two things journalists need a checker to handle that academic tools never bothered with.
Most plagiarism tools weren't built for newsrooms. They flag verbatim wire copy as plagiarism even though every paper running the same AP feed is using it under license — a false positive that's exhausting on volume. They miss the byline-theft case entirely because they only check what you submit, not the live web that's republishing you. Noplag's wire-service corpus subscription (AP, Reuters, AFP, Bloomberg, dpa, Kyodo, PA) classifies wire copy as WIRE rather than PLAGIARISM. The byline-monitoring service watches the live web for republished bylines. Both features are off by default — they're the journalism-specific layer you opt into when the standard plagiarism check isn't enough.
Read the wire-classification guideThe questions reporters actually ask.
- Will my draft be stored or used to train your models?
- No. Drafts are fingerprinted (winnowing hashes, non-reversible) and stored in your tenant-isolated database. The text itself is purged after 30 days unless you opt into retention. Detection models (Binoculars-based) are pre-trained on public corpora; we don't add reporter drafts to any training set. Explicit in the DPA.
- How does the wire-classification work?
- Premium tier with wire-service corpus subscription (AP, Reuters, AFP, Bloomberg, dpa, Kyodo, PA). Verbatim matches against wire feeds are tagged WIRE — not PLAGIARISM — because every paper running the feed is using it under license. Configure per-outlet policy: some outlets allow unlimited wire-quoting, others require explicit attribution. The classification is shown next to every interval.
- What happens with press releases?
- Matched against the public press-release index (PR Newswire, Business Wire, EurekAlert, PRWeb). Tagged REPRINT. Some outlets allow press-release quoting verbatim with attribution; others (most US dailies) require rewriting. Configure your outlet's policy in Settings; the report tells you what action your policy requires.
- How will the AI-detection handle freelancer submissions?
- AI detection is coming soon (on the v1.2 roadmap). When it ships, it will give a per-paragraph AI-likelihood with ESL-adjusted verdict — built for newsrooms vetting freelance pitches, since AI-generated submissions are increasingly common. The ESL-adjusted band matters because non-native English writing patterns overlap with AI signatures (Stanford 2023 found 61.3% FPR on naive detectors); our calibration is designed to bring that down substantially, with the residual rate published when the feature clears evaluation.
- Can I monitor my published bylines for scrapers?
- Premium tier with the byline-watch feature. Add your published pieces to the watch list; daily web-scan surfaces non-allowlisted domains publishing matched intervals. Each flag ships with a hashed HTML snapshot dated to the day of the scrape — DMCA-grade evidence. We don't file DMCAs (that's your legal counsel's call); we give them the paperwork.
- What about embargoes?
- Add embargoed materials to a temporary scope with the embargo date. Matches against embargoed content tag as EMBARGO-OK until the date, then revert to standard handling. Useful for science journalism (NIH / EurekAlert embargoes are routine) and policy journalism (Congressional hearing transcripts with embargo).
- Does it work for non-English reporting?
- Detection is calibrated for English, Spanish, Portuguese, Polish, and Ukrainian at launch, with more languages rolling out. Detection quality is benchmarked on PAN-PC-11; per-language evaluation is on the roadmap. Useful for international correspondents filing in multiple languages.
- What's the API like for newsroom CMS integration?
- REST + OpenAPI 3.0 spec. Python and Node SDKs ship from the spec. Webhooks for async checks on long-form features. POST /v1/checks from your CMS hook on save_draft; the report attaches to the article as audit-trail metadata. Sample WordPress + Ghost + Drupal integrations at github.com/NoplagLabs/noplag-engine.
- What if I'm a freelancer without a newsroom backing me up?
- Free tier covers 2,500 words per check, unlimited checks, no signup. Mobile-friendly (you can run a check on a phone before filing from the field). Premium tier at $79/mo for the wire-classification + byline-watch + folder-level features when you need them. Self-host is also Apache 2.0 free if you want to run everything locally.
- What happens if a source's quote turns out to be AI-generated?
- Once AI detection ships (coming soon), the per-paragraph AI verdict will surface the suspicious paragraph with an ESL-adjusted band. If the quote attributed to a source comes back AI-likely, that's an editorial flag — could be the source AI-generated their statement, could be the source paraphrasing AI output, could be a false positive. The verdict is a signal to verify the quote directly with the source, not an automatic disqualification.
Check it before the editor does.
Drop in the filed copy. Get a per-paragraph report in 47 seconds: quoted material, wire overlap, self-citation, with AI-likelihood coming soon. Free up to 2,500 words. Premium tier for the wire-classification corpus and byline-watch.