REPORTERPre-file self-check · Source verification · Byline-theft tracking

Plagiarism checker for journalists.

Pre-file self-check before the editor sees it. Detects paraphrased lifts and press-release boilerplate, with AI-generated-submission detection coming soon. Plus: track who's republishing your byline. Free up to 2,500 words; REST API for newsroom integration if you're filing into a CMS.

Pre-file self-checkSource verificationByline-theft detection
WHATEVER YOU'RE FILING

Four reporter workflows. One self-check.

BEAT REPORTER

Daily filing on deadline

Your editor sees it next — make sure it's clean first.

  • 30-second self-check on a 900-word filing
  • Press-release boilerplate auto-flagged
  • Wire-service overlap surfaced before publish
FEATURE WRITER

Long-form with reused research

5,000-word features pull from interview transcripts + prior reporting.

  • Self-citation: your own past features tagged SELF-CITE
  • Quoted-material classification (attribution vs paraphrase)
  • Embargo-aware: check against unembargoed materials only
INVESTIGATIVE

Months of source documents in one piece

Detect accidental quote-without-attribution drift across the draft.

  • Per-paragraph source attribution check
  • Document-cache scan against your folder of sources
  • DMCA-grade snapshot when your scoop gets republished
FREELANCER

Multi-outlet filings, no in-house QA

You're the editor; the self-check is the only check before the editor sees it.

  • Mobile-ready check on a phone (you're in the field)
  • Cross-publication self-recycling flagged with allowlist
  • Free tier covers a typical 2,500-word filing
PRE-FILE WORKFLOW

Three steps before the file button.

Embedded in your filing pipeline via API, or one-off in the dashboard. Same engine. Same report. Same 47 seconds.

01

01 · Self-check the draft

Paste or upload the filed copy. Add your byline-history allowlist (your past pieces, your wire-service syndication list). The engine catches accidental paraphrase of your own prior work (legitimate) separately from accidental paraphrase of someone else's (not).

02

02 · Verify the sources

The report surfaces every matched interval with the source URL, canonical-tag status, and snapshot date. Press-release boilerplate is auto-flagged as REPRINT (some news outlets allow press-release quoting verbatim, with attribution; others don't). Per-paragraph AI-generated-text classification with an ESL-adjusted band is coming soon.

03

03 · Decide before filing

Look at the classification: properly-attributed quotes (excluded from headline score), paraphrase-with-citation (low weight), uncited-paraphrase (the one you fix). Once AI detection rolls out, AI-likely paragraphs will need a disclosure or a rewrite per your outlet's policy. Once it's clean, file. The report attaches to the filing as an audit-trail artifact.

SOURCE VERIFICATION

Six things every paragraph needs checked before filing.

Every category is shown next to its matched interval. Auto-flag where attribution is required by your outlet's policy. The bar is “would a standards desk accept this?”

01 · DIRECT QUOTE

Quotation marks + speaker attribution in the same sentence. Counted as quoted, excluded from your similarity score. Detected against your transcript folder if you've uploaded interview recordings as primary source material.

02 · PRESS RELEASE

Verbatim text matched against a public press release. Tagged REPRINT. Some outlets allow press-release quoting with attribution; others (most US dailies) require rewriting. Configure per-outlet policy in Settings.

03 · WIRE COPY

Match against AP, Reuters, AFP, Bloomberg, dpa, Kyodo, PA wire feeds (Premium tier with wire-service corpus subscription). Wire copy is typically rewriteable under your outlet's license; the engine confirms the wire source so you know what's been licensed vs lifted.

04 · EMBARGO HOLDS

Add embargoed press materials to a temporary scope. Embargo-aware: matches against embargoed content tag as EMBARGO-OK until the embargo date, then revert to standard source classification. Useful for science journalism where embargoes are routine.

05 · AI-GENERATED (COMING SOON)

Per-paragraph AI-likelihood with ESL-adjusted verdict, on the roadmap. Built for vetting freelancer submissions — AI-generated pieces are common enough now that pre-publish checks will catch them before reviewer 2 (your standards desk) does. Open methodology (Binoculars, Hans et al. 2024).

06 · YOUR OWN PRIOR

Byline-history allowlist: your past published pieces, your portfolio site URL, your Substack handle, your Medium archive. Self-recycling tagged SELF-CITE — common in feature writing when you build on a prior story. Doesn't inflate your similarity score; you'll still see it flagged for review.

THE STORY-CHECK VIEW

What you see after 47 seconds on a 1,500-word filing.

Per-paragraph classification with source URLs, attribution status, and click-through to the matched intervals. Defensible at a standards-desk meeting; reproducible against the engine commit stamped on the report.

PRE-FILE CHECK · 1,632 words · 18 intervals classifiedState Senate budget showdown — what's next
Headline · 11.4%AI · 0.07
#MATCHED INTERVAL · SOURCECLASSIFICATION
04
“The current budget framework simply doesn't reflect the realities of our communities,” Senator Hadley told reporters Monday.Hadley press office release · 2026-05-19 · attribution present
QUOTED · CITED
07
The proposed cuts would reduce K-12 funding by approximately $340 million across the biennium.ap.org · State Senate Budget Roundup · 2026-05-20
WIRE · AP
11
Critics have called the cuts “a generational disinvestment in public education.”Marisol Chen (your own) · last week's piece on Senate Ed Committee
SELF-CITE
13
Hadley argued the framework prioritizes essential services while “living within our means as a state.”Hadley press release · attribution + quote marks
QUOTED · CITED
14
Education advocates have rallied at the Capitol every Tuesday for the past five weeks.your draft · original reporting · no source match
ORIGINAL
16
Hadley described the budget as “fiscally prudent, even if politically difficult.”TIMES wire · Hadley interview transcript · 2026-05-19 · same quote, different outlet
WIRE · TIMES
Headline 11.4% = wire overlap + self-cite weight; quoted intervals + originals excluded.Engine v0.4.2 · ready to file · attach report to filing
47smedian pre-file check
6wire feeds (AP/Reuters/AFP/Bloomberg/dpa/Kyodo)
0filings used for training
5languages at launch, expanding
Apache2.0 · auditable methodology
BYLINE THEFT

Three things to do when your scoop shows up under someone else's byline.

01 · MONITOR

Set your byline as a watch query

Paste your published piece (or list of pieces) into the watch list. Daily web-scan surfaces any non-allowlisted domain publishing matched intervals of your reporting — verbatim copies, light rewrites, AI-spun versions. Notification includes the offending URL, the matched paragraphs, and the timestamp of the scrape.

Catches the spammy aggregator within 24h; catches the rival outlet's “inspired by” rewrite within a week.
02 · EVIDENCE PACK

DMCA-grade snapshot, dated and hashed

Every flagged scrape ships with a HTML snapshot of the scraper's page, dated to the day it was crawled, hashed for tamper-evidence. Pair with the timestamp on your original byline (your CMS publish date, your tweet, the Wayback Machine capture) and you have the priority chain.

DMCA filings are a legal call. We give your lawyer the paperwork; we don't file for you.
03 · CANONICAL HEALTH

Your byline on legitimate syndication

If you syndicate to Medium, LinkedIn, Substack, or partner outlets, the engine tracks whether their canonical tag points back to your original. ORIGIN means SEO credit flows to you; LEAK means the syndication is splitting your ranking. Common LinkedIn issue: they strip canonical tags by default.

Add an in-body “This piece first appeared at…” link when canonical-tag drift is your only option.
WIRE & SYNDICATION

Two things journalists need a checker to handle that academic tools never bothered with.

Most plagiarism tools weren't built for newsrooms. They flag verbatim wire copy as plagiarism even though every paper running the same AP feed is using it under license — a false positive that's exhausting on volume. They miss the byline-theft case entirely because they only check what you submit, not the live web that's republishing you. Noplag's wire-service corpus subscription (AP, Reuters, AFP, Bloomberg, dpa, Kyodo, PA) classifies wire copy as WIRE rather than PLAGIARISM. The byline-monitoring service watches the live web for republished bylines. Both features are off by default — they're the journalism-specific layer you opt into when the standard plagiarism check isn't enough.

Read the wire-classification guide
JOURNALISM WORKFLOW NOTES
WIREAP / Reuters / AFP / Bloomberg / dpa / Kyodo / PA feed classified as WIRE — not PLAGIARISM. Configure per-outlet license.
PRESSVerbatim press release matched against the public release; tagged REPRINT. Outlet-policy-configurable handling.
EMBARGOEmbargoed-materials scope: matches tag as EMBARGO-OK until the date, revert to standard afterwards.
AIPer-paragraph AI-likelihood with ESL-adjusted band, coming soon. Built to catch AI-generated freelancer submissions before the editor sees them.
BYLINEByline-history allowlist: your past pieces tag as SELF-CITE. Cross-publication recycling expected at feature-length.
WATCHSet published pieces as a watch query; daily web-scan surfaces scrapers and rewriters. DMCA-grade evidence pack on demand.
FAQ

The questions reporters actually ask.

Will my draft be stored or used to train your models?
No. Drafts are fingerprinted (winnowing hashes, non-reversible) and stored in your tenant-isolated database. The text itself is purged after 30 days unless you opt into retention. Detection models (Binoculars-based) are pre-trained on public corpora; we don't add reporter drafts to any training set. Explicit in the DPA.
How does the wire-classification work?
Premium tier with wire-service corpus subscription (AP, Reuters, AFP, Bloomberg, dpa, Kyodo, PA). Verbatim matches against wire feeds are tagged WIRE — not PLAGIARISM — because every paper running the feed is using it under license. Configure per-outlet policy: some outlets allow unlimited wire-quoting, others require explicit attribution. The classification is shown next to every interval.
What happens with press releases?
Matched against the public press-release index (PR Newswire, Business Wire, EurekAlert, PRWeb). Tagged REPRINT. Some outlets allow press-release quoting verbatim with attribution; others (most US dailies) require rewriting. Configure your outlet's policy in Settings; the report tells you what action your policy requires.
How will the AI-detection handle freelancer submissions?
AI detection is coming soon (on the v1.2 roadmap). When it ships, it will give a per-paragraph AI-likelihood with ESL-adjusted verdict — built for newsrooms vetting freelance pitches, since AI-generated submissions are increasingly common. The ESL-adjusted band matters because non-native English writing patterns overlap with AI signatures (Stanford 2023 found 61.3% FPR on naive detectors); our calibration is designed to bring that down substantially, with the residual rate published when the feature clears evaluation.
Can I monitor my published bylines for scrapers?
Premium tier with the byline-watch feature. Add your published pieces to the watch list; daily web-scan surfaces non-allowlisted domains publishing matched intervals. Each flag ships with a hashed HTML snapshot dated to the day of the scrape — DMCA-grade evidence. We don't file DMCAs (that's your legal counsel's call); we give them the paperwork.
What about embargoes?
Add embargoed materials to a temporary scope with the embargo date. Matches against embargoed content tag as EMBARGO-OK until the date, then revert to standard handling. Useful for science journalism (NIH / EurekAlert embargoes are routine) and policy journalism (Congressional hearing transcripts with embargo).
Does it work for non-English reporting?
Detection is calibrated for English, Spanish, Portuguese, Polish, and Ukrainian at launch, with more languages rolling out. Detection quality is benchmarked on PAN-PC-11; per-language evaluation is on the roadmap. Useful for international correspondents filing in multiple languages.
What's the API like for newsroom CMS integration?
REST + OpenAPI 3.0 spec. Python and Node SDKs ship from the spec. Webhooks for async checks on long-form features. POST /v1/checks from your CMS hook on save_draft; the report attaches to the article as audit-trail metadata. Sample WordPress + Ghost + Drupal integrations at github.com/NoplagLabs/noplag-engine.
What if I'm a freelancer without a newsroom backing me up?
Free tier covers 2,500 words per check, unlimited checks, no signup. Mobile-friendly (you can run a check on a phone before filing from the field). Premium tier at $79/mo for the wire-classification + byline-watch + folder-level features when you need them. Self-host is also Apache 2.0 free if you want to run everything locally.
What happens if a source's quote turns out to be AI-generated?
Once AI detection ships (coming soon), the per-paragraph AI verdict will surface the suspicious paragraph with an ESL-adjusted band. If the quote attributed to a source comes back AI-likely, that's an editorial flag — could be the source AI-generated their statement, could be the source paraphrasing AI output, could be a false positive. The verdict is a signal to verify the quote directly with the source, not an automatic disqualification.

Check it before the editor does.

Drop in the filed copy. Get a per-paragraph report in 47 seconds: quoted material, wire overlap, self-citation, with AI-likelihood coming soon. Free up to 2,500 words. Premium tier for the wire-classification corpus and byline-watch.

Plagiarism checker for journalists — pre-file + byline theft