RESEARCHPre-submission · Self-citation aware · EU residency · IRB-reviewable

Plagiarism checker for researchers.

Pre-submission manuscript checks against 700M+ scholarly works. My Folders tracks your own prior publications so legitimate self-citation doesn't inflate the score. EU-residency endpoint available for sensitive pre-publication research. Apache 2.0 engine — IRB-reviewable.

700M+ scholarly worksSelf-citation allowlistEU residency · IRB-reviewable
WHEREVER YOU ARE IN THE PIPELINE

Four research roles. One self-check workflow.

PHD / DOCTORAL

Dissertation chapter pre-submission

Iterate across chapters. Track self-citation across drafts.

  • Per-chapter check + version history
  • Self-citation allowlist (your ORCID + DOI list)
  • Bibliography auto-excluded
POST-DOC

Manuscript before journal submission

iThenticate hits at submission; we run before.

  • 700M+ open-access scholarly index
  • Citation classification (cited vs uncited paraphrase)
  • Flat-rate Pro · iterate without per-credit cost
GRANT WRITER

Lit review + proposal vetting

Recycling your past grant boilerplate is expected — flag it.

  • Recognise your own prior publications + grants
  • Co-author allowlist (collaborator's prior work)
  • Export classification report to PDF for collaborators
LAB PI

Lab-wide pre-publication QA

Catch issues across the lab's submissions in one queue.

  • Folder per project + per-postdoc
  • EU residency for clinical / sensitive primary data
  • Apache 2.0 engine — defensible at IRB
PRE-SUBMISSION WORKFLOW

Three steps before the manuscript hits the journal portal.

Run iteratively. Catch self-cited intervals so they don't inflate your iThenticate / Crossref score at submission. AI-detection verdicts are coming soon, so you can preview them before reviewer 2 raises it.

01

01 · Draft + allowlist

Upload the chapter / manuscript draft. Add your ORCID + a list of your DOIs (the allowlist) — your own prior publications get tagged SELF-CITE, not flagged as plagiarism. Co-author DOIs included if you collaborated; the engine recognises legitimate scholarly recycling.

02

02 · Run the cascade

L1 winnowing → live-web today, with L2 MinHash (v1.1) and L3 vector embeddings (v1.2) to follow. Catches verbatim and near-verbatim copying; paraphrase coverage arrives with L2 and L3. Citation classification separates quoted material from uncited paraphrase. Multilingual, with detection calibrated for English, Spanish, Portuguese, Polish, and Ukrainian at launch and more languages rolling out. Median 47 seconds per manuscript.

03

03 · Review + iterate

The report flags intervals by category: SELF-CITE (your own, expected), CITED (quoted with attribution, expected), CITED-PARA (paraphrased with citation, low weight), UNCITED-PARA (the one that needs attention), and AI-LIKELY (with an ESL-adjusted band) coming soon. Reproducible — engine commit + corpus snapshot stamped on every report.

SELF-CITATION HANDLING

Six categories of legitimate scholarly recycling — none of which should inflate your score.

Black-box checkers flag any text-match as plagiarism. Researchers recycle text legitimately all the time. Here's how the classifier handles it.

01 · YOUR OWN PRIOR WORK

ORCID + DOI allowlist. Your previously-published papers tag as SELF-CITE when matched — not flagged as plagiarism. Common in lit reviews (“as I argued in Chen 2023…”) and grant proposals where you build on your published work.

02 · CO-AUTHORED WORK

Add collaborator ORCIDs to the allowlist. Co-authored papers tag as CO-CITE. Useful when your manuscript builds on a paper you co-wrote with a collaborator — the boilerplate methods section is legitimately shared.

03 · LAB HERITAGE TEXT

PI-folder shared corpus: standard methods sections, lab protocols, ethics statements that propagate across the lab's publications. Tag as LAB-CITE and exclude from the score. The IRB-approved boilerplate doesn't need to be re-written for every paper.

04 · EMBARGOED PRE-PRINTS

Your own arXiv / bioRxiv pre-print of the manuscript you're now submitting. The classifier recognises pre-print → journal-version evolution and tags as PREPRINT-SELF. Doesn't count toward the score. Journals that explicitly disallow pre-prints can configure stricter handling.

05 · RETRACTED CITATIONS

Crossref Retraction Watch integration flags any matched intervals citing retracted papers. Doesn't change your similarity score — but you'll want to know before submission. Catches the “I cited a paper that got retracted last month” problem before reviewer 2 does.

06 · IRB / SENSITIVE

EU-residency endpoint (api.eu.noplag.com) keeps the manuscript inside EU infrastructure. Self-host the engine entirely inside your network for HIPAA / GDPR / patient-data manuscripts. Apache 2.0 + Docker; no traffic leaves your VPC.

MY FOLDERS · LAB VIEW

Organized by project. Audit-trail by date.

Group your manuscripts by project, lab year, or grant cycle. Re-run drafts against the same engine commit later — reports are reproducible months out.

MY FOLDERS · Dr. Hirano lab · CompBio24 manuscripts · 6 projects
EU RESIDENCYNew folder
FOLDERDOCSLATEST ACTIVITYREGION
Project: PD-L1 expression cohort8 manuscriptsdraft v3 checked · 2h agoEU
Project: synthetic-bio safety review5 manuscriptssubmission-ready · yesterdayEU
Grant: NIH R01 renewal Y33 manuscriptslit-review v2 checked · 2d agoEU
Lab heritage: methods + ethics text4 docs · sharedlab-allowlist updated · 1w agoEU
Pre-prints: arXiv pending2 manuscriptsPREPRINT-SELF tagged · 1w agoEU
Archived (2022–2024)12 papersself-citation reference setEU
All folders EU-residency · encrypted at rest · IRB-reviewable engine commit on every reportreproducibility: engine v0.4.2 active
700M+scholarly works indexed
5detection languages at launch, expanding
EUresidency endpoint available
Apache2.0 · IRB-reviewable
Crossref+ Retraction Watch integration
DATA PROTECTION

Three things sensitive pre-publication research needs from a checker.

01 · EU RESIDENCY

Manuscripts never leave EU infrastructure

Pro+ provisioned to api.eu.noplag.com. Hosted on EU-based Hetzner data centres (Falkenstein + Helsinki). The submission, fingerprint store, and report stay in the EU. GDPR-compliant by default; DPA available on request.

Useful for collaborations with EU partners under Horizon Europe data-protection requirements.
02 · SELF-HOST

Run the engine inside your own institution

Apache 2.0 + Docker compose. The full detection pipeline runs entirely on your infrastructure. Bring your own Postgres + Redis; we ship the corpus as an indexable bundle on Enterprise. No traffic leaves your VPC. The right answer for clinical-trial data, patient case studies, or commercially-sensitive primary research.

Pairs cleanly with Slurm clusters and existing institutional auth (SAML, LDAP, OIDC).
03 · RIGHT-TO-ERASURE

Delete every fingerprint we hold for you

One click in Settings. Within 24h, every fingerprint, report, allowlist entry, and metadata record we hold for your account is deleted. GDPR Article 17 compliant. Self-host users have direct database access; cloud users have the deletion API endpoint documented.

Auditable deletion log retained for compliance reviews.
IRB-REVIEWABLE

An Apache 2.0 engine your IRB can actually review before approving.

Most institutional research-software reviews of plagiarism tools stall on the same questions: where does the manuscript go after submission, what's the matching algorithm doing, is patient/sensitive data ever transmitted, what's the right-to-erasure pathway, and can you reproduce the published accuracy claim. With closed-source vendors, those questions get answered in a sales call. With Apache 2.0, your IRB can clone the repo, run the engine on a sandbox, audit the data flow, verify the right-to-erasure works end-to-end, and reproduce the published per-language F1 scores against their own test set. We're not arguing for cheaper review — we're arguing for a review your IRB can actually conclude. That's why research-data offices that previously refused new plagiarism vendors are signing this one.

Read the IRB review guide
IRB QUESTIONS, ANSWERED
Q1Where does the manuscript go? — Tenant-isolated EU Postgres if you're on the EU-residency endpoint. Source text purged after 30 days unless retention is enabled.
Q2Trained on submissions? — No. Submitted manuscripts are never used to train detection models. Explicit in the DPA.
Q3Patient / sensitive data handling? — Self-host the engine entirely inside your network. No traffic leaves your VPC. HIPAA-compatible deployment.
Q4Algorithm review? — Apache 2.0 at github.com/NoplagLabs/noplag-engine. Readable, auditable, reproducible.
Q5Multi-institution access? — Per-tenant role-based access. Audit logs for every read. Or self-host so the data never leaves your institution.
Q6Right-to-erasure? — One-click in Settings. 24-hour completion. GDPR Art. 17 compliant. Audit log retained.
FAQ

The questions a research office actually asks.

Why does my similarity score change when I add my ORCID + DOI list?
Because the engine retroactively classifies every interval that matched a paper on your allowlist as SELF-CITE rather than UNCITED-PARA. The matches don't disappear — they're tagged differently. SELF-CITE intervals are weighted at near-zero in the headline score. The math: if your dissertation chapter quotes your own pre-print twice and a co-author's prior paper once, that's three intervals that go from “counted” to “expected scholarly recycling”.
Does the engine train on submitted manuscripts?
No. Submitted text is fingerprinted (winnowing hashes only — a non-reversible signature) and stored in your tenant-isolated database. The text itself is purged after 30 days unless retention is enabled. The Binoculars-based AI detector — coming soon — will be pre-trained on public corpora; we don't add customer manuscripts to the training set. Explicit in the DPA.
What about my pre-print on arXiv / bioRxiv?
If you list the pre-print DOI in your allowlist, matches to it tag as PREPRINT-SELF and don't count toward your similarity score. The engine recognises that journal-version evolution from pre-print is expected. Journals that explicitly disallow pre-prints (some still do) can have the allowlist configured stricter — but most major journals (Nature, Cell, NEJM, PLOS) accept pre-prints, and the SELF tagging matches that policy.
How will the AI detector handle ESL co-authors?
AI detection is coming soon — on the v1.2 roadmap. When it ships, it uses per-language calibration with an ESL-adjusted verdict. When the lexical signature suggests non-native English (common in multinational collaborations), the verdict ships as ESL ADJ with a documented residual error rather than a claim of zero. The Liang et al. 2023 paper measured 61.3% FPR on TOEFL essays for naive detectors; we publish our mitigation methodology at /docs/developers/benchmarks. ESL co-authors don't get unfairly flagged.
Can I run the engine on a Slurm cluster for batch checks?
Yes — Apache 2.0 + Docker. Run the engine container on your cluster's compute nodes; ingest from your shared filesystem; emit reports as JSON to wherever your pipeline reads. Useful for batch-checking a year's worth of lab output against the latest corpus snapshot. Sample Slurm sbatch script at github.com/NoplagLabs/noplag-engine.
What about clinical / patient data manuscripts?
Self-host inside your VPC. No traffic leaves your network. The engine doesn't need internet access for indexed-corpus checks (you'd lose the live-web verification layer, but that's optional). HIPAA-compatible deployment with proper infrastructure setup; your IT runs it like any other internal service.
How is this different from iThenticate at submission?
iThenticate hits at submission time and your editor sees it — you don't get to iterate. Noplag is your self-check before that submission. The cascade is comparable on open-access content (broader, actually, post-Plan-S); narrower on subscription/paywalled academic content for which iThenticate has licensing. Self-citation handling is what most researchers actually need that iThenticate doesn't do well.
Does it integrate with my reference manager (Zotero / Mendeley / EndNote)?
Allowlist ingestion supports BibTeX, RIS, and direct ORCID API. Drop your Zotero library export or paste your ORCID; the allowlist populates with all your DOIs in seconds. Real-time sync with Mendeley + EndNote is on the v1.2 roadmap; for now it's the export pathway.
What's the multilingual coverage like for research published in non-English journals?
Detection is calibrated for English, Spanish, Portuguese, Polish, and Ukrainian at launch, with more languages rolling out. Strongest on European languages (long calibration history). Detection quality is benchmarked on PAN-PC-11; per-language evaluation is on the roadmap. Useful for European and South American academic contexts in particular.
What happens if a cited paper is retracted after I submit?
The Crossref Retraction Watch integration flags it on every re-check of your manuscript. If you've already submitted, you'll want to notify the editor (some journals require explicit handling of retraction-after-submission). Doesn't change the similarity score, but you'll know before reviewer 2 does.

Run a chapter through the cascade. See the self-citation breakdown.

Drop in a manuscript draft, a dissertation chapter, or a grant lit-review. Add your ORCID. See which intervals are SELF-CITE, CITED, and UNCITED-PARA, with AI-LIKELY coming soon. Free up to 2,500 words. EU residency on Pro+ for sensitive pre-publication material.

Plagiarism checker for researchers — pre-submission — Noplag