Academic plagiarism checker.
Plagiarism detection built for scholarly writing — checks against 700M+ academic works across OpenAlex, CORE, arXiv, PubMed Central, and Crossref. Citation-aware classification. Apache 2.0 engine auditable by any institutional IT review.
Five sources. 700M+ works. All open-access or DOI-registered.
No paywalled subscription content — we work entirely off open-access indexes. Wider coverage for post-Plan-S research; narrower for older paywalled journal back-catalogues.
OpenAlex
250M+ worksOpen scholarly index — replaces the old Microsoft Academic Graph. Covers journal articles, conference papers, datasets, books across every discipline.
CORE
290M+ papersWorld's largest open-access research aggregator. Full-text indexed from 11,000+ repositories. Strong on European + UK research output.
arXiv
2.4M+ preprintsPhysics, mathematics, computer science, quantitative biology, statistics, economics. Cross-checks against the version your reviewer just saw on arXiv.
PubMed Central
10M+ papersOpen biomedical / life-sciences archive. Indexed full text + metadata. NIH-deposited research and journal mirror feeds.
Three workflows. One product.
Whether you're prepping a journal submission, defending a thesis chapter, or finalising a grant proposal — the checker fits the stage, not the other way around.
01 · Pre-submission journal check
Run the manuscript against the open-access universe before your journal does. iThenticate / Crossref Similarity Check will run it again at submission — you want to see what the editor sees, before the editor sees it. Flat-rate Pro means iterating across drafts without burning per-document credits.
02 · Dissertation / thesis chapter
Chunk by chapter or paste the whole thesis. Self-citation across chapters is flagged but classified separately (legitimate at thesis length). Bibliography auto-excluded from the score. Version history compares January's chapter against May's revision — see what you changed, what got cleaner, what's new.
03 · Grant proposal / lit review
Boilerplate from previous grants you co-wrote should not flag as plagiarism. Properly-cited literature review should not flag either. The classification model treats legitimate scholarly recycling (your own past work, with citation) differently from accidental copying. The score reflects what's actually problematic, not what's structurally unavoidable.
Six categories. Every matched interval lands in one.
A 28% similarity report on a heavily-cited literature review reads differently from a 28% on uncited paraphrasing. The classification is the difference.
01 · Direct quote (cited)
Quotation marks + parenthetical citation OR footnote. Counted as expected; excluded from similarity score. Multiple citation styles auto-detected (MLA / APA / Chicago / Harvard / Vancouver / IEEE).
02 · Block quote
Indented passage 40+ words with trailing citation. Counted as expected. The 0.5-inch left indent + citation pattern is recognised; missing-indent flags as direct-quote-formatting drift.
03 · Paraphrase (cited)
Sentence-shape match with attribution somewhere in the surrounding paragraph. Counted at reduced weight. The classification accepts “Smith (2023) argues...” as legitimate attribution preceding the paraphrase.
04 · Common knowledge
Phrases appearing in 2,000+ independent sources (“Water boils at 100°C”, standard technical definitions, well-known formulas). Excluded from similarity. Adjustable threshold for fields with denser terminology.
05 · Bibliography
References list auto-detected by heading + format. Excluded entirely from the similarity score. References themselves are NOT supposed to be original — flagging them inflates the score for no reason.
06 · Unattributed paraphrase
Sentence-shape match with no citation in the surrounding context, no quotation marks, no block-quote formatting. THE category that should drive your similarity score. Counted at full weight; the rest of the matched intervals don't.
Per-interval classification. Per-source URL. Auditable end-to-end.
Every matched interval ships with the source DOI / URL, the classification category, the corpus snapshot timestamp, and the engine commit. Defensible at a peer-review meeting, an integrity board, or a procurement IT audit.
Modernism is less a coherent movement than a constellation of rejections — of realism, of progress, of the unified self... (Eysteinsson, 1990, p.117).
doi.org/10.7551/mitpress/4596.001.0001 · MIT Press, 1990Eliot's emphasis on impersonality reframes the lyric as constructed artifact rather than confessional outpouring.
Eliot, T.S. (1919). Tradition and the Individual Talent · Egoist Press · doi.org/10.2307/24…The novel emerges from a fundamentally bourgeois worldview that took shape in the eighteenth century.
openalex.org/W12345678 · Watt (1957) — uncited matchJoyce's stream-of-consciousness owes much to William James's psychology of interior duration.
James, W. (1890). The Principles of Psychology, ch.IX (mentioned in citations list)The author's own previously published doctoral dissertation, Chapter 1.
ucl-discovery.com/dissertation/2024-Modernism-Chen.pdf · self-citation, no attribution this paragraphAll references — Eysteinsson, Eliot, Joyce, James, Watt, and 32 others.
Bibliography section, pp. 38–41- 700M+scholarly works indexed
- 6citation styles auto-detected
- 5languages calibrated at launch
- 3%ESL-bias residual (documented)
- Apache2.0 · IRB-reviewable
Three things a black-box checker can't survive in a faculty hearing.
Re-run the same paper against the same commit
Every report stamps X-Engine-Commit and a corpus snapshot timestamp. Months later, you can re-check the same manuscript against the engine version that produced the original score — same input, same algorithm, same result. Closed-source checkers can re-train silently and your January score won't reproduce in May.
Apache 2.0 source on GitHub — bring it to the meeting
When AI detection ships (coming soon, on the v1.2 roadmap) and a student disputes their AI verdict at an integrity board, you'll be able to cite the engine commit, point to the line of code that produced the threshold decision, and show the per-language ESL calibration documentation. Not a screenshot of a vendor support reply.
Liang et al. 2023 documented; mitigation on the roadmap
Stanford 2023 measured 61.3% FPR against TOEFL essays on naive detectors. When Noplag's AI detection rolls out (coming soon, v1.2 roadmap), the verdict will carry an ESL-adjusted band when the lexical signature suggests non-native English. Residual error will be documented, not claimed to be zero. Override path logged with engine commit + reason.
An Apache 2.0 engine your institution's IT can actually review before signing.
Most institutional procurement reviews of plagiarism software stall on the same questions: where does the document go after submission, what's the matching algorithm, who has access to the data, what happens at right-to-erasure, and can you replicate the vendor's accuracy claim. With a closed-source vendor, those questions get answered in a sales call. With Apache 2.0, IT can clone the repo, run the engine on a sandbox VM, audit the data flow, verify the right-to-erasure pathway, and reproduce the published F1 scores against their own test set. We're not arguing for cheaper review — we're arguing for a review that the IT team can actually conclude. That's why universities that previously refused to add new plagiarism vendors are signing this one.
Read the procurement guideThe questions academics actually ask.
Does this replace iThenticate for journal pre-submission?
For most cases, no — if your target journal mandates iThenticate (Elsevier, Springer, Wiley, IEEE workflows), they'll run it at submission regardless. What Noplag replaces is the iterative self-check across drafts before that submission. iThenticate at $125/document gets prohibitive after the 4th revision; flat-rate Pro doesn't. See the full vs-iThenticate comparison at noplag.com/vs-ithenticate.
What about ProQuest dissertations? You don't have those?
Correct — that gap is real. ProQuest's dissertation database is proprietary, paywalled, and not in OpenAlex/CORE. We cover the ~700M open-access scholarly works, which is broader than ProQuest for post-Plan-S research (most major journals 2018 onwards); narrower for older paywalled dissertations specifically. If cross-checking against US PhD dissertations is your dominant requirement, iThenticate's the right tool for that one job.
How does the engine handle citation styles other than the big six?
MLA, APA, Chicago / Turabian, Harvard, Vancouver, and IEEE auto-detect from the bibliography format. For less common styles (ACS, AMA, ASA, Bluebook, Oxford), the engine still classifies intervals based on quotation marks + nearby parenthetical patterns, but the style-specific in-text validator is on the v1.2 roadmap. Self-host users can patch in additional styles via the alignment/ module.
What's the false-positive rate on heavily cited literature reviews?
Lower than uncited-paraphrase reviewers expect, because the classification model excludes properly-cited intervals from the headline score. We measure ~3.4% false-positive rate on the PAN-PC-11 academic-prose subset (where most intervals are legitimate scholarly quoting). The closed-source competitors who don't classify cited intervals report 12–18% on the same corpus.
Can I exclude my own previous publications?
Pro tier — add a self-citation allowlist (your ORCID, your past DOIs, your institutional repository URL). Matches against allowlisted documents tag as SELF-CITE and don't count toward the score. Useful for academics whose grant proposals build on their previous published work; self-recycling with citation is expected, not plagiarism.
What's the multilingual coverage actually like?
Detection is calibrated for English, Spanish, Portuguese, Polish, and Ukrainian at launch, with more languages rolling out. Detection quality is benchmarked on PAN-PC-11; per-language evaluation is on the roadmap. Strongest on the European languages with the longest validation history; newer additions ship as their calibration matures.
What does “Apache 2.0” mean for institutional procurement?
Your IT / DPO / IRB can clone github.com/NoplagLabs/noplag-engine, run the engine on a sandbox, audit the data flow, verify right-to-erasure works end-to-end, and reproduce our published benchmark numbers against their own test corpus. Without signing anything, without an NDA, without trusting our marketing. For institutions whose procurement process requires this depth of review, that's the unlock.
Will the AI detector work on Claude-3 + Gemini text, or just GPT?
AI detection is coming soon — it's on the v1.2 roadmap, not live at launch. The approach is Binoculars (Hans et al. 2024), which is model-agnostic: it doesn't train per-LLM, it compares a small reference model's perplexity against the test text, which generalises. Accuracy figures publish when the feature clears evaluation, rather than being quoted ahead of it, on mixed-LLM content (vs ~70% on Claude-3 for GPT-trained detectors). The methodology paper documents the failure modes; we don't hide them.
What if my advisor disputes a flag?
Every flagged interval ships with the source DOI, the corpus snapshot timestamp, the engine commit, and the classification reason. Re-run the check against that engine commit and the result reproduces exactly. Bring the printed report to the meeting; the citation paths are click-through. We don't have a “trust the score” mode, because that's what black-box checkers have.
What about self-host for medical / clinical research with HIPAA?
Apache 2.0 + Docker — runs entirely inside your network. Pair with PostgreSQL + Redis on your own infra; no traffic leaves the VPC. Useful for clinical drafts, patient case studies, or any pre-publication research that can't go to a third-party cloud. Enterprise customers get the corpus shipped as an indexable bundle.
Run a chapter against 700M+ scholarly works.
Drop in a dissertation chapter, grant proposal, or manuscript draft. The citation-aware report shows what's expected, what needs a citation, and what's your own. Free up to 2,500 words. Apache 2.0 if your IRB needs the algorithm reviewed first.