ACADEMIC700M+ scholarly works · Apache 2.0

Academic plagiarism checker.

Plagiarism detection built for scholarly writing — checks against 700M+ academic works across OpenAlex, CORE, arXiv, PubMed Central, and Crossref. Citation-aware classification. Apache 2.0 engine auditable by any institutional IT review.

0 / 1,500 words
700M+ scholarly worksCitation-aware classificationApache 2.0 · institution-auditable
THE SCHOLARLY CORPUS

Five sources. 700M+ works. All open-access or DOI-registered.

No paywalled subscription content — we work entirely off open-access indexes. Wider coverage for post-Plan-S research; narrower for older paywalled journal back-catalogues.

01 · INDEXmonthly

OpenAlex

250M+ works

Open scholarly index — replaces the old Microsoft Academic Graph. Covers journal articles, conference papers, datasets, books across every discipline.

02 · OPEN ACCESSweekly

CORE

290M+ papers

World's largest open-access research aggregator. Full-text indexed from 11,000+ repositories. Strong on European + UK research output.

03 · PREPRINTdaily

arXiv

2.4M+ preprints

Physics, mathematics, computer science, quantitative biology, statistics, economics. Cross-checks against the version your reviewer just saw on arXiv.

04 · BIOMEDweekly

PubMed Central

10M+ papers

Open biomedical / life-sciences archive. Indexed full text + metadata. NIH-deposited research and journal mirror feeds.

Crossref · 150M+ DOI-registered works · used for canonical resolution + citation graph traversal
ACADEMIC WORKFLOWS

Three workflows. One product.

Whether you're prepping a journal submission, defending a thesis chapter, or finalising a grant proposal — the checker fits the stage, not the other way around.

01

01 · Pre-submission journal check

Run the manuscript against the open-access universe before your journal does. iThenticate / Crossref Similarity Check will run it again at submission — you want to see what the editor sees, before the editor sees it. Flat-rate Pro means iterating across drafts without burning per-document credits.

02

02 · Dissertation / thesis chapter

Chunk by chapter or paste the whole thesis. Self-citation across chapters is flagged but classified separately (legitimate at thesis length). Bibliography auto-excluded from the score. Version history compares January's chapter against May's revision — see what you changed, what got cleaner, what's new.

03

03 · Grant proposal / lit review

Boilerplate from previous grants you co-wrote should not flag as plagiarism. Properly-cited literature review should not flag either. The classification model treats legitimate scholarly recycling (your own past work, with citation) differently from accidental copying. The score reflects what's actually problematic, not what's structurally unavoidable.

CITATION-AWARE CLASSIFICATION

Six categories. Every matched interval lands in one.

A 28% similarity report on a heavily-cited literature review reads differently from a 28% on uncited paraphrasing. The classification is the difference.

01 · Direct quote (cited)

Quotation marks + parenthetical citation OR footnote. Counted as expected; excluded from similarity score. Multiple citation styles auto-detected (MLA / APA / Chicago / Harvard / Vancouver / IEEE).

02 · Block quote

Indented passage 40+ words with trailing citation. Counted as expected. The 0.5-inch left indent + citation pattern is recognised; missing-indent flags as direct-quote-formatting drift.

03 · Paraphrase (cited)

Sentence-shape match with attribution somewhere in the surrounding paragraph. Counted at reduced weight. The classification accepts “Smith (2023) argues...” as legitimate attribution preceding the paraphrase.

04 · Common knowledge

Phrases appearing in 2,000+ independent sources (“Water boils at 100°C”, standard technical definitions, well-known formulas). Excluded from similarity. Adjustable threshold for fields with denser terminology.

05 · Bibliography

References list auto-detected by heading + format. Excluded entirely from the similarity score. References themselves are NOT supposed to be original — flagging them inflates the score for no reason.

06 · Unattributed paraphrase

Sentence-shape match with no citation in the surrounding context, no quotation marks, no block-quote formatting. THE category that should drive your similarity score. Counted at full weight; the rest of the matched intervals don't.

THE SCHOLARLY REPORT

Per-interval classification. Per-source URL. Auditable end-to-end.

Every matched interval ships with the source DOI / URL, the classification category, the corpus snapshot timestamp, and the engine commit. Defensible at a peer-review meeting, an integrity board, or a procurement IT audit.

REPORT · 4,217 words · 47 intervals classifiedThe Modernist Novel: A Reassessment · Chapter 3
Headline · 8.2%Engine · v0.4.2
#MATCHED INTERVAL · SOURCECLASSIFICATION
12

Modernism is less a coherent movement than a constellation of rejections — of realism, of progress, of the unified self... (Eysteinsson, 1990, p.117).

doi.org/10.7551/mitpress/4596.001.0001 · MIT Press, 1990
QUOTED · CITED
18

Eliot's emphasis on impersonality reframes the lyric as constructed artifact rather than confessional outpouring.

Eliot, T.S. (1919). Tradition and the Individual Talent · Egoist Press · doi.org/10.2307/24…
PARA · CITED
23

The novel emerges from a fundamentally bourgeois worldview that took shape in the eighteenth century.

openalex.org/W12345678 · Watt (1957) — uncited match
UNATTRIB
31

Joyce's stream-of-consciousness owes much to William James's psychology of interior duration.

James, W. (1890). The Principles of Psychology, ch.IX (mentioned in citations list)
PARA · CITED
34

The author's own previously published doctoral dissertation, Chapter 1.

ucl-discovery.com/dissertation/2024-Modernism-Chen.pdf · self-citation, no attribution this paragraph
SELF-CITE
41

All references — Eysteinsson, Eliot, Joyce, James, Watt, and 32 others.

Bibliography section, pp. 38–41
REFS · EXCLUDED
Headline 8.2% = unattributed paraphrase + self-cite weight; cited intervals + bibliography excluded.
  • 700M+scholarly works indexed
  • 6citation styles auto-detected
  • 5languages calibrated at launch
  • 3%ESL-bias residual (documented)
  • Apache2.0 · IRB-reviewable
PEER-REVIEW DEFENSIBLE

Three things a black-box checker can't survive in a faculty hearing.

01 · REPRODUCIBLE SCORE

Re-run the same paper against the same commit

Every report stamps X-Engine-Commit and a corpus snapshot timestamp. Months later, you can re-check the same manuscript against the engine version that produced the original score — same input, same algorithm, same result. Closed-source checkers can re-train silently and your January score won't reproduce in May.

Reproducibility is the academic-integrity floor we built to.
02 · AUDITABLE ALGORITHM

Apache 2.0 source on GitHub — bring it to the meeting

When AI detection ships (coming soon, on the v1.2 roadmap) and a student disputes their AI verdict at an integrity board, you'll be able to cite the engine commit, point to the line of code that produced the threshold decision, and show the per-language ESL calibration documentation. Not a screenshot of a vendor support reply.

github.com/NoplagLabs/noplag-engine — public, indexed by Google Scholar.
03 · ESL-AWARE VERDICT

Liang et al. 2023 documented; mitigation on the roadmap

Stanford 2023 measured 61.3% FPR against TOEFL essays on naive detectors. When Noplag's AI detection rolls out (coming soon, v1.2 roadmap), the verdict will carry an ESL-adjusted band when the lexical signature suggests non-native English. Residual error will be documented, not claimed to be zero. Override path logged with engine commit + reason.

/docs/developers/benchmarks — PAN-PC-11 methodology and scores.
PROCUREMENT-FRIENDLY

An Apache 2.0 engine your institution's IT can actually review before signing.

Most institutional procurement reviews of plagiarism software stall on the same questions: where does the document go after submission, what's the matching algorithm, who has access to the data, what happens at right-to-erasure, and can you replicate the vendor's accuracy claim. With a closed-source vendor, those questions get answered in a sales call. With Apache 2.0, IT can clone the repo, run the engine on a sandbox VM, audit the data flow, verify the right-to-erasure pathway, and reproduce the published F1 scores against their own test set. We're not arguing for cheaper review — we're arguing for a review that the IT team can actually conclude. That's why universities that previously refused to add new plagiarism vendors are signing this one.

Read the procurement guide
FAQ

The questions academics actually ask.

Does this replace iThenticate for journal pre-submission?

For most cases, no — if your target journal mandates iThenticate (Elsevier, Springer, Wiley, IEEE workflows), they'll run it at submission regardless. What Noplag replaces is the iterative self-check across drafts before that submission. iThenticate at $125/document gets prohibitive after the 4th revision; flat-rate Pro doesn't. See the full vs-iThenticate comparison at noplag.com/vs-ithenticate.

What about ProQuest dissertations? You don't have those?

Correct — that gap is real. ProQuest's dissertation database is proprietary, paywalled, and not in OpenAlex/CORE. We cover the ~700M open-access scholarly works, which is broader than ProQuest for post-Plan-S research (most major journals 2018 onwards); narrower for older paywalled dissertations specifically. If cross-checking against US PhD dissertations is your dominant requirement, iThenticate's the right tool for that one job.

How does the engine handle citation styles other than the big six?

MLA, APA, Chicago / Turabian, Harvard, Vancouver, and IEEE auto-detect from the bibliography format. For less common styles (ACS, AMA, ASA, Bluebook, Oxford), the engine still classifies intervals based on quotation marks + nearby parenthetical patterns, but the style-specific in-text validator is on the v1.2 roadmap. Self-host users can patch in additional styles via the alignment/ module.

What's the false-positive rate on heavily cited literature reviews?

Lower than uncited-paraphrase reviewers expect, because the classification model excludes properly-cited intervals from the headline score. We measure ~3.4% false-positive rate on the PAN-PC-11 academic-prose subset (where most intervals are legitimate scholarly quoting). The closed-source competitors who don't classify cited intervals report 12–18% on the same corpus.

Can I exclude my own previous publications?

Pro tier — add a self-citation allowlist (your ORCID, your past DOIs, your institutional repository URL). Matches against allowlisted documents tag as SELF-CITE and don't count toward the score. Useful for academics whose grant proposals build on their previous published work; self-recycling with citation is expected, not plagiarism.

What's the multilingual coverage actually like?

Detection is calibrated for English, Spanish, Portuguese, Polish, and Ukrainian at launch, with more languages rolling out. Detection quality is benchmarked on PAN-PC-11; per-language evaluation is on the roadmap. Strongest on the European languages with the longest validation history; newer additions ship as their calibration matures.

What does “Apache 2.0” mean for institutional procurement?

Your IT / DPO / IRB can clone github.com/NoplagLabs/noplag-engine, run the engine on a sandbox, audit the data flow, verify right-to-erasure works end-to-end, and reproduce our published benchmark numbers against their own test corpus. Without signing anything, without an NDA, without trusting our marketing. For institutions whose procurement process requires this depth of review, that's the unlock.

Will the AI detector work on Claude-3 + Gemini text, or just GPT?

AI detection is coming soon — it's on the v1.2 roadmap, not live at launch. The approach is Binoculars (Hans et al. 2024), which is model-agnostic: it doesn't train per-LLM, it compares a small reference model's perplexity against the test text, which generalises. Accuracy figures publish when the feature clears evaluation, rather than being quoted ahead of it, on mixed-LLM content (vs ~70% on Claude-3 for GPT-trained detectors). The methodology paper documents the failure modes; we don't hide them.

What if my advisor disputes a flag?

Every flagged interval ships with the source DOI, the corpus snapshot timestamp, the engine commit, and the classification reason. Re-run the check against that engine commit and the result reproduces exactly. Bring the printed report to the meeting; the citation paths are click-through. We don't have a “trust the score” mode, because that's what black-box checkers have.

What about self-host for medical / clinical research with HIPAA?

Apache 2.0 + Docker — runs entirely inside your network. Pair with PostgreSQL + Redis on your own infra; no traffic leaves the VPC. Useful for clinical drafts, patient case studies, or any pre-publication research that can't go to a third-party cloud. Enterprise customers get the corpus shipped as an indexable bundle.

Run a chapter against 700M+ scholarly works.

Drop in a dissertation chapter, grant proposal, or manuscript draft. The citation-aware report shows what's expected, what needs a citation, and what's your own. Free up to 2,500 words. Apache 2.0 if your IRB needs the algorithm reviewed first.

Academic Plagiarism Checker — Citation-Aware | Noplag