Plagiarism checker for law firms.
Pre-filing check for briefs, memos, and motions. Detects unattributed lifts, recycled boilerplate, and AI-drafted text (with the hallucinated-case risk that comes with it). Privilege-preserving: self-host the engine inside your own network, or EU-residency for cross-border matters. Documents never used for training. Ever.
Four roles. One pre-filing self-check.
Pre-partner-review brief check
Catch the uncited paraphrase that's missing a Bluebook citation before the partner does.
- Self-check before the partner sees it
- Bluebook + ALWD citation-aware classification
- AI-likelihood per paragraph (if you used an LLM)
Filing-ready verification
The signature is yours. The originality should be too.
- Firm-wide self-citation allowlist (your prior briefs)
- Hallucinated-case detection (Mata v. Avianca was a warning)
- Audit trail by associate for malpractice prevention
Discovery + memo prep workflow
Bulk-check memos and discovery responses before they hit the case file.
- Folder per matter · case-file isolation
- EU residency for cross-border discovery
- Apache 2.0 engine — defensible at conflicts review
Firm-wide AI-policy enforcement
Verify the firm's AI-use policy is actually being followed.
- AI-detection across the firm's outgoing work product
- Right-to-erasure logs for client-confidentiality audits
- Self-host inside the firm's network — privilege preserved
Three steps before the e-file button.
Self-hosted inside the firm's network, EU-residency endpoint for cross-border matters, or cloud Pro with strict tenant isolation. Pick the deployment that fits the matter's confidentiality requirement.
01 · Draft inside the firm
Upload the brief / memo / motion (.docx most common; .pdf supported). The submission lives in the firm's tenant — privilege-preserving access controls, no shared corpus, no cross-tenant exposure. Self-host users: the document never leaves your VPC.
02 · Check against the right sources
The cascade scans against open web + Common Crawl + open opinion archives (CourtListener, Caselaw Access Project, Google Scholar, Justia). Citation classification distinguishes properly-Bluebooked case quotes from uncited paraphrase. AI-likelihood flagged per paragraph — relevant because hallucinated cases are a documented courtroom risk.
03 · Verify before filing
The report classifies every interval: QUOTED·CITED (Bluebook attribution, excluded), CASE·QUOTE (full case-name + reporter cite, excluded), PARA·CITED (paraphrase with citation, low weight), UNCITED·PARA (review required), AI·LIKELY (verify or rewrite), HALLUCINATED·RISK (cited case not found in any open opinion archive — investigate before filing).
Six things to verify before filing a brief that touched a chatbot.
Mata v. Avianca (S.D.N.Y. 2023) was the warning shot — attorneys cited six cases that ChatGPT invented; the court sanctioned them $5,000 and the firm publicly. AI-drafted legal text needs its own verification layer.
01 · CITATION EXISTENCE
Every cited case is cross-checked against the open opinion archive (CourtListener, Caselaw Access Project, Google Scholar opinions, Justia). If a cited case-name + reporter cite returns zero hits across all four sources, the report flags HALLUCINATED-RISK. Doesn't mean the case is fake (could be unpublished, sealed, or paywalled); does mean verify before filing.
02 · QUOTE FIDELITY
When a brief quotes a case, the engine checks the quoted text against the actual opinion. If the cited case exists but the quoted text doesn't appear in the opinion, that's QUOTE-DRIFT — a classic LLM hallucination pattern. Flagged separately from the citation-existence check.
03 · PINCITE ACCURACY
Pin-cite verification (page numbers within an opinion). The cited page either contains the quoted text or it doesn't. Common LLM failure mode: real case, real page number, but the quoted text is from a different paragraph. Flagged PINCITE-MISMATCH.
04 · HOLDING vs DICTUM
Beta feature (Pro+). The engine cross-references whether the cited proposition is part of the case's actual holding or a piece of dictum. Misrepresenting dictum as holding is a common LLM error and a malpractice-adjacent risk.
05 · OVERRULED / VACATED
Citation history cross-check via CourtListener's case-history API. If a cited case has been overruled, vacated, or distinguished by later precedent in the relevant circuit, that surfaces in the report. A correctly-cited opinion can still be the wrong opinion to cite.
06 · AI-LIKELY PARAGRAPHS
Per-paragraph AI-likelihood (Binoculars-based) is coming soon — on the roadmap, not live at launch. Flagged paragraphs will be the ones that need human verification before filing. ESL-adjusted for international firms with non-native-English attorneys — the verdict labels as ESL ADJ rather than LIKELY-AI when the lexical signature warrants it.
What you see before the e-file button.
Per-paragraph classification: cited cases verified, paraphrases attributed or not, AI-flagged paragraphs, and the all-important HALLUCINATED-RISK flag for any cite the engine couldn't find in the open opinion archive.
Three things attorney-client confidentiality needs from a checker.
The engine runs inside your firm's network
Apache 2.0 + Docker. Pair with your existing Postgres + Redis on your own infrastructure. The submission, fingerprint store, and report stay on your servers. No traffic leaves your VPC. The right answer for matters under protective order, sealed-discovery materials, or work-product-doctrine-sensitive drafts.
For cross-border matters touching EU jurisdiction
Pro+ provisioned to api.eu.noplag.com. Hosted on EU-based Hetzner data centres (Falkenstein + Helsinki). The submission, fingerprint store, and report stay in the EU. GDPR-compliant by default; DPA available on request. Useful for EU + UK cross-border representation, GDPR-related matters, and EU-domiciled clients.
Submitted briefs do not become training data
Drafts are fingerprinted (winnowing hashes, non-reversible signature) and stored in tenant-isolated database. The text itself is purged after 30 days unless retention is enabled. Detection models are pre-trained on public corpora — we don't add law firm documents to any training set. Explicit in the DPA.
Two things a multi-jurisdictional matter needs that single-jurisdiction tools miss.
Most plagiarism tools were built for one jurisdiction's case law — usually US federal. Cross-border matters touch multiple opinion archives (US federal + state + EU member-state + UK + Commonwealth), multiple citation conventions (Bluebook, OSCOLA, McGill Guide, AGLC), and multiple languages (legal drafting in French, German, Spanish, Portuguese for South American matters). The cascade handles all of that. The EU-residency endpoint handles the data location. Self-host handles the matters where neither cloud option is enough.
Read the cross-border matter guideWhat managing partners ask before procurement.
- Will our briefs become training data?
- No. Drafts are fingerprinted (winnowing hashes, non-reversible) and stored in tenant-isolated database. Text is purged after 30 days unless retention is enabled. Detection models are pre-trained on public corpora; we do not add law firm documents to any training set. Explicit in the DPA. Self-host users keep everything inside their own VPC — even fingerprints never leave.
- How does the hallucinated-case detection actually work?
- Every case-name + reporter citation in the brief is cross-checked against four open opinion archives: CourtListener, Caselaw Access Project (Harvard), Google Scholar opinions, and Justia. If a citation returns zero hits across all four, the report flags HALLUCINATED-RISK. Doesn't mean the case is fake — could be unpublished, sealed, or paywalled — does mean verify with Westlaw / Lexis before filing. Mata v. Avianca was the warning shot we built this for.
- Does it work for non-US case law?
- US federal + state, UK + Commonwealth (OSCOLA citation style), Canadian (McGill Guide), Australian (AGLC), and EU (ECJ + ECHR). The cascade indexes open archives from Bailii (UK), CanLII (Canada), AustLII, and the official EU Curia opinions. Older case law behind paywalls (Westlaw / Lexis / juris) won't surface as matches — same gap iThenticate has on academic paywalled content.
- What's the privilege story for sensitive matters?
- Three tiers. Cloud Pro: tenant-isolated database, EU residency available, 30-day text purge. EU residency endpoint (api.eu.noplag.com): hosted entirely in EU Hetzner data centres, GDPR Art.28 processor. Self-host: Apache 2.0 + Docker, runs inside your VPC, no traffic leaves the network. Pick the tier the matter's protective order requires.
- Does it handle Bluebook citation properly?
- Yes — Bluebook 21st edition, plus ALWD Guide, OSCOLA (UK), McGill Guide (Canada), AGLC (Australia). Citation format auto-detected from the brief's existing references. Properly-cited cases excluded from the headline similarity score. Incorrectly-formatted citations flagged separately — not as plagiarism, as citation-format drift you might want to fix before filing.
- What about discovery materials covered by protective order?
- Self-host the engine inside your firm's network — Apache 2.0 + Docker. Material under protective order, sealed deposition transcripts, and trade-secret discovery materials should never go to a third-party cloud regardless of the vendor's claims. The self-host option means “never leaves your VPC” — not “trust our cloud”. Most firms doing high-stakes commercial litigation use this deployment.
- Does the AI detector survive a sanctions hearing?
- AI detection is coming soon — on the roadmap, not live at launch — but the methodology is open (Binoculars, Hans et al. 2024, peer-reviewed paper). Every flagged paragraph will stamp the engine commit and corpus snapshot. Re-run later against the same commit and the verdict reproduces. Closed-source AI detectors fail this test — courts have started excluding their output as Daubert-inadmissible because the methodology isn't reviewable. Ours is reviewable; we're not claiming the verdict is perfect, but it's defensible.
- What's the pricing for a 40-attorney firm?
- Premium at $79/mo covers per-attorney usage with 5,000 words per check, 5,000 documents per month, REST API + LMS integrations. For firms with bulk discovery review (10,000+ documents/month), Enterprise contracts are negotiated on corpus + dedicated infrastructure dimensions. Per-attorney seats not part of the model; volume is what scales the price.
- Can we integrate with our document management system (iManage, NetDocuments, OpenText)?
- Pro+ tier provides webhook + REST API integration. iManage and NetDocuments both support outbound webhooks on document-save; configure the webhook to POST to /v1/checks; the report attaches to the document as a metadata field. OpenText eDOCS supports the same via Content Server hooks. Sample integration scripts for all three at github.com/NoplagLabs/noplag-engine/integrations.
- What about ediscovery — does this work for reviewing produced documents?
- Different use case from brief checking — ediscovery review is about classifying produced documents (responsive / non-responsive / privileged) rather than checking your own brief for plagiarism. Noplag isn't an ediscovery review platform (Relativity / Everlaw / Disco own that space). But for checking that your work-product responding to discovery is original and AI-free, Noplag fits.
Run the brief through the cascade. Before the e-file button.
Drop in the brief or motion. Get a per-paragraph report with citation verification, hallucinated-case detection, and AI-likelihood. Free tier up to 2,500 words; Pro for firm volume; Enterprise for self-host + dedicated infrastructure.