Understanding your similarity score
What the percentage actually measures, what 0% and very high scores mean, and how to read a score responsibly.
What the number is
The similarity score is the share of your text that overlaps with known sources — the deduplicated union of every matched passage, so overlapping matches aren't counted twice. It measures textual overlap. It does not measure intent, originality of ideas, or whether a passage was properly cited.
The score's job is to tell you where to look. The passages and sources are the evidence; judge those.
| Score | Usually means | First thing to check |
|---|---|---|
| 0% | No verbatim overlap found in the checked corpora | Was the likely source in scope? Paraphrase won't match |
| A few % | Incidental overlap — names, phrases, boilerplate | Filter short matches; is anything substantive left? |
| High | Real shared text — quotes, reuse, or copying | Open the sources; quoted vs unquoted, cited vs not |
| ~100% | The document itself is published somewhere | Which source — an earlier version, or someone else's? |
Reading a 0% score
0% is not a certificate of originality
0% means no verbatim or near-verbatim overlap was found against the corpora your check ran on — nothing more.
Two gaps to keep in mind:
- Text plagiarized by heavy rewriting won't match — the shipped engine detects verbatim and near-verbatim copying, and says so (honest scope).
- Text copied from a source that isn't in the corpus won't match. The live-web layer narrows this gap for recent web content on Premium plans.
Reading a low score (a few percent)
Usually incidental: names, set phrases, boilerplate, common definitions. Use match sensitivity to exclude short matches and see whether anything substantive remains. A low score with one long matched passage matters more than the same score spread over twenty four-word fragments.
Reading a high score
Open the sources and look at what actually matched before drawing conclusions:
- Quotes and bibliographies match by design — quoted material is someone else's text. The quotes filter separates properly quoted passages from unquoted ones.
- Your own earlier drafts or previously published versions match if they're in a corpus your check ran against.
- Template language (legal boilerplate, methods sections, standard disclaimers) matches across many documents legitimately.
A high score built from one unquoted, uncited source is the case that deserves attention.
What about 100%?
A score at or near 100% almost always means the document itself is published somewhere the corpus covers — a webpage you pasted, a published article, an earlier submission. That's confirmation the engine found the original, not necessarily evidence of wrongdoing by whoever handed you the text — though it may be exactly that evidence, too.
If the score still looks wrong
Re-check the match sensitivity filters (an aggressive threshold hides real matches; Off shows everything), read the report guide, and if a specific match looks like an engine error, we'd genuinely like to see it — contact support with the check and the passage.