ToolMint

PDF Redaction Checker

Runs in your browser

This tool runs entirely in your browser. Your file is never uploaded and never leaves your device.

A black box over a name does not remove the name. In most leaked documents the text is still sitting in the file, and anyone can copy it out in seconds. This tool reads a PDF in your browser and tells you whether any text is still recoverable β€” before you send the file, not after.

Loading redaction checker…

Who this is for

Anyone about to publish a document that has been redacted by someone else, or by themselves in a hurry: lawyers filing exhibits, journalists handling leaked material, FOI and public-records officers, compliance and HR teams, and researchers sharing data. The check takes seconds and answers one question β€” is the removed text really gone?

What it checks for

Text under a black box

The classic failure: a filled rectangle is drawn over a name or number, but the text underneath is never deleted. Selecting the area reveals it. This is the mistake behind most published redaction leaks.

Text under a pasted image

Same failure, but the cover is an opaque image rather than a drawn shape β€” what you get when someone patches a screenshot over a line of text.

Invisible text

Text set to render invisibly. It never appears on screen or in print, but any extractor finds it. Scanned pages with a normal OCR layer are excluded, because that is how searchable scans work.

Same-colour text

White text on a white page, or any text painted the same colour as what sits behind it. It looks blank and is completely intact in the file.

Document properties and attachments

Title, Author, Subject, Keywords and embedded files travel with the document. These are reported for review rather than as a failure β€” nearly every PDF has them.

How to check a redacted PDF

1

Choose a PDF

Drop in the document you are about to share or publish.

2

Run the check

Every page is parsed in your browser. Nothing is uploaded.

3

Read the result

Any recoverable text is shown with the page it came from.

4

Fix and re-check

Redact properly, then run the file through again to confirm.

What this tool cannot check

A security tool that overstates what it verifies is worse than no tool at all. These are the things this checker does not and cannot establish:

  • Non-text content hidden under a box. A signature, photograph, chart or map is invisible to text analysis β€” this checker reads text only.
  • Whether the text that is visible should have been removed. It finds concealed content; it cannot judge sensitivity.
  • Content preserved in an earlier incremental revision inside the file. A PDF can retain previous saved states, and reading those is outside these checks.
  • Anything inside a password-protected PDF it cannot open.
  • Vertical writing modes, which the text-position calculation does not model.

A clear result means β€œthese checks found nothing recoverable”. It does not mean the document is safe to publish.

Methodology

The checker walks each page's content stream in draw order and rebuilds every text run with its position, colour and rendering mode. Draw order is the whole game: a filled box only conceals text that was painted before it. An early revision of this engine ignored ordering and flagged sixteen cells of an ordinary shaded table, plus a report cover's own title, as β€œhidden text”. For a tool that makes a security claim, a false alarm on a normal document is more damaging than a missed edge case.

Only opaque, axis-aligned rectangles and images larger than four points a side count as covers. Semi-transparent fills are highlights, not redactions. Rotated text is measured by transforming all four corners of its box rather than padding a baseline, because a rotated chart label otherwise collapses into a sliver that any nearby marker appears to cover.

It is validated against sixteen purpose-built fixtures β€” nine leak types, four legitimate documents that must not flag, two flattened scans, and one case it is expected to miss β€” every one of which is independently verified with PyMuPDF, a different PDF library, so the tests are not circular. Three further fixtures cover encrypted, corrupt and 300-page documents. It is additionally measured against nine real published PDFs (US government publications, IRS forms, arXiv papers and a shareholder letter; 169 pages, roughly 26,700 text runs), on which it currently reports zero false positives.

Results were also compared against x-ray, the Free Law Project's open-source bad-redaction detector, which is built on a different parser. The two agree on every case inside x-ray's documented scope of rectangles over text. This tool additionally reports invisible text, same-colour text, image overlays and document properties, which x-ray does not claim to cover.

How this was tested

Fixtures are built and verified with PyMuPDF, a different PDF library from the one this checker runs on, so the expectations are not derived from the code being tested. These figures describe how it performed on that corpus. They are not an accuracy guarantee for your document, and they do not extend to the limitations listed above.

Measured results on the test corpus
ResultCountBasis
Leaks correctly found9Each independently confirmed recoverable by PyMuPDF
Correctly reported clean15Six benign fixtures plus nine real published PDFs
False alarms0On that corpus, after fixing the causes described below
Known misses1A signature under a box β€” non-text content, kept as a deliberate failing case

The real-world half of that corpus was nine published documents from different generators β€” US government publications, IRS forms with heavy shading and form fields, two LaTeX papers with figures, and a shareholder letter β€” totalling 169 pages and roughly 26,700 text runs.

Two false-alarm sources were found during that measurement, and both mattered more than the detections. A four-page scanned government memo produced 878 findings, one for every word: a perfectly normal OCR layer. Uncorrected, this tool would have fired on every scanned document in existence. Separately, an early revision flagged sixteen cells of an ordinary shaded table because it ignored draw order.

How this compares with x-ray

x-ray is the Free Law Project’s open-source bad-redaction detector, built on a different PDF library and run across a very large corpus of court filings. Several of the rules used here were adopted from it. The two tools have different scopes, not different quality: x-ray documents its scope as rectangles over text, and anything outside that is simply not what it sets out to cover.

Scope comparison, verified on identical fixtures
Casex-rayThis checker
Text under a rectangleDetectsDetects
Ordinary documents (no false alarm)CleanClean
White-on-white textOutside its scopeDetects
Text under a pasted imageOutside its scopeDetects
Document propertiesOutside its scopeReports for review
Runs without installing anythingNo β€” Python library and CLIYes β€” in the browser

On identical fixtures the two agree on every case inside x-ray’s documented scope. For processing PDFs in bulk or inside a pipeline, x-ray is the better fit; this tool exists for the person with one document and no ability to install software.

How to redact a PDF properly

The principle, stated by the NSA in its guidance on sanitising documents, is that sensitive information must be removed, not merely hidden from view. Drawing a shape over text, highlighting it in black, or recolouring it to match the page all leave the original characters in the file.

Practical approach: redact with a tool that deletes the text or flattens the page to an image, clear the document properties, then re-open the finished file and try to select the redacted area. If you can copy anything out, the redaction failed. Running the result back through this checker is the same test, automated.

ToolMint's Redact PDF rebuilds each redacted page as an image, so no text layer survives β€” then check the output here to confirm.

For the failure patterns themselves, and the manual tests that catch each one, see how to tell if a PDF redaction failed.

Frequently Asked Questions

Does this upload my PDF?
No. The document is read inside your browser tab using a local copy of pdf.js, and no part of it is sent anywhere. You can verify this yourself: open your browser's Network tab before running the check and confirm no request carries your file. This matters because a document you are redacting is, by definition, sensitive.
What does 'no recoverable text was detected' actually mean?
It means the specific checks this tool performs found nothing recoverable. It is not a certificate that the document is secure. In particular, non-text content under a box, an earlier revision inside the file, or sensitive information that is simply visible on the page will all pass this check.
Why did it flag text on a page that looks fine?
Because the text is present in the file even though it is not visible to you. Text under an opaque box, invisible text, and text painted the same colour as its background all render as nothing but stay fully extractable. Select the reported area in your PDF reader and you should be able to copy it out.
Why did my scanned document not get flagged for invisible text?
Searchable scans are an image with an invisible text layer on top β€” that is how the text is selectable at all. When essentially all of a page's text is invisible and a large image covers the page, the tool treats it as an OCR layer and says so, rather than reporting every word as hidden content.
How do I redact a PDF so it passes?
Use a tool that removes the underlying text rather than drawing over it. ToolMint's Redact PDF rebuilds each page as an image, so the text layer does not survive at all. Whatever you use, re-run the finished file through this checker before you send it.
Is this a complete PDF security audit?
No, and it is not offered as one. It tests for text that is present but concealed, plus document properties worth reviewing. The limitations listed on this page are part of the tool, not a disclaimer bolted on afterwards.

Related Tools