Of the 4,058 Prevention of Future Death reports published between 2013 and 2022, 2,112 of them (52.0%) are image-only scans containing not one character of extractable text. Not sparse text, or badly encoded text. Zero. Any analysis that screens these reports by reading the PDF's text layer could only ever have seen about half of them.
This is the single most consequential thing we found while building a full-text index of the archive, and it has a direct bearing on the growing body of published research that uses PFD reports as a data source.
Across all 6,283 reports in the corpus, 2,179 (34.7%) are image-only. But the proportion is not spread evenly. It collapses almost to nothing once the judiciary moved to publishing born-digital PDFs:
| Published | Reports | Image-only | Share |
|---|---|---|---|
| 2013–2019 | 2,903 | 1,955 | 67.3% |
| 2013–2022 | 4,058 | 2,112 | 52.0% |
| 2023–2026 | 2,225 | 67 | 3.0% |
So the blind spot is concentrated precisely where the historical record lives. For the years 2014–2018 the image-only share runs between 65% and 73%.
Our pipeline records, for every one of the 14,143 documents it processed, which method produced
its text: a PDF text-layer extractor (pdfminer or pdfplumber) or OCR.
A file is classed image-only when the text-layer extractors returned nothing usable and OCR was
required.
Counts here are of reports, deduplicated by the coroner's reference number, not of files: a few hundred reports have more than one document attached, and counting files instead would inflate every total by about 5%. It barely moves the percentages, but the absolute numbers should mean what they say.
To confirm that this classification means what we think it means, we re-ran a plain text-layer extraction over a random sample of files from each group:
The split is binary. These are not documents with degraded text; they are photographs of paper. A keyword search over their text layer does not return a poor result. It returns nothing at all, silently, and the report simply never appears in the result set.
A number of studies have now analysed PFD reports at scale by acquiring the PDFs from judiciary.uk and screening them for a topic (haemorrhage, thromboembolism, a drug, a device) using automated code. That is exactly the method the blind spot bites.
In our corpus, restricted to the 2013–2022 window that several of these case series use:
| Reports mentioning | Total | Image-only | Share |
|---|---|---|---|
| haemorrhage / hemorrhage | 253 | 132 | 52% |
| thromboembolism / pulmonary embolism | 135 | 71 | 53% |
And by category, again for 2013–2022. A single report often carries several categories at once, so a report counts here if the category appears anywhere in its classification. The clinical categories are the worst affected:
| Category | Reports | Image-only | Share |
|---|---|---|---|
| Hospital Death (clinical procedures and medical management) | 1,754 | 952 | 54% |
| Community health care and emergency services | 475 | 265 | 56% |
| State Custody related deaths | 240 | 125 | 52% |
| Care Home Health related deaths | 280 | 145 | 52% |
| Alcohol, drug and medication related deaths | 291 | 118 | 41% |
| Mental Health related deaths | 443 | 153 | 35% |
An important caveat. A simple keyword match is not the same thing as a study's eligibility screen, and we are not claiming that any particular published paper missed a specific number of eligible reports. Every report also has an HTML page on judiciary.uk carrying structured metadata (deceased name, coroner, area, category) which is machine-readable and which a screening pipeline may well have used. The claim here is narrower and, we think, harder to argue with: the body text of about half the pre-2023 archive cannot be read by software unless someone has run OCR over it. Whether that changed a given result is a question only that study's authors can answer, and it seems worth their asking.
Correction, 9 September 2026. The category table first published here counted only reports whose classification was exactly one category, which missed the many reports carrying two or more. That undercounted every row and overstated every percentage, because single-category reports skew older and so are likelier to be scans. Mental health, for instance, was shown as 33 of 57 (58%) when the correct figure is 158 of 458 (34%). The table above is the corrected version. The headline finding is unaffected: it is counted per report, not per category. Every figure on this page was also switched on 10 September from counting files to counting reports deduplicated by reference number, which moved the headline from 51.8% to 52.0%.
We OCR'd the lot. Of roughly 4,500 image-only PDFs, all but one now have full recovered text. The exception is a file that judiciary.uk serves as 0 bytes. That text is what the free search on this site queries, and it is why searching here returns reports that a text-layer search cannot surface.
You can see the effect directly. Search for a clinical term and compare the pre-2020 results with what you get elsewhere: reports that exist only as scans appear here with their full text, quoted and highlighted.
The underlying reports are public records and stay free to read and search. If you are working on PFD reports and want to compare notes on the OCR, or check any number in this post against your own pipeline, we would genuinely like to hear from you.