When I built the viewer and the OCR tool in the lab on this site, I thought the main problem with PDFs was compatibility between programs. I was wrong: the most common problem is that a huge number of PDFs contain no text at all. They're photos of pages. They look perfect, but to the "Find" function, to a search engine and above all to people using a screen reader, they're empty.
From then on I started looking at PDFs differently. In this article I bring together what really makes a PDF accessible and the practical checks I use.
The check that opened my eyes
In my OCR tool, the first step for each page is to ask the PDF for its text layer. If there is one, I use it as is: it's instant and accurate. Only pages without text are turned into images and sent to optical recognition:
const content = await page.getTextContent();
if (content.items.length === 0) {
// No text layer: this page is just a picture.
// Search, copy-paste and screen readers get nothing.
queueForOcr(page);
}
That check, written to save time, is also the simplest accessibility test there is. In the viewer I made it visible: when a document has no extractable text, a notice suggests OCR. Without it, users search for a word, get zero results and assume search is broken; I wrote about it in this article.
What a PDF needs to be accessible
Real text is only the first requirement. The reference standard is PDF/UA (ISO 14289), consistent with the WCAG guidelines. In practice, an accessible PDF has:
- Real text, not images of text.
- Structure tags: headings, paragraphs, lists and tables marked as such, so a screen reader knows what it is reading.
- A correct reading order: in a two-column layout without tags, text may be read line by line straight across the columns.
- Alternative text for images that carry information, such as charts and diagrams.
- The document language set, so text-to-speech pronounces it correctly.
- A title in the metadata, not just the file name.
- Labelled form fields, if the PDF is fillable.
Since June 2025, under the European Accessibility Act, accessibility is mandatory for many consumer-facing digital services, and the documents those services provide, such as PDF contracts and invoices, are part of the experience.
The practical checks I run (in five minutes)
- I try to select text and search for a word with Ctrl+F. If nothing happens, the PDF is an image.
- I open the document properties in Acrobat Reader: "Tagged PDF" must say "Yes", and the title and language must be filled in.
- I use the PDF reader's Read Out Loud on a two-column page: if it jumps between columns, the reading order is wrong.
- I copy a table and paste it into a spreadsheet: if it arrives as a single block of text, there's no structure.
- For a full check I use PAC (PDF Accessibility Checker), a free tool that verifies the PDF/UA requirements one by one.
A typical case: the scanned form
The case I see most often is a printed form, filled in by hand and scanned back in. From an accessibility point of view it's the worst possible: no text, no structure, no fields. OCR makes it searchable and copyable, which is already a huge step forward, but it doesn't make it an accessible PDF: tags, reading order and labels are still missing. My OCR tool, for instance, extracts the text but doesn't rebuild the structure: it's the first step, not the last.
The real fix is upstream: publish the original document, not its scan.
How to create accessible PDFs from the start
- In Word or LibreOffice, use styles (Heading 1, Heading 2, real lists) instead of enlarged bold text: they become the PDF's tags.
- Add alternative text to images before exporting.
- Export with tags enabled: in Word it's the "Document structure tags for accessibility" option when saving as PDF.
- Avoid "printing" to PDF through a virtual printer or from an image: the structure is often lost.
- Set the title and language in the file properties.
In short
An accessible PDF isn't a "special" PDF: it's a well-made PDF, with real text, structure, reading order, alternative text and a language. Building my OCR tool taught me that the most revealing test is also the most trivial: try searching for a word. If you can't find it, neither can a screen reader. For documents you only have as scans, OCR is the first step to make them at least readable; for everything else, accessibility is decided the moment you create them.