<link rel="stylesheet" href="/assets/fonts/jetbrains-mono/jetbrains-mono.css" />
All posts

Accessible PDFs: what you really need (what I learned building an OCR tool)

When I built the viewer and the OCR tool in the lab on this site, I thought the main problem with PDFs was compatibility between programs. I was wrong: the most common problem is that a huge number of PDFs contain no text at all. They're photos of pages. They look perfect, but to the "Find" function, to a search engine and above all to people using a screen reader, they're empty.

From then on I started looking at PDFs differently. In this article I bring together what really makes a PDF accessible and the practical checks I use.

The check that opened my eyes

In my OCR tool, the first step for each page is to ask the PDF for its text layer. If there is one, I use it as is: it's instant and accurate. Only pages without text are turned into images and sent to optical recognition:

const content = await page.getTextContent();
if (content.items.length === 0) {
  // No text layer: this page is just a picture.
  // Search, copy-paste and screen readers get nothing.
  queueForOcr(page);
}

That check, written to save time, is also the simplest accessibility test there is. In the viewer I made it visible: when a document has no extractable text, a notice suggests OCR. Without it, users search for a word, get zero results and assume search is broken; I wrote about it in this article.

What a PDF needs to be accessible

Real text is only the first requirement. The reference standard is PDF/UA (ISO 14289), consistent with the WCAG guidelines. In practice, an accessible PDF has:

  • Real text, not images of text.
  • Structure tags: headings, paragraphs, lists and tables marked as such, so a screen reader knows what it is reading.
  • A correct reading order: in a two-column layout without tags, text may be read line by line straight across the columns.
  • Alternative text for images that carry information, such as charts and diagrams.
  • The document language set, so text-to-speech pronounces it correctly.
  • A title in the metadata, not just the file name.
  • Labelled form fields, if the PDF is fillable.

Since June 2025, under the European Accessibility Act, accessibility is mandatory for many consumer-facing digital services, and the documents those services provide, such as PDF contracts and invoices, are part of the experience.

The practical checks I run (in five minutes)

  1. I try to select text and search for a word with Ctrl+F. If nothing happens, the PDF is an image.
  2. I open the document properties in Acrobat Reader: "Tagged PDF" must say "Yes", and the title and language must be filled in.
  3. I use the PDF reader's Read Out Loud on a two-column page: if it jumps between columns, the reading order is wrong.
  4. I copy a table and paste it into a spreadsheet: if it arrives as a single block of text, there's no structure.
  5. For a full check I use PAC (PDF Accessibility Checker), a free tool that verifies the PDF/UA requirements one by one.

A typical case: the scanned form

The case I see most often is a printed form, filled in by hand and scanned back in. From an accessibility point of view it's the worst possible: no text, no structure, no fields. OCR makes it searchable and copyable, which is already a huge step forward, but it doesn't make it an accessible PDF: tags, reading order and labels are still missing. My OCR tool, for instance, extracts the text but doesn't rebuild the structure: it's the first step, not the last.

The real fix is upstream: publish the original document, not its scan.

How to create accessible PDFs from the start

  • In Word or LibreOffice, use styles (Heading 1, Heading 2, real lists) instead of enlarged bold text: they become the PDF's tags.
  • Add alternative text to images before exporting.
  • Export with tags enabled: in Word it's the "Document structure tags for accessibility" option when saving as PDF.
  • Avoid "printing" to PDF through a virtual printer or from an image: the structure is often lost.
  • Set the title and language in the file properties.

In short

An accessible PDF isn't a "special" PDF: it's a well-made PDF, with real text, structure, reading order, alternative text and a language. Building my OCR tool taught me that the most revealing test is also the most trivial: try searching for a word. If you can't find it, neither can a screen reader. For documents you only have as scans, OCR is the first step to make them at least readable; for everything else, accessibility is decided the moment you create them.

💬 Reader notes

0 notes

Write a note

Share your opinion, a suggestion or a compliment

Latest notes

No notes yet. Be the first to comment!