Why can't I select the text in this PDF?

Search for a way to extract PDF text and you get a box of words and a copy button. You paste it somewhere, it comes out empty or as noise, and you are back where you started. The question that would actually have helped — which of five things is stopping this file? — is not answered by any of them. So it is answered here, on your file, in your browser, in one click.

  • Nothing is uploaded
  • Tells you the cause, not just the result
  • No OCR, no upload, no upsell
  • A second reader confirms every count

Drop the PDF you can't select

or paste from the clipboard ·

PDF only · one file at a time · the whole check runs on your device

Every extractor in the search results will extract text from a PDF for you, and many will copy text from a PDF and drop it in a PDF to text box on the same tab. None of them says why one particular file hands nothing back, so you are left trying to highlight a line that will not highlight, pasting into something, and getting an empty box with no explanation. This page starts one step earlier and answers the question underneath: is there a text layer in this file at all, or is the page an image only? It opens your PDF in this tab, reads the drawing operators on every page, and reports what is really there — selectable text you can take, or pixels that only look like characters. Nothing is uploaded to make that happen.

Five things stop a page from having selectable words

They are not equally common, and they do not have the same fix. Reading the list is enough to work out which one you have before uploading anything, and the check below tells you which one your file actually is.

1. The page is a picture, so there is nothing to select

A scanned PDF, a camera photo of a screen, a file that was printed and then re-scanned. The words exist as pixels the way a photograph of a sign exists as pixels. There are no text objects in the file, no font, no operator saying "draw a character here" — only a picture of some letters. Highlighting does nothing and there is nothing to copy, because the information is not in the file in the first place. This is the case, by far the most common one, and it is the only one of the five that a person can usually see for themselves by zooming in.

2. The page's own objects are packed away

Since the mid-2000s writers have been allowed to pack small objects — dictionaries, and sometimes the content that draws them — into a single compressed container called an object stream. A reader that does not decompress those sees a file with no pages in it. Most readers decompress them; some do not, and a page that does not can look like an empty one. This is a difference between readers rather than a fault in your file, which is why the check below counts them and reports the number instead of quietly failing.

3. The font knows its glyphs but not their letters

This is the nastiest of the five, because the text layer is genuinely there. The file draws glyph 0x1F2, and glyph 0x1F2 exists — the right shape for the character you wanted — but nothing in the file records that the shape is a letter. Without a character map, the only way to name a glyph is the encoding the font claims. When the font is a subset font — one that keeps only the glyphs this document uses — its codes were invented for this one document, so guessing costs you a letter a time and produces a line of noise that is the right length. The report names the font when it meets this, so you can tell a missing map apart from a broken tool.

4. The document is encrypted

An encrypted PDF does not keep its words in readable form. Deciding whether that is the case takes one look at the file's structure, so the check below looks first and says "this one is encrypted" before it reads anything else. It will not try to open it without your password, and it will not pretend to.

5. The structure does not survive the file

A download that came back truncated, a disk that was pulled mid-write, a converter that stopped halfway. Part of the page tree is missing, and the page that survived may have lost the content stream it needs. The check below says which stream it could not find rather than reporting the page as empty.

What the checker on this page actually looks at

It is not a renderer, and it is not a wrapper around somebody else's PDF library either. It reads the same object list our PDF compressor already reads, inflates each page's content stream with the browser's own decompression, and walks the drawing operators in the order they were written — Tf to set a font, Td to move down the page, Tj to show a string — translating each string through that font's character map. Five things come back:

  • how many characters each page carries, and the verbatim words where there are any
  • which pages have no text, and the reason each one gives up
  • every font on the page, and whether it has a character map or not
  • how many object streams the file packs its objects into
  • whether an encryption dictionary is present

Those numbers are not taken on trust. On a two-page test document — one page of body text around a photograph, one short paragraph after it — the checker counted 2,432 characters on page one and 76 on page two. A text-only page came out at 603, and four photo pages at six apiece. All of those were matched exactly by a second, independent PDF reader running separately on the same files, and the comparison is rerun every time the module changes. Where the two disagree on line breaks they are expected to: one watches the vertical position drop, the other has its own idea of where a paragraph ends, and neither claim to reflow the text.

What it deliberately does not do

It does not run OCR. Reading letters back out of a picture is not something a browser can do without a model measured in megabytes, and shipping one here would either cost you that download or make a network call with your document attached — the same trade the extractors in the search results mostly make, one way or the other. This page is built to answer the other question: the file has words, something is stopping you, and here is which something.

It also does not rewrite your file. Once you know the cause, the next move belongs to whoever made the document — a writer that did not write character maps, an OCR pass you have not run, a copy that was not truncated — and the tools here are for what you do afterwards: shrink a PDF without flattening the words, hit an upload limit and keep the text, or put a corrected page back with the rest.

How to read the result

The summary line answers "can anything be copied out of this?" and the notes under each page answer why not. A page with a character count and a list of words is a file where the text was always there. A page that says the words are pixels is a scan, and copying harder will not change it. A page naming a font and telling you the map is missing is a file you should take back to whatever produced it. A page reporting an encryption dictionary is a file you need a password for. And a page reporting a stream it could not find is a damaged file you should get again from wherever it came from.

Each of those is a sentence about your own document, printed on the page, from a file that never left your machine.

Frequently asked questions

Why can't I select the text in my PDF?

One of five things is true, and which one it is changes what you should do next. The page is a picture rather than text — a scan or a photograph, with the words drawn as pixels and no text objects anywhere in the file. The words are inside compressed object streams. The font that draws them has no character map, so the file knows it drew a glyph but not which letter it was. The document is encrypted. Or the structure is damaged. The tool on this page separates those cases for your own file and names the one that applies, page by page, instead of handing you a blank box.

How do I know whether my PDF is a scan?

Open the file in any viewer and zoom in hard. If the letters have soft, shimmering edges and you can see the paper grain through them, you are looking at a picture of a page and there is no text in the file. The check that does not depend on your eyes is whether the content stream for the page contains any text-drawing operators at all. A page that answers “zero characters” with no reason attached is the one to distrust; a page that says “this page has no text objects, its words are pixels” is telling you the same thing a scan would, and it is also telling you that no amount of copying will produce a word.

Can you run OCR on it for me?

No, and this page is deliberately built around what it does not do. Reading a picture of a page back into letters is OCR, every part of it runs on pixels, and doing it here would mean either shipping a font and a model measured in megabytes or sending your file somewhere for the job. Both are trade-offs this page is not going to make silently: the first costs the visitor a long download, the second costs the visitor their confidentiality. So this tool answers the other question — the one that does have a browser-only answer — which is whether the words are already in the file and something is stopping you from having them. When they are not, it says so, and then you know that OCR is the move rather than another extractor.

The text comes out as gibberish. Why?

Almost always a font that was embedded as a subset without a character map. When a writer embeds a font to draw one page, it only carries the glyphs that page needs; the codes it uses to name those glyphs are then local to that file. Without a ToUnicode map recording which code is which letter, the file can say it drew a shape at code 0x1F2 but not that the shape is a letter. Readers fill the gap with the encoding the font claims — WinAnsi, MacRoman, Standard — and produce nonsense. The report names the font for you, so you can see that the file itself is missing the map rather than the tool failing.

My PDF is password protected. Can you read the text?

No. An encrypted document does not store its words in readable form, and decrypting them needs the password. This tool looks for an encryption dictionary and reports the file as encrypted before reading anything else, rather than trying and failing halfway. If you have the password, the usual route is to open the document in the tool that made it or another reader that accepts a password, save an unlocked copy, and run this page on that copy.

Is my PDF uploaded to a server?

No. The file is opened with the browser's own reader and every step — reading the object list, inflating the content streams, reading the character maps, walking the operators — happens on your machine, in this tab. Worth comparing against the search results for this query: the extractors that rank for it mostly run in the browser too, and several of them say outright that a page with no text layer gets sent to a larger OCR service for a second attempt. That is a legitimate trade, but it is a different one, and it is the sort of thing a page that only sells you the output will not mention.

The words really are in there. Why did this find none?

Read the notes on the page rather than the character count, because a zero is a result here, not a failure. If the note names a font, the file is missing its character map. If it names a filter, the page's stream is compressed with something the browser cannot undo and the tool will say so rather than guess. If it says the stream is missing, the file is damaged. Any of those three means a different reader may still do better; what none of them means is that the words were never there.

Can I fix the file once I know what is wrong?

It depends entirely on which of the five it is, which is the reason this page reports the cause instead of just the text. A missing character map is fixed by whoever produced the file — a writer that writes ToUnicode maps. A scan is fixed by OCR into a searchable copy first. A damaged file is fixed by the copy that did not get truncated. What is never fixed by an extractor is the selection problem itself: once you have the words out, the useful next steps are ones this site already covers, such as shrinking the file without flattening the words, or merging a corrected page back in.