Skip to content
files.co

How to extract text from a PDF

Need the words out of a PDF as plain, editable text? Pull them out in your browser with files.co. The document stays on your machine the whole time.

AGAntonia González · July 29, 2026 · 5 min read

You’ve got a PDF and you need the words out of it. Not the layout, not the styling, just the text, so you can paste it into an email, drop it into a document, search it, or feed it somewhere else. Copy-pasting page by page works until it doesn’t: the selection jumps around, columns come out scrambled, and a long file turns into twenty minutes of fiddly dragging.

Pulling the text out in one go is faster and cleaner. Here’s how to do it without sending the PDF anywhere.

Two kinds of PDF, and it changes everything

Before you start, it helps to know which kind of PDF you’ve got, because they behave very differently.

A text-based PDF was made from a digital document, like exporting from a word processor. The text is real, selectable characters underneath. If you can highlight a sentence with your cursor and it selects cleanly, it’s this kind. Extracting from these is instant and accurate.

A scanned PDF is really a photo of a page. The “text” is just an image of letters. Your cursor can’t select it because, to the computer, there are no words there, only pixels. Try to extract text and you’ll get nothing, because there’s nothing to extract yet.

Quick test: open the PDF and try to select a line of text. Selects? Text-based. Won’t select? It’s a scan, and you’ll need a different first step.

Extracting text in your browser

For a text-based PDF, files.co has an extract text tool that runs on your device:

  1. Open the extract text tool and drop your PDF in.
  2. Let it read the file. Your browser pulls the text content out of the document, in reading order.
  3. Grab the result. Copy the text straight out, or download it as a plain text file you can use anywhere.

No page-by-page selecting, no scrambled columns to untangle by hand. The whole document’s text comes out in one move.

If it’s a scan, OCR first

If your PDF is a scan and the extract tool comes back empty, you’re not stuck, you just need an extra step. OCR (optical character recognition) looks at the image of the page and works out what letters it’s seeing, turning the picture of text into actual text.

Run the file through the OCR tool first. It reads the images and produces a PDF with a real, selectable text layer. Then send that result to the extract text tool and the words come out like any text-based file. Two steps, both in the browser, nothing uploaded.

Why it doesn’t leave your machine

Text extraction is one of those tasks where the privacy angle is easy to overlook, but the words in a PDF are often the whole point of keeping it private. A contract’s clauses. The figures in a financial statement. The body of a confidential report. Uploading the file to pull its text out hands every one of those words to a server you don’t control.

files.co reads the text locally. The tool runs in your browser with JavaScript: the PDF is loaded into memory on your device, the text is read out there, and the result is yours to copy or download. Nothing is sent anywhere. Want proof? Open DevTools (F12), watch the Network tab while you extract, and you’ll see no upload. Or turn off your connection and do it offline, it works the same.

What you might do next

Once you’ve got the text, the original PDF is often still in play. A few neighboring jobs, all browser-based on files.co:

For the everyday case it’s one step. Open the extract text tool, drop your PDF, copy the words out. Clean text, no manual selecting, and the document never left your computer.

Explore by category