How to extract selectable text from a PDF
Save selectable text as a TXT file. Learn the available controls, a practical example and what to check.
Open Extract text →Follow these steps
- Choose the supported source file or files from your device.
- Use a PDF containing selectable text; run OCR first for a scan.
- Export the TXT file and check reading order and missing sections.
What this tool actually does
Extraction reads available text objects. Multi-column layouts, unusual character mappings and untagged reading order can produce a sequence different from what the page looks like.
Choose the controls for your document
Choose a PDF containing selectable text and inspect page boundaries in the result. Keep the PDF alongside the TXT so quotations can be checked against the source.
Before your first export
For extracted notes, compare the first paragraph, a multi-column page and the final paragraph with the original. Reorder text manually when the source’s drawing order differs from its human reading order.
The TXT is empty for a scan
The page is an image. Run OCR first and review the recognized words before relying on extracted text.
A checked example you can try
Extracted content contains the source Alpha text.
Settings used in this engine check
- quality
- compact
- angle
- 90
- margin
- 20
- start
- 1
- columns
- 2
Check the result
Columns and positioned text may come out in an unexpected order. Images and layout are not preserved.
Keep the original. When this task produces a file, reopen the download in the application where you will use it and inspect the affected pages. An illustration explains the operation; it is not a recorded processing test.
Practical tips
- Choose a PDF containing selectable text and inspect page boundaries in the result. Keep the PDF alongside the TXT so quotations can be checked against the source.
- The page is an image. Run OCR first and review the recognized words before relying on extracted text.
- Plain-text extraction does not reconstruct the original layout. Reorder the relevant passage manually after comparing it with the visible page.
Questions
The TXT is empty for a scan?
The page is an image. Run OCR first and review the recognized words before relying on extracted text.
Columns appear mixed together?
Plain-text extraction does not reconstruct the original layout. Reorder the relevant passage manually after comparing it with the visible page.
What is this tool useful for?
Copy document text into a plain-text workflow.
What should I check in the result?
Columns and positioned text may come out in an unexpected order. Images and layout are not preserved.