Extract selectable PDF text into a plain-text file
A researcher needs words from a report for quotation review.
Open Extract text →Follow these steps
- Choose the supported source file or files from your device.
- Use a PDF containing selectable text; run OCR first for a scan.
- Export the TXT file and check reading order and missing sections.
The document and the intended outcome
A researcher needs words from a report for quotation review. Extract text, open the TXT file and compare several paragraphs with their original pages.
Make this example work for your file
For extracted notes, compare the first paragraph, a multi-column page and the final paragraph with the original. Reorder text manually when the source’s drawing order differs from its human reading order.
Choose only the settings this case needs
Choose a PDF containing selectable text and inspect page boundaries in the result. Keep the PDF alongside the TXT so quotations can be checked against the source.
When the example needs a different approach
Plain-text extraction does not reconstruct the original layout. Reorder the relevant passage manually after comparing it with the visible page.
A checked example you can try
Extracted content contains the source Alpha text.
Settings used in this engine check
- quality
- compact
- angle
- 90
- margin
- 20
- start
- 1
- columns
- 2
Check the result
Columns and positioned text may come out in an unexpected order. Images and layout are not preserved.
Keep the original. When this task produces a file, reopen the download in the application where you will use it and inspect the affected pages. An illustration explains the operation; it is not a recorded processing test.
Practical tips
- Choose a PDF containing selectable text and inspect page boundaries in the result. Keep the PDF alongside the TXT so quotations can be checked against the source.
- The page is an image. Run OCR first and review the recognized words before relying on extracted text.
- Plain-text extraction does not reconstruct the original layout. Reorder the relevant passage manually after comparing it with the visible page.
Questions
The TXT is empty for a scan?
The page is an image. Run OCR first and review the recognized words before relying on extracted text.
Columns appear mixed together?
Plain-text extraction does not reconstruct the original layout. Reorder the relevant passage manually after comparing it with the visible page.
What is this tool useful for?
Copy document text into a plain-text workflow.
What should I check in the result?
Columns and positioned text may come out in an unexpected order. Images and layout are not preserved.