Extract text: common problems and how to fix them
Resolve the txt is empty for a scan and columns appear mixed together. Understand the limits before retrying.
Open Extract text →Follow these steps
- Choose the supported source file or files from your device.
- Use a PDF containing selectable text; run OCR first for a scan.
- Export the TXT file and check reading order and missing sections.
The TXT is empty for a scan
The page is an image. Run OCR first and review the recognized words before relying on extracted text.
Columns appear mixed together
Plain-text extraction does not reconstruct the original layout. Reorder the relevant passage manually after comparing it with the visible page.
Check whether this is the right operation
Extraction reads available text objects. Multi-column layouts, unusual character mappings and untagged reading order can produce a sequence different from what the page looks like.
A more useful retry
For extracted notes, compare the first paragraph, a multi-column page and the final paragraph with the original. Reorder text manually when the source’s drawing order differs from its human reading order.
A checked example you can try
Extracted content contains the source Alpha text.
Settings used in this engine check
- quality
- compact
- angle
- 90
- margin
- 20
- start
- 1
- columns
- 2
Check the result
Columns and positioned text may come out in an unexpected order. Images and layout are not preserved.
Keep the original. When this task produces a file, reopen the download in the application where you will use it and inspect the affected pages. An illustration explains the operation; it is not a recorded processing test.
Practical tips
- Choose a PDF containing selectable text and inspect page boundaries in the result. Keep the PDF alongside the TXT so quotations can be checked against the source.
- The page is an image. Run OCR first and review the recognized words before relying on extracted text.
- Plain-text extraction does not reconstruct the original layout. Reorder the relevant passage manually after comparing it with the visible page.
Questions
The TXT is empty for a scan?
The page is an image. Run OCR first and review the recognized words before relying on extracted text.
Columns appear mixed together?
Plain-text extraction does not reconstruct the original layout. Reorder the relevant passage manually after comparing it with the visible page.
What is this tool useful for?
Copy document text into a plain-text workflow.
What should I check in the result?
Columns and positioned text may come out in an unexpected order. Images and layout are not preserved.