“Extract text,” I told the tool, pointing at a 10-page flat scan. It returned ten blank pages with polite per-page flags. Correct behavior — and the moment I finally internalized the distinction that saves hours: selectable text extracts; scanned images need OCR. They look identical on screen and behave oppositely under extraction. This guide teaches telling them apart in seconds, extracting with page ranges, and knowing exactly when to reach for OCR software instead.
Part of the PDF workflow guide. Extract in PDF to text; split large sets first in PDF split. Scans to PDF background in scans guide.
Selectable layer vs scanned image (the 5-second test)
| Test | Digital (extractable) | Scanned (needs OCR) |
|---|---|---|
| Drag-select text | Selects cleanly | Selects nothing / whole page |
| Zoom to 400% | Edges stay sharp | Pixels blur |
| Search (Ctrl+F) | Finds words | Finds nothing |
| File size per page | Kilobytes | Hundreds of KB+ |
Run the drag-select test first — five seconds that prevent ten minutes of confused re-extraction. Hybrid documents mix both types unpredictably; review per-page flags instead of assuming uniformity.
Extract with ranges: 1-3,5 chapters to .txt
Same range grammar as splitting: type 1-3,5 to pull an introduction, or leave blank for all pages (cap 200). Output arrives as a plain-text preview with --- Page N --- separators — empty pages flagged (no extractable text) rather than silently skipped, so you know exactly which sheets need OCR. Copy to clipboard for Word/Notes, or download chapter.txt. Academic hygiene: retain original pagination, author and range when quoting (“pp. 10–20 of the manual”), archive source PDFs beside notes.
When you actually need OCR (and preprocessing that helps)
- Image-only scans: dedicated OCR software (desktop or service) — extraction tools correctly return empty here by design.
- Before OCR: unlock secured files, split 300-page bundles into smaller chunks, straighten skewed scans — recognition accuracy tracks input quality linearly.
- After OCR: proofread proper nouns and numbers (OCR confuses 0/O, 1/l, 5/S); searchable PDFs from good OCR then extract normally.
- Handwriting: specialized engines only; general OCR fails on cursive — budget human transcription for critical passages.
Preprocessing checklist for the 120-page manual: split 1-3,5 plus 10-20 test batches through extraction first (free, instant) to map which pages are digital vs scanned — then OCR only the scanned subset instead of paying for all 120.
General guidance only. Cite extracted passages with original pagination — copy-paste without attribution is still plagiarism with extra steps.