By Jetformat
PDF text extraction reads text objects that already exist in a PDF. Optical character recognition, or OCR, recognizes characters in page images and can create searchable text from scans. Jetformat performs the first operation; it does not perform OCR.
Compare what the file contains
| Source | What a text extractor can use | Review needed |
|---|---|---|
| Digital PDF with text | Existing text objects | Reading order and tables |
| Scanned page image | No usable text layer | OCR in another tool |
| Mixed PDF | Text on some pages only | Page-by-page completeness |
Jetformat does not perform OCR. Converting a scan with pdf to-word does not create recognized text that was absent from the PDF. Use an OCR workflow separately if the source requires it, then review recognition quality before downstream extraction.
Start with a free read
jetformat text source.pdf --format markdown -o source.md
jetformat pdf info source.pdf --json
Open the extracted text and compare it with several source pages. Check a page with dense columns, a page with a table, and any page that looks like a scanned insert. Nonempty output from the first page is not evidence that the entire document has usable text.
Editable output still needs review
For a digital PDF, jetformat pdf to-word source.pdf -o source.docx rebuilds an editable document. The text layer can preserve words while losing the logical ordering implied by the visual layout. Tables are especially sensitive to spacing and merged cells.
Before treating the result as a source of record, compare headings, key amounts, names, and paragraph order. Keep the original PDF alongside the editable derivative so discrepancies can be traced.
Use the PDF-to-Word task guide for conversion and the table extraction guide when you only need structured rows. For image inspection, render the original pages.
Inspect a PDF step by step
- Run
pdf infoto read the source page count and properties. - Extract Markdown or plain text to a new file.
- Compare the extracted content with the original pages, including tables and scanned inserts.
- If words exist only in images, use a separate OCR tool and review recognition errors before continuing.
- Create a DOCX only after deciding that the available text is suitable for the requested editable output.
The presence of some text does not settle the whole document. A scanned appendix can coexist with a digital cover sheet, and an old OCR layer can contain inaccurate text beneath a readable image.
What to check in the extracted result
Compare names, dates, totals, and paragraph order. A two-column page may require more careful reading-order review than a simple letter. Tables can lose the visual relationships suggested by borders and alignment, so verify that an amount remains associated with the correct row label.
Do not infer exact source layout from Markdown. It is a useful reading representation, not a visual reproduction. Use page rendering when the question concerns overlap, clipped text, or the position of a signature line.
A mixed-document example
Imagine a 10-page PDF containing 8 digital pages and 2 scanned pages. Text from the 8 digital pages is not proof that the other 2 pages have been captured. Keep a page-level checklist so the missing content does not silently disappear from a summary or editable derivative.
Reading text and running pdf info use 0 credits. Rendering 3 selected pages uses 3 credits; converting the entire 10-page PDF with pdf to-word uses 10 credits. These are example operation counts. Paying for conversion does not add OCR support or repair absent source text.
When a separate OCR workflow is appropriate
Use OCR when the required words are stored only as images. Adobe’s scan-to-PDF documentation describes OCR options in Acrobat. Review the resulting text against the scan, especially short names, numeric amounts, and similar-looking characters. OCR output should not be assumed correct merely because it becomes searchable.
Retain the original scan with the reviewed derivative. If later extraction disagrees with the visible source, that original provides the evidence needed to resolve the discrepancy. For tasks that require only a table, compare PDF table extraction with full PDF-to-DOCX conversion before creating an unnecessary intermediate document.