PDF Conversion

PDF Conversion and OCR: Choose the Right Workflow

Cover illustration for: PDF Conversion and OCR: Choose the Right Workflow
Quick answer

Choose a conversion workflow based on what the PDF contains and what you need afterwards. Selectable text, scanned pages and complex tables need different handling. OCR recognises characters in images; it does not automatically restore formulas or perfect layout. Always compare the exported file with the source. This guide is the hub for our conversion tools, including PDF to Word, PDF to Excel, PDF to PowerPoint, PDF to JPG, JPG to PDF, Word to PDF, Excel to PDF and OCR PDF.

"Convert" is one word for many different jobs. A converter that works well for one file can fail on another because the file contains something else: real text, a picture of text, or a table drawn as loose words. The fastest way to a good result is to identify what you have, decide what you need, and pick the route that fits. Below, we describe each step and note what the iLikePDF pages currently do.

Identify text PDFs versus scans

Open the PDF, search for a word you can see and try to select a line. If the word is found and text highlights, the PDF has real text. If not, it is a scan: a picture of a page. Everything that reads text out of a PDF, including PDF to Word and PDF to Excel, needs real text. Scans need text recognition first. The article scanned PDF versus text PDF explains the difference in detail.

Choose the required output format

You needRouteKeepsLoses
Editable paragraphsPDF to WordTextTables, columns, images, layout
Rows in a spreadsheetPDF to ExcelText rows, page numbersBorders, merged cells, number formats
Slides to presentPDF to PowerPointExact page lookEditable text
Images of pagesPDF to JPGPage lookSearchable text
A PDF from photosJPG to PDFImagesSearchable text

Decide the output before starting, then check the tool's stated limits. The detailed articles cover each: PDF to Word, PDF to Excel, PDF to PowerPoint, PDF to JPG and JPG to PDF.

Try it: OCR PDFText recognition for scanned PDFs.
Open OCR PDF β†’

Convert Office files to PDF

For Word and Excel documents, the best route is the Save as PDF or Export command in the program that created them. It uses the same fonts, styles and print settings as your screen. Set print areas and scaling in spreadsheets before exporting. The iLikePDF Word to PDF and Excel to PDF pages currently read plain text and CSV-style text only, so they suit simple text, not formatted workbooks; see Word to PDF and Excel to PDF for the full workflow and what to check.

Extract tables and check data types

A table in a PDF is only positioned words. Extraction tools infer rows and columns, so cells that wrap onto two lines can become two rows. After extracting, add up a column and compare it with the total in the PDF. Import CSV files as text so leading zeros and long numbers survive. The article converting PDF tables to Excel gives a step-by-step method.

Run OCR on clear scans

OCR gives best results on straight, high-contrast scans of about 300 dpi with the correct language selected. Rotate sideways pages first with Rotate PDF. Then verify the recognised text: search known words, and compare every number with the picture. The iLikePDF OCR PDF page currently re-saves the file without adding recognised text, so use a scanning app or desktop OCR program for this step and test the output with search.

Create images or presentation slides

PDF to JPG draws each page at twice its standard size and saves JPEG images, delivering a ZIP for multi-page files. PDF to PowerPoint places a picture of each page on a slide, at 16:9 or 4:3. Neither gives editable text. For editable content, extract the text with PDF to Word and rebuild the slide. If you only need to show pages, a PDF viewer's full-screen mode is often simpler.

Review conversion errors

SymptomLikely causeFix
Empty outputThe PDF is a scanRun OCR first
Numbers changedSpreadsheet reinterpreted textImport as text
Layout lostText extraction does not rebuild layoutRecreate tables and columns
Sideways pagesOrientation flag ignoredRotate before converting
Huge fileHigh-resolution page imagesSplit the PDF, or reduce pages
Characters wrongRecognition or encoding errorsProofread against the source

Keep the original PDF next to every converted file. When something looks off, it is the reference you compare with.

A worked example

A finance assistant receives a supplier's price list as a PDF and needs it in a spreadsheet. She searches for a word in the PDF and finds it, so the file has real text. She chooses PDF to Excel with CSV output and imports it as text, so item codes keep their leading zeros. She adds up the price column and compares it with the total in the PDF. They match. A second document from the same supplier is a scanned delivery note; searching finds nothing, so she knows it needs text recognition first, and she uses a scanning app with OCR, checks the numbers by eye and only then extracts the table.

The same input type, a PDF, led to two different workflows because the contents were different. Identifying scan or text was the decision that mattered.

Choosing quickly

QuestionAnswerNext step
Can I select and search the text?YesUse extraction tools directly
Can I select and search the text?NoRun OCR first
Do I need to edit the words?YesPDF to Word, then fix the layout
Do I need numbers in cells?YesPDF to Excel, then import as text
Do I just need to show the pages?YesImages or slides, or a PDF viewer
Is the source an Office file I still have?YesExport from the original program

Keep the original and verify

Conversion is a way of making a copy in another shape. The original PDF remains the reference for layout, numbers and wording. Every conversion needs a check: compare a few pages, add up a column, read the important numbers aloud and look for missing sections. If the document has legal or financial weight, do this check properly rather than glancing at it. And remember the browser tools on this site process files in your browser without an upload request from their code, but the converted files are saved on your device, so treat them with the same care as the originals.

Frequently asked questions

Related tools

Keep reading

Try it: OCR PDFText recognition for scanned PDFs.
Open OCR PDF β†’