calibre PDF Conversion Problems: Why the Output Looks Wrong

  • Identify whether PDF extraction, OCR, layout, or conversion settings caused the damage.
  • Fix broken paragraphs, repeated headers, missing chapters, and mixed column order.
  • Learn when manual editing or keeping the original PDF is the smarter choice.

Common calibre PDF conversion problems include broken paragraphs, repeated headers inside the text, page numbers appearing between sentences, missing chapter breaks, scrambled columns, and pages read in the wrong order. These symptoms usually do not mean that calibre is broken. They occur because PDF is a final-layout format that records where content appears on a page, while EPUB, MOBI, and AZW3 are primarily reflowable formats that need meaningful paragraphs, headings, and a reliable reading sequence.

The most likely cause falls into one of four categories: the source PDF has weak or missing text information, the page layout is too complex to reconstruct reliably, a conversion option is unsuitable for that document, or the converted book needs manual structural cleanup. The steps below help you identify which category applies without reinstalling calibre, deleting your library, or changing unrelated settings.

Fixed PDF pages being transformed into a reflowable ebook during a controlled conversion test.

1. Confirm the Symptom With a Small Safe Test

Before adjusting conversion settings, make a controlled test. Work with a copy of the PDF if it is stored in a synchronized folder or shared location. Add that copy to calibre as a separate test book, or temporarily add it to a small test library. This keeps your existing formats and metadata safe while you experiment.

1.1 Determine Whether the PDF Contains Usable Text

Open the original PDF in a PDF reader and try to select one ordinary paragraph. Copy it into a plain-text editor such as Notepad, TextEdit in plain-text mode, or a Linux text editor.

  • If the paragraph pastes in the correct order, the PDF contains a usable text layer.
  • If each line pastes separately, conversion may produce hard line breaks.
  • If columns become mixed, the PDF lacks a clear reading sequence.
  • If the result is empty or nonsensical, the pages may be scans or the text encoding may be defective.
  • If letters are replaced with unusual symbols, the embedded character mapping may be unreliable.

This simple copy test often predicts the conversion result. Success at this stage means that a normal paragraph can be selected and pasted in the intended reading order. If it cannot, stop changing EPUB output settings. The main problem is already present in the PDF source.

1.2 Convert a Short Representative Section

If possible, create a short PDF containing several representative pages, such as a chapter opening, a normal text page, and a page with a header or illustration. Alternatively, run a normal conversion but inspect those pages first.

Convert the test to EPUB before trying several output formats. EPUB is convenient for diagnosis because calibre can open it in both the viewer and Edit book. If the EPUB has the same defects as an AZW3 or MOBI conversion, the problem is probably in PDF extraction or structure detection rather than the device profile.

Success means ordinary paragraphs reflow when the viewer window changes width, sentences remain in order, and chapter headings appear near their original locations. Once those conditions are met, stop changing PDF input settings and test the required final format.

2. Check the Conversion Options Directly Related to the Problem

Open the book's individual conversion dialog rather than changing global preferences immediately. Confirm that PDF is selected as the input format and choose EPUB, AZW3, or the format actually required by your reading device as the output.

2.1 Fix Broken Paragraphs and Bad Line Breaks

A PDF normally stores text as lines positioned on fixed pages. calibre must infer whether the end of each line is a real paragraph boundary or merely the right edge of the printed page. The PDF Input section includes a line unwrapping factor that influences this decision.

If nearly every printed line becomes a separate line in the converted book, lower the line unwrapping factor in small steps. This encourages calibre to join more lines. If separate paragraphs are being merged into long blocks, increase it slightly so fewer lines are joined.

Change only this setting, convert again, and compare the same two or three pages. Do not combine the test with multiple heuristic options because you will not know which change helped. Success means lines within a paragraph join naturally while genuine paragraph endings remain separate. Stop when the representative pages are readable, even if a few exceptional paragraphs still need manual repair.

2.2 Remove Headers, Footers, and Page Numbers

Running titles, author names, page numbers, and footnotes are often stored as ordinary page text. They can therefore appear in the middle of paragraphs after conversion. Repeated headers and footers can also interfere with paragraph unwrapping.

Check the PDF Input options available for header and footer detection. For document-specific repeated text, use the Search and replace section of the conversion dialog. calibre provides a testing interface that can highlight matches before removal. Build the narrowest possible pattern around the repeated header, footer, or page-number line.

Test carefully before applying a replacement. A broad rule that removes every number, for example, could erase dates, list numbers, or references in the main text. Success means the repeated page furniture disappears while legitimate body content remains. Once the test highlights only the unwanted material, stop expanding the rule.

2.3 Handle Missing Chapters and a Broken Table of Contents

A printed chapter heading may only be larger text at a particular page coordinate. It does not necessarily contain the semantic heading information expected in a reflowable e-book. As a result, calibre may fail to detect chapters, create breaks in the wrong places, or generate an incomplete table of contents.

Inspect Structure detection and Table of Contents settings, but do not assume one detection expression will work for every PDF. If chapter titles survive as recognizable text, structure detection may be able to identify a consistent pattern. If titles have been split, converted into images, or mixed with page headers, manual cleanup after conversion will usually be more reliable.

Success means chapter starts are consistently identifiable and navigation entries open the correct locations. If automatic detection still misses irregular headings after one or two controlled tests, stop tuning detection rules and plan to repair the EPUB or AZW3 in Edit book.

2.4 Diagnose Columns and Wrong Reading Order

Multi-column PDFs are among the hardest sources to convert. A page may visually show the left column followed by the right column, but its internal text objects may be stored line by line across both columns or in the order they were added by the publishing software.

If copied text already alternates between columns, ordinary conversion options may not restore the intended sequence. Cropping or splitting columns in a suitable PDF preparation tool before conversion may help when you have permission to modify the document. For a short document, manually rebuilding the reading order may be faster. For a complex textbook, journal, brochure, or illustrated manual, keeping the PDF is often the better choice.

Success means the full left column is read before the full right column on each test page. If different pages use different column arrangements, stop searching for one universal conversion setting. The source layout is not consistently reflowable.

2.5 Verify Metadata Without Confusing It With Text Structure

Correcting title, author, language, tags, or cover metadata can improve library organization, but it does not repair paragraph extraction or page reading order. Likewise, downloading metadata from another source will not reconstruct missing chapters inside the PDF.

Confirm that you are converting the intended PDF format attached to the intended calibre book record. A book record can contain multiple formats, and an older output may remain attached after a test. Remove only an obsolete generated format when you are certain it is not needed, then run the conversion again from the original PDF.

Success means the newly generated format has the expected modification time and opens with the correct title and content. Stop editing metadata once the right source and output files are confirmed.

3. Check Source Quality, Reader Limitations, and Operating System Interference

3.1 Distinguish Scanned PDFs From Text PDFs

A scanned PDF contains page images. It may look perfectly readable to a person while containing no machine-readable text. calibre is an e-book manager and converter, not a full OCR workflow for rebuilding scanned books. If text cannot be selected, the document generally needs optical character recognition before meaningful reflowable conversion is possible.

Some scanned PDFs already contain an OCR text layer behind the images. That layer can still contain misspelled words, incorrect punctuation, merged columns, and misplaced headings. The visible page image may look correct even when the hidden OCR text is poor. The copy-and-paste test reveals what calibre is likely to receive.

If you run OCR with software you are authorized to use, review its output before importing it into calibre. A DOCX, HTML, or clean EPUB produced from corrected OCR text is usually a better conversion source than the original scanned PDF. Success means copied text is accurate, ordered correctly, and separated into sensible paragraphs before calibre conversion begins.

3.2 Try a Better Source Format

PDF should be treated as a last-choice source for reflowable conversion. If the same book is legally available to you as EPUB, DOCX, HTML, ODT, RTF, or another structured text format, use that version instead. A source containing real paragraphs and headings gives calibre much better information than fixed page coordinates.

If you created the PDF yourself, return to the original word-processing or page-layout document. Export clean DOCX, HTML, or EPUB when the application supports it, then add that file to calibre. Do not repeatedly convert PDF to EPUB, EPUB to MOBI, and MOBI to AZW3. Each conversion can add more structural loss.

Success means the replacement source retains paragraph boundaries, chapter headings, emphasis, and navigation with little or no cleanup. At that point, stop troubleshooting the PDF.

3.3 Rule Out File Access and Cloud Sync Problems

Permissions, antivirus software, and cloud synchronization are not common causes of scrambled paragraphs, but they can explain failed jobs, incomplete output, or files that appear not to update. Copy the source PDF to a normal local folder that your user account can write to. Avoid converting directly from a removable drive, network share, or actively synchronized folder during the test.

On Windows, macOS, or Linux, confirm that the PDF opens normally and is not zero bytes, partially downloaded, or locked by another application. If security software reports an action, review that report rather than disabling protection broadly. Grant access only to the specific trusted calibre executable or working folder when appropriate.

Success means the conversion job completes, the generated format has a current timestamp, and reopening it shows the latest test. Once that happens, stop changing permissions or security settings because they will not improve the internal reading order.

3.4 Separate Conversion Defects From Device or Viewer Defects

Open the converted EPUB or AZW3 in calibre's viewer before sending it to a device. Resize the viewer and navigate through the same problem pages. If the text is already wrong, changing USB mode, email delivery, Content server options, or device settings will not fix it.

If the file looks correct in calibre but wrong on one device, test a format natively supported by that device and select an appropriate output profile. Older readers may have limited CSS, font, table, or image support. MOBI can also impose more formatting limitations than newer formats, so use EPUB or AZW3 when the destination supports it.

Success means the book is correct in calibre and in at least one compatible reader. If only one device fails, stop modifying PDF extraction settings and focus on that device's supported formats and rendering limitations.

Document stages revealing where paragraph structure first breaks during conversion.

4. Use Conversion Debug Output and Job Details

4.1 Read the Conversion Job Log First

After conversion, open calibre's Jobs area and view the details for the completed or failed conversion. The log records the selected input and output, processing stages, warnings, and errors. Save the log if you need to compare tests or request help.

A completed job can still produce poor content, so distinguish technical completion from conversion quality. Errors about opening the source, writing output, loading a plugin, or accessing a path point to an operational problem. A clean completion followed by scrambled text usually points to source structure rather than a damaged calibre installation.

4.2 Generate Conversion Debug Output

The conversion dialog can write intermediate files to a debug folder. calibre's conversion pipeline first extracts the input into HTML-like content, then parses and transforms that content before producing the final e-book. Debug output lets you see where the defect first appears.

  • input: Shows the material produced by the PDF input stage.
  • parsed: Shows content after parsing and preprocessing.
  • structure: Shows the result after structural detection.
  • processed: Shows content shortly before it reaches the output plugin.

Open the relevant intermediate HTML files in a browser or text editor. If paragraphs, columns, or chapter titles are already wrong in the input folder, the PDF extraction stage could not recover the intended structure. If input looks reasonable but later stages become wrong, review structure detection, search and replace, heuristic processing, and output options.

Success means you can identify the first stage where the content changes incorrectly. Stop changing unrelated settings after locating that stage.

4.3 Use Edit Book for Targeted Cleanup

Convert the PDF to EPUB or AZW3, select the resulting book, and open Edit book. This is appropriate when the automatic conversion is mostly correct but leaves a manageable number of defects.

Typical repairs include joining split paragraphs, separating merged paragraphs, deleting repeated headers, marking chapter titles as headings, splitting files at chapter boundaries, correcting the table of contents, and adjusting CSS. Use search and replace cautiously, preview changes, and run the editor's checks before saving.

If hundreds of pages require individual reconstruction, manual cleanup may cost more time than obtaining a better source or retaining the PDF. Success means the edited book reflows cleanly, navigation works, and only minor visual differences remain.

4.4 Reserve calibre-debug for Actual Application Problems

The calibre-debug command can start calibre with diagnostic output, and calibre also provides device-detection debugging. These tools are useful when the application crashes, a plugin fails, the GUI behaves unexpectedly, or a connected device is not detected. They are not usually necessary for ordinary PDF paragraph problems.

If calibre itself is not working, start it in debug mode, reproduce one problem, close the application, and inspect or save the resulting output. Temporarily disabling a recently added third-party plugin can also be a useful comparison test. Do not remove all plugins or reset every preference unless the evidence points there.

Success means the problem can be reproduced with a concise log or disappears when one specific plugin is disabled. If conversions complete normally and only PDF layout is wrong, stop application-level debugging.

5. Run a Clean Temporary Test Before Reinstalling

Reinstalling calibre rarely repairs a PDF whose text order is defective. Deleting a library is even less appropriate because the library contains your books, formats, covers, and metadata. Instead, isolate the conversion in a temporary test.

  1. Create a new empty test library from calibre's library menu.
  2. Copy one non-sensitive PDF to a local folder.
  3. Add only that PDF to the test library.
  4. Use default conversion settings and produce an EPUB.
  5. Inspect the EPUB in calibre's viewer.
  6. Repeat once with only the setting directly related to the symptom.
  7. If needed, temporarily disable a relevant third-party plugin and repeat.

If the clean conversion has the same broken columns or text order, the source PDF is the likely cause. If it works, compare the original book's saved conversion settings, attached formats, and any plugin involvement. Keep the original library untouched until you know the cause.

Success means you can state whether the defect follows the source file, a specific setting, a plugin, or the original library record. Once isolated, stop broad troubleshooting and apply the narrow fix.

6. Quick Fix Checklist

  • Copy one paragraph from the PDF to verify that usable text exists.
  • Check whether the file is a scan, a text PDF, or a scan with OCR.
  • Convert one representative sample to EPUB before testing multiple formats.
  • Adjust the PDF line unwrapping factor in small steps for broken paragraphs.
  • Remove repeated headers, footers, and page numbers with narrowly tested rules.
  • Inspect copied text for mixed columns before blaming the output profile.
  • Use conversion debug output to find the first damaged pipeline stage.
  • Repair limited structural defects in Edit book.
  • Use a better source such as EPUB, DOCX, or HTML when available.
  • Test from a local writable folder if jobs fail or output does not update.
  • Check the converted file in calibre before sending it to a device.
  • Keep the PDF when fixed layout is essential or cleanup is excessive.

7. Frequently Asked Questions

7.1 Why Does calibre Put a Line Break After Every PDF Line?

The PDF probably stores each printed line as a separate positioned text object. calibre must guess which line endings belong inside a paragraph. Adjust the PDF input line unwrapping factor gradually and compare the same sample pages. If the hidden text is badly fragmented, a better source or manual cleanup may be necessary.

7.2 Why Are Page Numbers and Headers Inside Sentences?

Headers, footers, and page numbers are often ordinary text objects in a PDF. After reflow, they lose their fixed position and can land inside body paragraphs. Remove them during conversion with PDF input controls or carefully tested search-and-replace patterns. Removing them can also improve paragraph joining.

7.3 Can calibre Convert a Scanned PDF to EPUB?

Not reliably when the PDF contains only page images. The document first needs a usable OCR text layer, and that OCR output must be checked for recognition and reading-order errors. Even after OCR, complex layouts may require substantial editing.

7.4 Why Are Two-Column Pages Read in the Wrong Order?

The PDF's internal text sequence may not match its visual arrangement. calibre can only work with the text objects and coordinates it receives. If a copy-and-paste test also mixes the columns, prepare a simpler source, split the columns before conversion when permitted, rebuild the document manually, or keep the PDF.

7.5 Should I Reinstall calibre When Every PDF Conversion Looks Wrong?

Usually not. First test a simple text PDF and inspect conversion debug output. If normal conversions complete and the defects vary by document, the issue is source quality or layout complexity. Reinstallation is reasonable only when calibre files are damaged, the program cannot start, or logs indicate an application-level failure.

7.6 When Is Keeping the PDF Better Than Converting?

Keep the PDF when exact pagination, columns, diagrams, forms, equations, footnotes, print fidelity, or page references matter more than font resizing and text reflow. It is also the better choice when the document is image-based or when correcting the reading order would require rebuilding most pages. A readable original PDF is preferable to a reflowable e-book with missing or rearranged content.


Citations

  1. Official guidance on calibre conversion, PDF input limitations, line unwrapping, structure detection, and debug output. (calibre E-book Conversion Manual)
  2. Official answers covering preferred conversion sources and the known difficulties of converting PDF files. (calibre Frequently Asked Questions)
  3. Official command reference for PDF input controls, header and footer handling, and line unwrapping. (calibre ebook-convert Documentation)
  4. Official reference for running calibre diagnostics and debugging device detection. (calibre-debug Documentation)
Cindy, ContentBASE creator assistant

MEET CINDY

Your ContentBASE creator assistant

Cindy helps creators find Canva templates, content ideas, and simple ways to make better social media posts faster.

Want ready-to-use templates? Claim the free Canva bundles or browse the full bundle store.