calibre PDF to EPUB Line Breaks Wrong: How to Fix It

When a PDF converted to EPUB in calibre breaks every sentence at the original page margin, splits normal paragraphs into short blocks, or joins unrelated text, the problem is usually paragraph reconstruction rather than EPUB display. PDF files store text according to fixed page coordinates, while EPUB files require reflowable paragraphs. calibre must infer which PDF lines belong together, and that inference can be disrupted by unusual line lengths, headers, footers, columns, OCR errors, or a poor text layer. The safest approach is to test a small sample, adjust only the relevant PDF input options, inspect the conversion output, and stop as soon as paragraphs reflow correctly.

Fixed PDF lines transforming into reflowable paragraphs on an ebook reader.

1. Confirm the Symptom With a Small Safe Test

Before changing settings, confirm that you are dealing with PDF line wrapping rather than a viewer, device, or stylesheet problem. Choose two or three pages containing ordinary body text. Avoid beginning with a table, index, footnotes, poetry, or a multicolumn page because those layouts are intentionally difficult to reconstruct.

1.1 Check How Text Behaves in the PDF

Open the original PDF in a PDF reader and copy one normal paragraph into a plain-text editor. Examine the pasted result.

  • If each visual line becomes a separate line, the PDF exposes hard line boundaries that calibre must unwrap.
  • If words appear out of order, the PDF's internal reading order is unreliable.
  • If no text can be selected, the pages may be image-only scans.
  • If selected text differs from the visible text, an inaccurate OCR text layer may be present.
  • If headers, page numbers, or footers appear inside the pasted paragraph, they can interfere with paragraph detection.

This simple test matters because calibre converts the PDF's available text structure, not merely its visible appearance. A page can look perfect while containing fragmented, duplicated, or incorrectly ordered text.

1.2 Convert One Clean Excerpt

If possible, make a temporary PDF containing a few representative pages that you are permitted to copy. Add that test file to calibre, select Convert books, choose EPUB as the output, and initially use the normal conversion defaults.

Open the resulting EPUB in calibre's E-book viewer. If lines still end near the PDF's original right margin even when the viewer window is resized, paragraph reconstruction failed during conversion. If paragraphs reflow correctly in calibre but break on a device, the conversion may already be good and the remaining issue is device-specific rendering or a transferred copy that was not refreshed.

Success looks like this: resizing the viewer changes where lines wrap, while genuine paragraph boundaries remain intact. Once that happens, stop adjusting PDF unwrapping settings.

2. Adjust the PDF Input Options That Control Paragraph Reconstruction

For this symptom, the most relevant settings are under PDF Input in the conversion dialog. Metadata downloads, library columns, email delivery accounts, Content server settings, and device profiles normally do not decide whether PDF lines become paragraphs. Do not change unrelated options simply because the finished EPUB was delivered through a server or sent to a device.

2.1 Tune the Line Un-Wrapping Factor

calibre uses a line un-wrapping factor to estimate whether a line is long enough to be joined to the following line. The value is a decimal between 0 and 1. A lower value makes calibre include more lines in unwrapping, while a higher value makes it unwrap fewer lines.

If almost every source line becomes a separate EPUB paragraph or break, lower the factor in small steps. For example, begin with a modest adjustment rather than jumping to an extreme. Convert the same short sample after each change and compare the same two paragraphs.

If calibre starts joining real paragraph endings, headings, list items, or other separate blocks, the value has gone too low. Move it back upward. There is no universal ideal value because PDFs use different page sizes, margins, columns, type sizes, and line-length patterns.

Success looks like this: wrapped prose lines are joined, but blank-line paragraphs, headings, and list items remain separate. Stop when the representative body text is correct, even if a few specialized passages still need editing.

2.2 Compare PDF Engines When Available

The PDF input options may provide a choice of PDF engine. The calibre engine is designed to handle PDF conversion and includes header and footer controls. If the current engine produces fragmented or badly ordered text, run a controlled comparison with the alternative engine while leaving other settings unchanged.

Do not combine an engine change with a new unwrapping value, heuristic processing, and several search-and-replace rules in the same test. If the output improves, you need to know which change caused the improvement.

2.3 Use Heuristic Line Unwrapping Carefully

calibre's heuristic processing includes an Unwrap lines function that looks for hard line breaks using line length and punctuation clues. It can help when the intermediate text contains hard breaks that survive the PDF input stage.

Enable heuristic processing and line unwrapping only as a separate test. Heuristics rely on patterns rather than knowledge of the document's meaning. They can incorrectly join headings, dialogue, poetry, addresses, bibliographies, code, or short list items.

If enabling heuristics improves ordinary prose without damaging other structures, keep it. If it joins unrelated blocks or changes already-correct sections, disable it and concentrate on PDF Input settings or manual cleanup.

2.4 Remove Headers and Footers Before Judging Unwrapping

Running titles, author names, chapter labels, page numbers, and footer notices can interrupt the sequence of lines at every page boundary. calibre may then treat the last body line, footer, header, and first line of the next page as separate paragraphs or combine them incorrectly.

First try the automatic header and footer handling available with the calibre PDF engine. If repeated text remains, use the PDF header and footer controls or the conversion dialog's search-and-replace panel. Build a rule around text that is genuinely repeated and distinctive. Test the rule before applying it to the entire document.

Be especially careful with changing chapter names, Roman numeral page numbers, and alternating left-page and right-page headers. A broad rule could remove legitimate text elsewhere in the book.

Success looks like this: recurring page furniture disappears, and paragraphs crossing page boundaries no longer contain page numbers or running titles. Retest the unwrapping factor after removing these interruptions because cleaner input may change the best result.

2.5 Reset Saved Conversion Settings When Results Seem Inconsistent

calibre can remember conversion choices for an individual book. If repeated attempts produce unexpected results, review the conversion dialog for previously saved PDF input, heuristic, search-and-replace, or structure-detection options. Reset the book-specific conversion settings or create a fresh temporary book entry for the test.

Replacing only the PDF file inside the same record may leave prior conversion choices associated with that record. Metadata such as title and author does not usually cause incorrect wrapping, but saved conversion options can.

3. Check Source Quality and Rule Out Unrelated Delivery Problems

3.1 Identify Image-Only and OCR-Based PDFs

An image-only PDF has no usable text for calibre to reconstruct. It must first be processed with legitimate OCR software capable of producing an accurate text layer or, preferably, an editable source document. OCR is also relevant when text can be selected but contains obvious substitutions, missing spaces, incorrect punctuation, or scrambled reading order.

Test several pages, not just the title page. Some PDFs contain searchable introductory pages followed by scanned images. Others contain two text layers, which can create duplicated sentences during conversion.

OCR errors affect line reconstruction because punctuation and spaces help calibre decide whether adjacent lines belong to the same paragraph. A period recognized as a comma, a missing final character, or random line-end spaces can change the result.

Success looks like this: copied text reads in the correct order, contains sensible spacing and punctuation, and does not duplicate visible content. If the text layer is badly corrupted, stop tuning calibre. Repair the OCR or obtain a better source.

3.2 Watch for Columns, Sidebars, Tables, and Footnotes

PDF is a fixed-layout format, so two columns may be represented as many positioned text fragments rather than two logical reading streams. A conversion might read across both columns, place sidebars inside paragraphs, or insert footnotes between body lines.

Line unwrapping settings cannot reliably repair incorrect reading order. If plain-text copying already interleaves the columns, look for a single-column source, split or crop the pages into logical regions using an appropriate document tool, or obtain the book in EPUB, DOCX, HTML, ODT, or another structured format.

Tables, mathematical layouts, plays, poetry, and heavily footnoted academic pages may require manual reconstruction. Do not force aggressive paragraph joining across the entire book to repair one difficult section.

3.3 Separate Conversion Errors From Viewer and Device Errors

Open the EPUB in calibre's viewer before sending it anywhere. Then resize the window and inspect the same paragraph in multiple font sizes. If it reflows correctly, the conversion is probably not the source of a later device problem.

If the file is correct in calibre but wrong after transfer, delete the old device copy and send the newly converted EPUB again. Confirm that you are opening the new file rather than a cached or previously converted edition. A Content server, email account, USB connection, firewall, cloud-synced library, or antivirus product can affect access or transfer, but these components do not normally reconstruct PDF paragraphs.

Operating system permissions become relevant only if calibre cannot read the PDF, write the EPUB, create debug files, or save edits. USB mode matters only when transferring the finished book. Do not troubleshoot firewalls or device detection when the EPUB is already broken inside calibre's local viewer.

3.4 Avoid Converting From a Cloud-Synchronized Working Copy

If the source PDF or temporary output is being actively synchronized, copy the PDF to a normal local folder for the test. This prevents incomplete downloads, placeholder files, file locks, and conflicting updates from complicating the diagnosis. Keep the main calibre library in a stable location and avoid manually modifying files inside its managed folders.

Cloud synchronization is unlikely to create systematic early line breaks, but it can produce stale or incomplete files that make comparisons unreliable.

Conversion stages revealing where a paragraph becomes fragmented.

4. Use Conversion Debug Output to Find the Failing Stage

When repeated tests do not explain why the calibre PDF to EPUB line breaks are wrong, inspect the conversion pipeline instead of guessing. calibre can save intermediate XHTML from multiple conversion stages.

4.1 Generate Debug Output

In the conversion dialog, select the debug option and choose an empty folder you can find easily. Run the conversion again. calibre creates stage folders commonly named input, parsed, structure, and processed.

  • input: The HTML produced by the PDF input plugin.
  • parsed: The preprocessed and normalized XHTML.
  • structure: Content after structure detection but before later presentation transforms.
  • processed: Content immediately before the EPUB output plugin receives it.

Open the relevant HTML files in a browser or text editor and search for a paragraph that converts badly. You do not need to be a developer to compare whether adjacent lines are already separate in the input stage or become damaged later.

4.2 Interpret What You Find

If the lines are already fragmented or out of order in the input folder, focus on the PDF engine, unwrapping factor, header and footer removal, OCR quality, and the original PDF's reading order.

If the input stage looks acceptable but a later folder introduces bad joins, disable heuristic processing, structure-detection rules, or custom search-and-replace expressions one at a time. If the processed XHTML is correct but the final EPUB looks wrong only in one reading application, test another viewer and inspect the EPUB's styling in calibre's editor.

Job details can also reveal which input and output formats were used and whether warnings occurred. Command-line users can run ebook-convert with debug-pipeline and verbose options. calibre-debug is useful for launching calibre components in debug mode, but it is rarely necessary for a straightforward PDF paragraph problem. Device-detection debugging is not relevant unless calibre also fails to recognize a connected reader.

4.3 Preserve a Useful Diagnostic Sample

Keep the original sample PDF, the settings used, the EPUB result, and the debug folder together. Record one sentence describing the outcome, such as “lowering the unwrapping factor joined body lines but merged headings.” This makes the next test intentional and prevents you from cycling through the same settings.

5. Perform Manual Cleanup After the Best Automatic Conversion

Automatic conversion does not have to be perfect before it becomes useful. Once the majority of paragraphs reflow correctly, manual editing is often safer than increasingly aggressive global settings.

5.1 Edit the EPUB Rather Than Repeatedly Converting It

Right-click the EPUB in calibre and choose Edit book. The editor can search across the book's HTML files, display a live preview, and create checkpoints before automated changes.

Typical cleanup tasks include removing repeated page-number paragraphs, joining paragraph tags that split a sentence, restoring blank lines, and correcting hyphenated words broken across source lines. Start with a small section and inspect every replacement pattern before applying it to all files.

Avoid a global rule that joins every adjacent paragraph. Real paragraph boundaries, dialogue, lists, headings, captions, and quotations may use the same markup as the unwanted breaks. Search-and-replace works best when the unwanted pattern has an additional clue, such as a lowercase continuation after a split, a repeated class, or a known header phrase.

5.2 Know When Another Source Format Is Required

Stop adjusting calibre when the debug input shows fundamentally unusable text: columns are interleaved, words are missing, OCR is inaccurate, every character is positioned separately, or paragraph order changes unpredictably. These are source-reconstruction failures, not ordinary EPUB styling problems.

Whenever possible, obtain the original EPUB, DOCX, HTML, ODT, or another structured source. Those formats preserve paragraphs and headings directly, while PDF primarily preserves page appearance. If you control the document, exporting from the original word-processing or publishing file will usually produce a better EPUB than converting its PDF derivative.

6. Run a Clean Temporary Test Before Reinstalling calibre

Reinstalling calibre rarely repairs one PDF's paragraph reconstruction because the behavior usually comes from the source file or conversion settings. A clean test provides better evidence.

  1. Copy a representative PDF or permitted excerpt to a local non-synchronized folder.
  2. Add it as a new temporary calibre record.
  3. Convert it to EPUB with default settings.
  4. Inspect it in calibre's viewer.
  5. Change only the PDF unwrapping factor and convert again.
  6. If necessary, test header and footer removal separately.
  7. Test heuristic line unwrapping as its own controlled comparison.
  8. Generate debug output if the cause remains unclear.

If the clean conversion works, the original book record probably has saved conversion settings or you were opening an older output. If the clean conversion fails in the same way, focus on the PDF's structure and text layer.

Do not delete your calibre library, move managed book folders manually, remove every plugin, or reset all preferences for this symptom. A third-party conversion-related plugin is worth disabling only if it participates in the conversion or post-processing workflow. Metadata, news, device, and interface plugins generally do not control PDF paragraph reconstruction.

7. Quick Fix Checklist

  • Copy a paragraph from the PDF and check its text order and line endings.
  • Test a few ordinary prose pages rather than the entire book.
  • Confirm the EPUB is broken in calibre's viewer before troubleshooting a device.
  • Lower the PDF line un-wrapping factor gradually when lines remain too short.
  • Raise the factor if legitimate paragraphs or headings are being joined.
  • Remove repeated headers, footers, and page numbers before final tuning.
  • Test heuristic line unwrapping separately rather than combining many changes.
  • Compare PDF engines while keeping every other option unchanged.
  • Inspect the OCR text layer when copied text is inaccurate or scrambled.
  • Use conversion debug output to locate the first broken pipeline stage.
  • Manually clean the EPUB once most prose has been reconstructed correctly.
  • Use a structured source format when the PDF's reading order is unusable.

8. Frequently Asked Questions

8.1 Why does calibre put a break after every PDF line?

The PDF likely stores each visual line as a separate positioned text fragment. calibre must estimate which fragments form a paragraph. Adjust the line un-wrapping factor, remove headers and footers, and check whether the PDF text layer has a sensible reading order.

8.2 Should I enable heuristic processing?

Test it only after a normal PDF conversion. The Unwrap lines heuristic can repair surviving hard breaks, but it can also join headings, lists, poetry, dialogue, or unrelated blocks. Keep it only if a controlled sample becomes clearly better.

8.3 Can metadata or a device profile cause early line breaks?

Ordinary title, author, cover, and identifier metadata does not reconstruct PDF paragraphs. Device profiles can influence output sizing and presentation, but systematic breaks at the PDF's original margin usually originate in PDF input processing. Verify the EPUB locally before changing device settings.

8.4 Why do page numbers appear inside paragraphs?

The page number is part of the PDF's text content and interrupts the sequence at a page boundary. Use the PDF header and footer controls or a carefully tested search-and-replace rule. Remove it before judging whether paragraph unwrapping is working.

8.5 Can calibre fix a scanned PDF automatically?

Not if the scan lacks a usable text layer. The document first needs accurate OCR performed with an appropriate tool, or you need another text-based source. If OCR exists but is inaccurate, correcting or recreating it is usually more effective than repeatedly changing calibre settings.

8.6 When should I stop troubleshooting and find another format?

Stop when copied text and debug output show interleaved columns, missing content, severe OCR errors, duplicated layers, or fundamentally incorrect reading order. An original EPUB, DOCX, HTML, or other structured source will usually produce a much better result than attempting to reconstruct complex pages from PDF.


Citations

  1. Official guidance on PDF conversion, paragraph unwrapping, headers, footers, heuristics, and debug pipeline stages. (calibre E-book Conversion Manual)
  2. Official command reference for PDF input engines, unwrapping, header and footer removal, and pipeline debugging. (calibre ebook-convert Documentation)
  3. Official instructions for editing EPUB files and using search and replace for manual cleanup. (calibre E-book Editor Manual)
  4. Official format-conversion FAQ explaining PDF limitations and preferred source formats. (calibre Frequently Asked Questions)
  5. Official reference for calibre debugging commands and component debug modes. (calibre-debug Documentation)
Cindy, ContentBASE creator assistant

MEET CINDY

Your ContentBASE creator assistant

Cindy helps creators find Canva templates, content ideas, and simple ways to make better social media posts faster.

Want ready-to-use templates? Claim the free Canva bundles or browse the full bundle store.