- Distinguish incorrect encoding from missing font glyphs before changing conversion settings.
- Test UTF-8, cp1252, HTML declarations, and a short representative chapter.
- Inspect logs and pipeline output before reinstalling calibre or deleting anything.
- Confirm the Symptom With a Small Safe Test
- Set the Correct Input Character Encoding
- Separate Encoding Damage From Missing Font Glyphs
- Choose a Better Source Format When Possible
- Rule Out Library, Plugin, Delivery, and Device Effects
- Use Conversion Logs and Debug Output
- Run a Clean Temporary Test Before Reinstalling
- Quick Fix Checklist
- Frequently Asked Questions
When accented letters, Cyrillic, Asian text, smart quotes, mathematical symbols, or other special characters become garbled after an e-book conversion, the problem usually falls into one of two categories. Either calibre decoded the source using the wrong character encoding, or the converted book is using a font that cannot display the required glyphs. The fastest solution is to determine whether the damage exists in the source, appears during conversion, or occurs only when the finished book is displayed on a particular device or app. Work through the checks below in order, and stop as soon as a small test conversion displays correctly.

Start with free Canva bundles
Browse the freebies page to claim ready-to-use Canva bundles, then get 25% off your first premium bundle after you sign up.
Free to claim. Canva-ready. Instant access.
1. Confirm the Symptom With a Small Safe Test
Do not begin by reinstalling calibre, deleting your library, or changing several conversion options. First, isolate the failure with a small test that contains the exact characters causing trouble.
1.1 Identify What the Broken Text Looks Like
The appearance of the damaged text provides useful clues:
- Sequences such as é, ñ, or ’: UTF-8 text was probably decoded as a legacy encoding such as Windows-1252, also called cp1252.
- Replacement diamonds or question marks: the source may already have lost characters, the decoder could not interpret the bytes, or the display font lacks suitable glyphs.
- Empty squares, boxes, or tofu symbols: the characters may be intact, but the selected font or reading device cannot display them.
- Only smart quotes and dashes are broken: the likely cause is a mismatch involving cp1252, UTF-8, or an incorrect encoding declaration in HTML.
- Correct text in calibre but broken text on a device: investigate fonts, the output format, and the device's rendering limitations before changing the input encoding.
1.2 Check the Source Before Converting
Open the original document in an application appropriate for its format. Use a text editor that can report the encoding for TXT or HTML files. For DOCX, open the file in a word processor. For EPUB or AZW3, use calibre's Edit book tool and inspect the affected passage in both the code pane and preview.
If the original already displays incorrect characters, conversion settings cannot reliably reconstruct the intended text. Return to an undamaged source, correct the source in its native application, or obtain a properly encoded copy. Success at this stage means the original passage displays the intended characters before calibre processes it.
1.3 Create a Representative Test Chapter
Make a copy containing one short chapter or several paragraphs. Include ordinary English text and representative characters such as é, ñ, Ł, Ж, 中文, 日本語, curly quotation marks, an em dash, and any specialist symbols used in the book. Save it in the same format and encoding as the real source.
Add the test as a separate book and convert it to the intended output format. A small test makes each attempt faster and prevents confusion caused by cached files, device copies, or older converted formats. If it works, apply the same setting to the full book and stop experimenting.
2. Set the Correct Input Character Encoding
The input character encoding tells calibre how to translate the source file's stored bytes into Unicode characters. It is especially important for TXT and older HTML documents that lack a reliable encoding declaration.
2.1 Find the Input Character Encoding Setting
Select the test book, choose Convert books, and open the Look & feel section followed by its Text tab. Locate Input character encoding. Menu wording or placement can vary slightly, so use the conversion dialog's search or setting tooltips if necessary.
Leave this field empty when calibre detects the source correctly. If the test output is garbled, enter a known encoding rather than trying unrelated appearance settings. This option can override an absent or incorrect declaration in the source.
2.2 Try UTF-8 When the Source Is Modern
UTF-8 can represent the full Unicode character range and is the normal choice for newly created multilingual TXT and HTML files. Select UTF-8 when your editor reports UTF-8, the HTML declares UTF-8, or you deliberately saved the file that way.
Convert the small test again and inspect the same passage. Success means accented letters, Asian scripts, Cyrillic text, punctuation, and symbols appear exactly as they do in the source. If they do, keep UTF-8 for this source and stop changing encoding options.
2.3 Try cp1252 for Older Western European Files
Windows-1252, commonly written as cp1252, is frequently found in older text and HTML created by Windows software. It covers Western European accented letters and typographic punctuation, but it cannot represent scripts such as Chinese, Japanese, or Cyrillic comprehensively.
Try cp1252 when an older Western-language file contains broken curly quotes, apostrophes, dashes, euro signs, or accented Latin letters. Do not use it as a universal fix for multilingual content. If UTF-8 text is mistakenly read as cp1252, familiar corruption such as é or ’ can appear.
Success means the Western European text and smart punctuation are restored without damaging other characters. If the book contains scripts outside cp1252 and those scripts are missing in the source, locate a Unicode version instead of repeatedly converting the damaged file.
2.4 Handle HTML Import Encoding Separately
HTML can be interpreted when it is first added to the library and again during conversion. An incorrect declaration inside the HTML, such as a wrong charset value, can therefore create problems before the conversion dialog is opened.
Inspect the source HTML's charset declaration and compare it with the encoding reported by a capable text editor. Correct the declaration or resave the files consistently as UTF-8. calibre also provides customization for the HTML-to-ZIP file-type plugin under Preferences, Advanced, and Plugins. This can help when adding HTML that does not identify its encoding correctly.
Because HTML from different sources may use different encodings, do not treat a plugin override as a permanent global cure. Test it with one copied file, then restore the setting if other HTML imports need different handling.
3. Separate Encoding Damage From Missing Font Glyphs
Encoding determines which characters the book contains. Fonts determine how those characters look. Changing the input encoding will not fix a character that is stored correctly but displayed as an empty square because the active font lacks that glyph.
3.1 Inspect the Converted Text in Edit Book
Convert the test to EPUB or another editable format supported by calibre, then open Edit book. Find an affected phrase in the HTML source.
- If the HTML contains the correct character but the preview shows a box, suspect the font or CSS.
- If the HTML itself contains sequences such as é, suspect incorrect decoding during import or conversion.
- If the HTML contains a question mark where a character should be, check whether the source was already damaged.
This inspection prevents an endless cycle of encoding changes when the real issue is display support.
3.2 Test Without the Book's Forced Font
A stylesheet may force a decorative or highly limited font throughout the book. In Edit book, inspect the CSS for font-family and @font-face rules. As a temporary test, remove or disable the forced family in a copy and allow the viewer or device to use its default font.
If the characters appear with the default font, the text is correctly encoded. The original font lacks the necessary glyphs, is incorrectly referenced, or is unsupported by the destination reader. Stop changing encoding settings and correct the font configuration instead.
3.3 Understand Embedded Fonts
An embedded font travels inside supported e-book formats, which can make specialized scripts or symbols display more consistently. calibre can embed referenced fonts found on the computer when the output format supports embedding. You must also have permission to embed the font.
Embedding is useful only when the selected font actually contains every required glyph. Embedding a Latin-only font will not add Chinese, Japanese, Cyrillic, or specialist mathematical characters to it. Choose a font with the appropriate script coverage and test the actual output on the target device.
Be cautious with subset fonts. A subset contains only selected glyphs to reduce file size. If text was added after subsetting, the new characters may be absent. Subsetting should normally be the final production step, after all text has been checked.
3.4 Compare calibre's Viewer With the Target Device
Open the converted file directly in calibre's viewer before sending it anywhere. If it is correct there but wrong on an e-reader or reading app, the conversion may already be successful. The destination could be ignoring embedded fonts, substituting its own font, or lacking support for the script or output format.
Try the device's built-in font choices and a widely supported output format. Test the same file in another independent reader application. Success means the characters remain correct in the file and at least one capable renderer. At that point, focus on the destination device rather than reconverting repeatedly.
4. Choose a Better Source Format When Possible
Source quality strongly affects conversion reliability. A clean DOCX or well-formed UTF-8 HTML file usually preserves multilingual text more predictably than plain text with an unknown encoding.
4.1 Prefer DOCX Over Ambiguous TXT
DOCX stores text as Unicode within a structured document package. It also preserves paragraphs, headings, italics, and other formatting. TXT has no universal internal marker that guarantees how its bytes should be decoded, so calibre may have to guess.
If you control the manuscript, open the original in a word processor, confirm every character, and save a clean DOCX. Convert that DOCX rather than exporting an ambiguously encoded TXT file. Do not create a DOCX by merely renaming the TXT extension.
Success means the test DOCX converts correctly without forcing an input encoding and retains its structure. If so, use DOCX for the full conversion.
4.2 Use UTF-8 When TXT Is Required
If TXT is necessary, explicitly save it as UTF-8 in a text editor. Close and reopen the saved file to verify it before adding it to calibre. Then set the conversion's input character encoding to UTF-8 if automatic detection still fails.
This approach removes guesswork. If the reopened source and converted test both show the correct characters, the encoding path is working.
4.3 Be Careful With PDF Sources
A PDF may display letters correctly while storing incomplete font mappings, positioned glyphs, or image-only pages underneath. Copy a problematic sentence from the PDF and paste it into a plain-text editor. If the pasted text is scrambled, missing, or replaced with unrelated characters, the PDF's text layer is unreliable.
Changing calibre's input encoding is unlikely to repair a broken PDF character map. Use the original DOCX, HTML, EPUB, or other source from which the PDF was produced when available. If the pages are scans, text recognition may be required before conversion, followed by careful proofreading.

5. Rule Out Library, Plugin, Delivery, and Device Effects
Once a small local conversion is correct, check whether another step is modifying what you inspect or causing an older file to be opened.
5.1 Confirm Which Format You Are Viewing
A calibre book record can contain several formats. Verify that you opened the newly converted EPUB, AZW3, or other output rather than the original or an older conversion. Check the conversion job's completion time and remove obsolete test copies from the device before transferring the new file.
Changing metadata such as the title, author, language, or tags does not repair corrupted body text. However, metadata imported from a damaged source can itself display broken characters. Correct metadata separately after confirming that the book's content is sound.
5.2 Test Without Optional Plugins
A file-type or conversion plugin can alter the import path. If the problem appears only for files processed by a particular plugin, update or temporarily disable that plugin and repeat the small test using calibre's built-in handling. Change one plugin at a time.
Success means the same source converts correctly without the optional plugin. Keep the plugin disabled for that workflow or consult its documentation before restoring it.
5.3 Test the Local File Before Email or Server Delivery
Email delivery and the calibre Content server are not the first places to troubleshoot encoding if the locally converted file is already broken. Open the output from the library first. If it is correct locally, download or transfer it and compare the exact file rather than relying on a cached copy in the receiving app.
For USB transfers, safely eject the device and confirm that the newly transferred file replaces the old one. If a device performs an additional automatic conversion, compare a direct transfer in a supported format. Cloud synchronization can also leave duplicate or stale copies, so use a distinctive test title and remove earlier versions.
5.4 Consider Security and File Access Only When Jobs Fail
Permissions, antivirus tools, and protected folders are relevant when calibre cannot read the source, write the output, or finish the conversion job. They are less likely to be responsible when conversion completes but specific characters are garbled.
If the job reports access errors, copy the test source to a normal local folder, such as a temporary folder inside your user profile. Avoid testing from a cloud-synced, network-mounted, or read-only location. Success means the job completes and produces a readable local output. If it completes but the text remains corrupted, return to encoding and font checks.
6. Use Conversion Logs and Debug Output
Logs help identify whether the problem begins during input decoding, an intermediate transformation, or final output generation.
6.1 Read the Completed Job Details
After conversion, open calibre's jobs list and view the details for the completed job. Look for encoding warnings, input-plugin messages, missing-font reports, parsing failures, and references to malformed HTML. Save the log before running more tests.
A completed job without errors does not prove that the detected encoding was correct. Encoding guesses can be technically valid while producing the wrong characters. Use the log together with the visible symptom and source inspection.
6.2 Save the Conversion Pipeline
calibre's conversion tools can save output from different stages of the conversion pipeline. This is useful when the source looks correct but the final result does not. The command-line option is --debug-pipeline, followed by a folder where the intermediate files will be stored.
Inspect the input-stage HTML first. If the characters are already broken there, focus on source encoding or the input plugin. If they are correct in the early stage but wrong later, inspect conversion transformations, fonts, and the output format.
6.3 Use Command-Line Testing Only When Needed
Most readers can solve this issue through the graphical conversion dialog. For a reproducible advanced test, use ebook-convert with the relevant source, output, input encoding, verbose logging, and debug-pipeline options. Run the command against copied test files, not the only copy of a book.
The purpose of command-line testing is visibility, not complexity. Once the pipeline stage containing the first broken text is identified, return to the corresponding source or conversion setting. There is no benefit in continuing to change unrelated options.
7. Run a Clean Temporary Test Before Reinstalling
A clean test separates a damaged source or saved conversion preference from a wider calibre problem.
7.1 Use a New Test Record
- Copy a small, verified source file to a local folder.
- Add it to calibre as a new book with a distinctive title.
- Open the conversion dialog and restore conversion settings to their defaults where appropriate.
- Set only the confirmed input encoding, such as UTF-8 or cp1252.
- Convert to EPUB and open the exact new output locally.
- Inspect both the HTML text and its rendered appearance.
If this works, calibre's core conversion process is functioning. Compare the successful test with the original book's source format, saved conversion options, CSS, fonts, and plugin path. Do not reinstall calibre or delete the library.
7.2 Test a Temporary Library When Necessary
If a new record still inherits confusing behavior, create a temporary calibre library and add only the verified test file. Disable optional plugins for the test and avoid cloud-synced storage. Keep the normal library untouched.
Success in the temporary library points to a saved per-book conversion setting, customized plugin, or workflow difference in the normal environment. Failure with a verified UTF-8 source and default settings provides a much cleaner case for further investigation.
7.3 Know When Reinstallation Is Unlikely to Help
Reinstallation normally does not change the encoding of a damaged source, add missing glyphs to a font, or repair a device that cannot display a script. Reinstall only when calibre itself cannot start, essential components are missing, or tests indicate damaged application files. Preserve the library and configuration until the problem has been clearly identified.
8. Quick Fix Checklist
- Open the original and verify that the affected characters are correct before conversion.
- Test one short chapter containing every problematic script or symbol.
- For modern TXT or HTML, save and convert as UTF-8.
- For older Western European TXT or HTML, test cp1252.
- Check whether HTML declares an encoding that matches the file's actual encoding.
- Inspect the converted HTML to distinguish damaged text from missing font glyphs.
- Temporarily remove forced fonts and test with the reader's default font.
- Use fonts that contain the required glyphs and embed them only when permitted.
- Prefer a clean DOCX over TXT with an unknown encoding.
- Open the local output before testing USB, email, cloud, or Content server delivery.
- Check job details and save conversion pipeline output when the failure stage is unclear.
- Stop changing settings as soon as the representative test displays correctly.
9. Frequently Asked Questions
9.1 Should I Choose UTF-8 or cp1252?
Choose the encoding that the source actually uses. UTF-8 is the better default for newly created multilingual files and can represent scripts from across Unicode. cp1252 is mainly relevant to older Western European files produced by Windows software. Test a copied chapter and compare it directly with the verified source.
9.2 Why Are Characters Correct in calibre but Broken on My E-Reader?
The device may be substituting a font, ignoring an embedded font, or lacking glyphs for the script. Try another device font, inspect the book's CSS, and test the same output in an independent reading app. If the underlying HTML contains the correct characters, do not keep changing the input encoding.
9.3 Can Font Embedding Fix Garbled Sequences Such as é?
No. That sequence normally indicates that text was decoded using the wrong encoding. A font can change how a stored character is drawn, but it cannot reverse incorrect byte decoding. Correct the source declaration or select the proper input character encoding, then reconvert from the original source.
9.4 Why Does DOCX Work When TXT Does Not?
DOCX stores structured Unicode text and formatting. A TXT file is only a stream of bytes and may not clearly identify its encoding. If calibre guesses incorrectly, non-English characters can be corrupted. Saving the source as a verified DOCX or explicit UTF-8 TXT removes much of that ambiguity.
9.5 Will Changing the Book Language Metadata Fix the Text?
No. Language metadata can help readers, spelling tools, and accessibility features understand the book, but it does not normally reinterpret already corrupted body text. Set the language correctly, but repair the source encoding or font separately.
9.6 What Should I Do if Only One Book Is Affected?
Assume the problem is specific to that source until testing proves otherwise. Check its original text, encoding declaration, saved conversion settings, CSS, embedded fonts, and input format. A successful conversion of a verified test file shows that calibre is working and prevents unnecessary application-wide changes.