Conversion
Honzo includes conversion tools for EPUB, MOBI, and PDF formats. Conversion preserves chapter structure, metadata, images, stylesheets, and fonts where possible.
Supported formats
EPUB 2/3
Table of contents, metadata, images, CSS, and fonts supported.
MOBI
Text, metadata, and basic formatting.
PDF
Text extraction only. No reflow preservation.
Converting files
CLI
Rust API
# EPUB to Honzo
honzo-cli convert book.epub book.hzo
# MOBI to Honzo
honzo-cli convert book.mobi book.hzo
# PDF to Honzo
honzo-cli convert book.pdf book.hzo
use honzo_convert::convert_epub;
let hzo = convert_epub("book.epub").unwrap();
std::fs::write("book.hzo", hzo).unwrap();
What gets converted
EPUB
| Source | Honzo target | Details |
|---|---|---|
| OPF metadata | META section | Title, creator, language, identifiers, subjects |
| NCX / nav | CHAP chunks + TOC | Chapter splitting follows the EPUB spine |
| XHTML content | CHAP chunks (HTML) | Full HTML preserved, including in-line images |
| Images | IMG_ chunks | JPEG, PNG, WebP embedded |
| CSS | CSS_ chunks | Stylesheets preserved |
| Fonts | FONT chunks | Embedded fonts carried over |
| Cover | COVR + COVT | Full cover + thumbnail |
| Page breaks | EXTRA (via pagebreaks.rs) | EPUB pagebreak markers converted |
MOBI
| Source | Honzo target | Details |
|---|---|---|
| Metadata | META section | Title, author, language |
| Text content | CHAP chunks (HTML) | Basic HTML conversion |
| Images | IMG_ chunks | Embedded images |
| Source | Honzo target | Details |
|---|---|---|
| Extracted text | CHAP chunks (Markdown) | Text flow, no layout preservation |
| Page breaks | Chapter splitting | One PDF page per chapter |
Page break detection
For EPUBs without explicit pagebreak markers, Honzo uses a character count heuristic:
use honzo_convert::pagebreaks::estimate_page_breaks;
let breaks = estimate_page_breaks(&content, None);
// Returns around 2000 character intervals for English text
Explicit EPUB pagebreak patterns are also detected:
<!-- These are all recognized -->
<span epub:type="pagebreak" id="pg42" title="42" />
<span class="pagebreak" id="page-42" />
<a id="page42" class="pagebreak"></a>
<hr class="pagebreak" />
<div class="pagebreak" title="42" />
Conversion options
Override metadata during conversion:
honzo-cli convert book.epub book.hzo \
--title "Custom Title" \
--author "Custom Author" \
--language "fr"
Next Steps
Detailed guides
For format-specific conversion details: