PDF conversion provides basic text extraction from PDF files. PDF is fundamentally a print-oriented format with fixed layout, so the conversion focuses on recovering readable text rather than preserving visual positioning.

PDF conversion is the roughest path in Honzo. Images, formatting, tables, and fonts are not carried over. For structured documents, use EPUB instead.

How it works

  1. Extract text. PDF text content is extracted page by page using text show operators.

  2. Detect reading order. The converter applies heuristics to reconstruct paragraph and chapter boundaries from positioned text.

  3. Split by pages. Each PDF page becomes a separate Honzo CHAP entry.

  4. Extract metadata. Title, author, and subject come from the PDF Info dictionary or XMP metadata.

  5. Build. All extracted content is assembled into the final Honzo file.

Metadata mapping

Title

Maps to Honzo title in META section.

Author

Maps to Honzo creator.

Subject

Maps to subject[] array.

Language

Maps to Honzo language (BCP 47).

Preserved features

Page structure

One CHAP chunk per PDF page.

Text content

Extracted as Markdown text.

Metadata

Title, author preserved in META.

Limitations
  • PDF layout and formatting are not preserved. Expect plain text without bold, italic, or font styling.
  • Images embedded in PDFs are not extracted.
  • Tables and multi-column layouts may lose their structural relationships.
  • Embedded fonts are not carried over.
  • PDF forms, annotations, and interactive elements are ignored.
  • OCR-based PDFs (scanned documents) produce no usable text output.

Example

honzo-cli convert book.pdf book.hzo

Next Steps