Section: User Commands (1)
Return to Main Contents
pdftohtml - program to convert PDF files into HTML, XML and PNG images
[options] <PDF-file> [<HTML-file> <XML-file>]
This manual page documents briefly the
This manual page was written for the Debian GNU/Linux distribution
because the original program does not have a manual page.
is a program that converts PDF documents into HTML. It generates its output in
the current working directory.
A summary of options are included below.
- -h, -help
Show summary of options.
- -f <int>
first page to print
- -l <int>
last page to print
do not print any messages or errors
print copyright and version info
exchange .pdf links with .html
generate complex output
generate single HTML that includes all pages
generate no frames. Not supported in complex output mode.
use standard output
- -zoom <fp>
zoom the PDF document (default 1.5)
output for XML post-processing
- -enc <string>
output text encoding name
- -opw <string>
owner password (for encrypted files)
- -upw <string>
user password (for encrypted files)
force hidden text extraction
image file format for Splash output (png or jpg).
If complex is selected, but -fmt is not specified,
-fmt png will be assumed
do not merge paragraphs
override document DRM settings
- -wbt <fp>
adjust the word break threshold percent. Default is 10.
Word break occurs when distance between two adjacent characters is
greater than this percent of character height.
outputs the font name without any substitutions.
Pdftohtml was developed by Gueorgui Ovtcharov and Rainer Dorsch. It is
based and benefits a lot from Derek Noonburg's xpdf package.
This manual page was written by Søren Boll Overgaard <email@example.com>,
for the Debian GNU/Linux system (but may be used by others).
- SEE ALSO
This document was created by
using the manual pages.
Time: 19:38:56 GMT, January 18, 2019