convert pdf to html in python
Converting PDFs to HTML in Python enables web‑friendly rendering, preserving layout and text flow. Popular tools include PDFMiner for pure text extraction, Apryse SDK for rich features like images and annotations, and Spire.PDF for quick, license‑free conversion. Each offers distinct trade‑offs. Learn now

Why Convert PDFs to HTML in Python?
PDF documents are ubiquitous in business, education, and research, yet their fixed layout hinders web integration, search engine indexing, and accessibility compliance. Converting PDFs to HTML with Python unlocks the underlying text and structure, enabling responsive rendering across devices, embedding interactive elements, and applying CSS styling for brand consistency. Python’s ecosystem offers several mature libraries—PDFMiner for low‑level text extraction, Apryse SDK for advanced features such as image handling, annotations, and PDF/A compliance, and Spire.PDF for quick, license‑free conversion. These tools support batch processing, allowing thousands of files to be transformed automatically with a single script, dramatically reducing manual effort and human error. Automated conversion also facilitates downstream analytics; extracted text can feed natural language processing pipelines, sentiment analysis, or compliance monitoring systems. Moreover, HTML output can be further processed into Markdown, JSON, or plain text, providing flexibility for content management systems, static site generators, or API services. Python’s integration with cloud services (AWS Lambda, Azure Functions, Google Cloud Functions) enables scalable, serverless workflows that scale with demand. The ability to programmatically adjust extraction parameters—such as page ranges, text extraction modes, or layout preservation—gives developers fine‑grained control over output quality. Finally, converting PDFs to HTML supports accessibility standards (WCAG) by producing semantic markup that screen readers can interpret, improving inclusivity for users with disabilities. In short, Python‑based PDF to HTML conversion is essential for modern digital workflows, enhancing discoverability, user experience, and operational efficiency across industries. The process also integrates seamlessly with version control systems, allowing teams to track changes, revert to previous states, and maintain a clear audit trail for compliance purposes daily.

Key Libraries and SDKs
When converting PDFs to HTML in Python, developers typically rely on a handful of mature libraries and SDKs that balance ease of use, feature breadth, and performance. The most common options include the open‑source PDFMiner.six, which excels at extracting raw text and layout information; the commercial Apryse SDK, which offers a rich API for image extraction, annotation handling, and PDF/A compliance; and the lightweight Spire.PDF for Python, which provides a straightforward command‑line interface and a permissive license for rapid prototyping. In addition, the PyMuPDF (fitz) library offers fast rendering and high‑quality image extraction, while pdfplumber focuses on table extraction and structured data. Each of these tools supports batch processing, custom output formatting, and integration with cloud services, making them suitable for large‑scale workflows. Choosing the right library depends on factors such as the complexity of source PDFs, required fidelity of images and annotations, licensing constraints, and the need for automated scaling. By evaluating these criteria, teams can select a solution that delivers accurate HTML output while fitting within their development and deployment pipelines.
Installation and Setup
Before converting PDFs to HTML, set up a clean Python environment. Create a virtual environment with python -m venv venv and activate it. Install PDFMiner.six via pip install pdfminer.six, which provides the pdf2txt.py utility for basic extraction. For richer output, add Apryse SDK: download the Python wheel from Apryse and install with pip install apryse-sdk‑python‑x.x.x‑py3‑win‑amd64.whl (replace version numbers). Apryse requires an API key; set it in the environment with export APRYSE_API_KEY=your_key or set APRYSE_API_KEY=your_key on Windows. Spire.PDF for Python can be installed via pip install spire.pdf and does not need external dependencies. Verify installations by running pdf2txt.py --version, import apryse in Python, and import spire.pdf. Keep libraries updated with pip install --upgrade commands. Finally, organize your scripts in a src folder and use requirements.txt to lock versions for reproducibility.
To streamline dependency management, requirements.txt with pip freeze > requirements.txt use pip‑compile from pip‑tools to lock dependencies. and lock!!.
Remember to test the conversion locally before scaling. Use a small sample set to validate layout fidelity, image quality, and text encoding. Capture any errors in a log file and iterate so until output meets standards.

Converting a Single PDF with PDFMiner
PDFMiner.six is a pure‑Python library that parses PDF files and exposes their textual content, layout, and font information. To convert a single PDF to HTML, the standard command‑line tool pdf2txt.py is used. The basic invocation looks like this:
pdf2txt.py -o output.html -t html input.pdf
Here -o specifies the output file, -t html tells the tool to emit HTML, and input.pdf is the source document. The resulting HTML preserves line breaks and basic formatting, but complex layouts may require additional tweaking. For finer control, the pdfminer.high_level.extract_text_to_fp function can be called directly from a Python script, passing a io.StringIO object and a laparams dictionary to adjust spacing and character extraction rules. Example code:
from pdfminer.high_level import extract_text_to_fp
from pdfminer.layout import LAParams
import io
output = io.StringIO
with open('input.pdf', 'rb') as fp:
extract_text_to_fp(fp, output, laparams=LAParams, output_type='html')
html_content = output.getvalue
After generating the HTML string you can write it to a file or serve it directly in a web application. Note that PDFMiner focuses on text extraction; images, vector graphics, and annotations are omitted unless additional libraries such as pdf2image are combined. For most simple documents, the command‑line approach is sufficient and requires no extra dependencies beyond PDFMiner itself.
When dealing with multi‑page PDFs, PDFMiner automatically iterates over each page, inserting <div class="page"> wrappers in the HTML output. If you need to preserve page breaks as <hr> elements, you can post‑process the HTML string with a regex or use BeautifulSoup to replace the wrappers. Additionally, PDFMiner’s LAParams class offers fine‑grained control over line spacing, word spacing, and character grouping. Setting all_texts=True forces extraction of all text, even if it is hidden or in a different layer, which can be useful for PDFs with annotations.
For large PDFs, the command‑line tool may consume significant memory because it loads the entire document into memory. To mitigate this, use the PDFPage.get_pages iterator to process pages sequentially and write incremental HTML fragments to disk. This approach keeps memory usage low and allows you to resume processing if the script crashes.
PDFMiner.six is distributed under the BSD license, making it suitable for commercial projects without licensing fees. Ensure you comply with the license when redistributing the library or its output and requires no extra dependencies.
When thousands of PDFs must be converted, a single‑pass script that iterates over a directory is essential. PDFMiner’s PDFPage.get_pages iterator allows page‑by‑page extraction, keeping memory usage low. The following pattern demonstrates a robust batch pipeline:
from pathlib import Path
from pdfminer.high_level import extract_text_to_fp
from pdfminer.layout import LAParams
import io
def pdf_to_html(pdf_path,out_path):
out=io.StringIO
with open(pdf_path,'rb') as fp:
extract_text_to_fp(fp,out,laparams=LAParams,output_type='html')
with open(out_path,'w',encoding='utf-8') as out_file:
out_file.write(out.getvalue)
def batch_convert(src_dir,dst_dir):
src_dir=Path(src_dir)
dst_dir=Path(dst_dir)
dst_dir.mkdir(parents=True,exist_ok=True)
for pdf in src_dir.rglob('*.pdf'):
rel=pdf.relative_to(src_dir)
out_path=dst_dir/rel.with_suffix('.html')
out_path.parent.mkdir(parents=True,exist_ok=True)
pdf_to_html(pdf,out_path)
if __name__=='__main__':
batch_convert('incoming','converted')
Key points:

- Use
Path.rglobto discover PDFs recursively. - Maintain directory structure by mirroring relative paths.
- Write each HTML file atomically to avoid partial writes on interruption.
- Wrap the conversion in a try/except block to log failures without stopping the whole run.
- Leverage
multiprocessing.Poolfor parallelism on multi‑core machines, passingpdf_to_htmlas the worker function.
For very large corpora, store intermediate results in a database to decouple ingestion from conversion. PDFMiner’s pure‑Python nature simplifies deployment on CI/CD pipelines cloudfunctions. process can be automated with cron jobs.!

Using Apryse SDK for Advanced Features

Apryse SDK provides a RESTful API that can be called from Python. Authentication is via API key or OAuth token passed in the Authorization header. The convertToHtml endpoint accepts a PDF file or URL and returns a fully rendered HTML document that preserves layout, fonts, images, and annotations. Optional parameters such as pageRange, imageQuality, and outputFormat allow fine‑tuning.
- Image and Annotation Support: All raster and vector graphics, as well as annotations, are extracted and embedded in the HTML. Images can be embedded as base64 data URIs or written to a separate folder via the
embedImagesflag. - Performance and Scaling: The SDK runs on microservices, enabling horizontal scaling. For batch jobs, the
asyncConvertendpoint queues requests and returns the final job ID. PollinggetJobStatusretrieves the processed HTML once complete. - Customizable Output: CSS classes can be injected via the
customCssparameter, and thelayoutModeoption (e.g.,singlePage,continuous) controls page stitching.
from apryse_sdk import ApryseClient
client = ApryseClient(api_key='YOUR_KEY')
response = client.convert_to_html(file_path='sample.pdf', page_range='1-5', embed_images=True)
with open('output.html', 'w', encoding='utf-8') as f:
f.write(response.content)
Apryse SDK provides PDF‑to‑HTML conversion with image handling and annotation support, batchnow processing.

Leveraging Spire.PDF for Simple Conversion
Spire.PDF for Python offers a lightweight, license‑free library that instantly converts PDFs to clean HTML. Install via pip, then load the PDF, call to_html, and write the output. No external dependencies, making it ideal for quick, reliable batch jobs. API returns bytes that can be written.
PDFMiner: Installation and Basic Commands
To begin, install the actively maintained fork with pip install pdfminer.six. Once available, the command‑line tool pdf2txt.py can be invoked directly: pdf2txt.py -o output.html -t html input.pdf. This produces a minimal HTML file containing extracted text wrapped in <pre> tags, preserving whitespace. For programmatic use, import the high‑level API: from pdfminer.high_level import extract_text_to_fp. Create a BytesIO buffer, call extract_text_to_fp(buffer, pdf_file, codec='utf-8', laparams=None), and then write buffer.getvalue to an .html file. The laparams argument can be tuned with LAParams to influence layout detection. Basic error handling can be added by wrapping calls in try/except blocks. For batch processing, loop through a directory of PDFs, applying the same command or API call to each file. This approach keeps dependencies minimal and works well for simple, text‑heavy documents. Additionally, the library supports custom layout analysis via the LAParams object, allowing fine‑grained control over text extraction thresholds, word spacing, and line gaps. For PDFs with complex columnar layouts, setting laparams.word_margin and laparams.line_margin can help preserve column integrity. When dealing with scanned images, pdfminer can be paired with OCR libraries such as Tesseract to first render text before conversion. Use context managers! close.
PDFMiner: Handling Text Extraction and Layout
PDFMiner’s core strength lies in parsing the PDF object model and reconstructing logical text streams. The high‑level extract_text function traverses page layouts, grouping LTTextBox and LTTextLine objects into coherent strings. By default, the library preserves line breaks and indentation, but developers can fine‑tune the LAParams object to adjust word_margin, char_margin, and line_margin. These parameters control how closely characters are considered part of the same word or line, which is crucial for multi‑column documents or tables. When word_margin is set too low, words that are actually separated may merge; too high, and legitimate spaces may be lost. For complex layouts, LAParams can be combined with PDFPageAggregator to retrieve the full layout tree, enabling custom rendering logic. PDFMiner exposes LTChar objects, giving access to font information and positioning, which can be used to reconstruct CSS styles for a faithful HTML output. In practice, a typical workflow iterates over PDFPage.get_pages, feeds each page to a PDFPageInterpreter, and collects the resulting LTPage objects. The extracted text can then be wrapped in HTML tags, such as <h1> for headings detected by font size or <table> for tabular data inferred from spatial coordinates. While PDFMiner excels at text extraction, its layout reconstruction is heuristic and may struggle with stylized PDFs, complex graphics, or scanned images without OCR. In those cases, integrating OCR or switching to a commercial SDK may be necessary now!!!
PDFMiner: Limitations with Complex PDFs
Despite its robust parsing engine, PDFMiner often falters when faced with documents that deviate from the “plain‑text” paradigm. Scanned PDFs, for instance, contain raster images rather than vector text; PDFMiner will treat these as empty LTTextBox objects, yielding no output unless an OCR layer is added. Complex layouts with multiple columns, overlapping text, or embedded graphics confuse the default LAParams heuristics, causing words to merge or split incorrectly. Tables are especially problematic: the library cannot reliably detect cell boundaries, so data that should appear in a structured <table> is rendered as a flat stream of text. Additionally, PDFMiner’s reliance on font metrics means that documents using custom or missing fonts may produce garbled characters or placeholder glyphs. When PDFs contain annotations, hyperlinks, or form fields, the extraction process ignores these interactive elements, resulting in a loss of semantic information. Performance-wise, PDFMiner can be slow on large files because it parses every object in the PDF stream; memory consumption spikes when handling high‑resolution images embedded within the document. Finally, the library lacks built‑in support for CSS styling; recreating the original visual appearance requires manual mapping of font sizes, colors, and positions, which is error‑prone and time‑consuming. For projects demanding high fidelity, developers often supplement PDFMiner with OCR tools like Tesseract recovers text from scanned pages, improving conversion accuracy full
Apryse SDK: API Overview and Authentication
The Apryse Server SDK exposes a RESTful interface that accepts PDF documents and streams back HTML, CSS, and image assets. Endpoints are versioned under /v1 and support multipart/form‑data uploads. A typical conversion request looks like:
POST https://api.aprise.com/v1/convert/pdf-to-html
Headers: Authorization: Bearer <token>
Body: file=@document.pdf
Authentication is token‑based. After registering an account, you receive an API key and secret. The SDK provides helper functions to generate a signed JWT or HMAC token. For example, in Python:
import aprise
client = aprise.Client(api_key="YOUR_KEY", api_secret="YOUR_SECRET")
response = client.convert_pdf_to_html("sample.pdf")
Requests can be customized with query parameters such as pageRange, includeImages, and layoutMode. The service returns a JSON payload containing html, css, and an array of image URLs. Error handling follows standard HTTP status codes; 401 indicates missing or invalid credentials, 429 signals rate limits, and 500 denotes server‑side failures.
Token generation follows the OAuth 2.0 standard. The SDK offers a helper that accepts a client ID and secret, then exchanges them for an access token via a POST request to the token endpoint. The returned JSON includes an access_token and an expires_in field. The SDK refreshes token when it nears expiry conversions without intervention.
Apryse SDK: Image and Annotation Support
Apryse SDK preserves visual fidelity by extracting raster and vector images as separate assets. Small images are inlined as base64 URIs; large ones are served from a CDN. The SDK exposes an imageQuality option to balance resolution and file size. Vector graphics are rendered as crisp .svg files, ensuring scalability.
Annotations are rendered as interactive HTML overlays. The SDK maps PDF annotation types to HTML elements: highlight becomes a <span> with a background color; note becomes a tooltip using <div> with a data-tooltip attribute; checkbox and radio fields become <input type="checkbox"> or <input type="radio"> elements.
To extract annotations, set includeAnnotations=true in the request. The response JSON contains an annotations array with type, coordinates, and content. Developers can apply custom JavaScript to enhance interactivity, such as opening a modal for a note or toggling field visibility.
Image extraction is GPU‑accelerated, reducing conversion time for high‑resolution PDFs. The SDK supports batch processing of images, queuing multiple PDFs and providing progress callbacks. This feature is critical when handling thousands of documents, preventing image handling from becoming a bottleneck.
Apryse’s image and annotation handling delivers faithful HTML for complex PDFs!!!!!!!!
Apryse SDK: Performance and Scaling
Apryse SDK uses async I/O and multi‑threaded rendering powered by native C++ libraries. Python bindings offload CPU‑intensive tasks to background workers. The maxConcurrency option tunes parallel conversions, balancing memory and speed.
For batch jobs, submit PDFs via the job queue API. The service returns a job ID and status endpoint for progress. This decouples client requests from long‑running tasks, enabling thousands of files to be processed without blocking.
Memory usage stays low by streaming PDFs directly from disk or network, rendering one page at a time. The SDK discards intermediate data, keeping RAM below 200 MB per conversion even for large documents.

Latency drops with caching of images and fonts. The SDK stores rendered assets, so repeated conversions hit the cache. CDN caching further reduces network round‑trips.

Scalability is achieved by horizontal scaling. Deploy multiple SDK instances behind a load balancer, each with maxConcurrency. Stateless design lets you spin up or down instances based on demand.
In production, monitor metrics like requestsPerSecond, averageConversionTime, and memoryUtilization. The SDK exposes Prometheus‑compatible metrics for Grafana dashboards, giving real‑time visibility.
Overall, Apryse SDK delivers sub‑second conversions for small PDFs and under a minute for large, complex documents, making it ideal for enterprise‑grade batch processing.!
Spire.PDF: Quick Setup and Licensing
Spire.PDF for Python offers a plug‑and‑play installation: pip install spire.pdf pulls a single wheel that contains all native binaries. After installation, a quick import test verifies the license key: import spire.pdf as sp; sp.set_license("YOUR_KEY"). The free tier supports up to 50 pages per document and is ideal for proof‑of‑concepts. For production, the commercial license unlocks pages, batch processing, and advanced features like form handling and signatures. The license key is a 32‑character string that can be embedded in code or stored to avoid hard‑coding. Spire.PDF ships with a license manager that validates the key against a remote server; if the key is invalid or expired, the library falls back to a restricted mode. The licensing API is lightweight, swift! requiring only a single HTTP call during initialization, so it does not add noticeable latency to conversion pipelines. To streamline deployment, the library supports Docker images that include the license file, enabling consistent behavior across CI/CD environments. The licensing model is per‑user, not per‑instance, so you can spin up multiple containers without purchasing additional keys. The documentation recommends keeping the license key out of source control and using sec tools and Vault or AWS Secrets Manager. With these practices, you can integrate Spire.PDF into automated workflows, from nightly data extraction to real‑time web services, without worrying about licensing constraints or performance bottlenecks.
Spire.PDF: Customizing HTML Output
Spire.PDF lets developers customize the HTML via HtmlSaveOptions. Set options.EmbedImages = true to embed images as base64, making the file self‑contained. Toggle options.EmbedFonts = false to link external fonts and reduce size. For layout, use options.PageMode = PageMode.SinglePage and options.PageLayout = PageLayout.SinglePageContinuous. Inline CSS is enabled with options.UseCssStyles = true, useful for emails. Multi‑column PDFs can be mimicked by options.ColumnCount = 2. The options.OnElementCreated event lets you alter elements—e.g., replace <img> tags or inject scripts after <body>. Export to HTML5 or HTML4 via options.HtmlVersion = HtmlVersion.Html5. Choose options.ExportMode = ExportMode.Full to keep annotations, or ExportMode.Simple for a clean view. Combine these options programmatically or via a JSON file for precise control.
Deploying Spire.PDF in a microservice, a FastAPI endpoint can accept a PDF, stream HTML, and set Content-Type: text/html headers. This isolates conversion and scales with Kubernetes. Persisting HtmlSaveOptions in CI/CD pipelines automates preview generation. Caching results in Redis or a CDN cuts CPU usage. Spire.PDF’s logs can be routed to a log aggregator for easier debugging.
By exposing HtmlSaveOptions as JSON, developers can version configurations, builds in CI pipelines.!