boxpdf

Open-source PDF-to-HTML for JavaScript

Turn PDF pages into HTML.
Keep the text, images, and layout.

@boxpdf/reader and @boxpdf/html-writer convert PDFs into visual or reflowable HTML, preserving text, images, fonts, and page layout. They run in Node, edge functions, and the browser.

Works with large PDFs Predictable memory use Fast first page Visual HTML Reflowable HTML
import { httpSource, openPdf } from "@boxpdf/reader";
import { pageToHtml } from "@boxpdf/html-writer";

const pdf = await openPdf(httpSource(url), {
  maxBytes: 2 * 1024 * 1024,
  maxObjectCacheBytes: 2 * 1024 * 1024
});

const page = await pdf.getPage(0);
viewer.innerHTML = await pageToHtml(page, { profile: "visual" });

Why @boxpdf/reader

Use PDF content in a web app.

Plain-text extractors lose columns, labels, images, and other parts of the page. Page renderers give you an image but not much structure. boxpdf gives you the content and where it appeared, as JavaScript data and HTML.

Open large files

The reader fetches requested byte ranges and keeps them within cache limits set by the application.

Run it almost anywhere

Use it with an HTTP URL, a browser File, bytes already in memory, or your own data source. It is pure JavaScript, with no native binary or headless browser.

Start on any page

Go directly to page one, page 10,000, or any page in between. The rest of the document is read only when requested.

Keep the page layout

Visual HTML includes positioned text, embedded fonts, images, vector graphics, opacity, rotation, clipping, and paint order.

Visual and semantic HTML

Visual mode recreates the PDF page for viewers and previews. Semantic mode turns the same content into headings, paragraphs, lists, and tables that reflow like a web page.

Live comparison

The PDF on the left. Reconstructed HTML on the right.

Choose a document, including a 100 MiB, 1,000-page PDF. The left side uses your browser's PDF viewer. The right side reads page one and turns it into HTML. The numbers below are measured live in your browser.

Browser PDF viewercomplete · scrollable
boxpdf visual HTMLpage 1 only · DOM + SVG
—open
—first page → HTML
—bytes transferred
—range requests
—peak reader cache

Loading the live demo…

Large PDFs

Open any page directly.

Opening the 1,000-page example reads the PDF's page index and page one. The other 999 pages stay on the server until your app asks for them.

openPdf()~1–6 ms
first page extracted~3–20 ms
pages 2–1,000not loaded yet

Local Node 24 benchmark, seven warm/cold runs, synthetic 1,000-page PDF. Hardware and document complexity vary.

10,000jump directly to a page

The large-file test opens page 10,000 while keeping the same memory limits.

Measured process memory

Memory use stays steady as the PDF grows.

Each tool reads the same test document in its own Node 24 process. PDF.js and unpdf receive the entire file in memory; @boxpdf/reader reads only the ranges it needs.

140 KiBsource data read from a 100 MiB input
77 KiBArrayBuffer growth
3.64 MiBRSS growth at both 10 and 100 MiB

The process-memory numbers include about 46 MiB used by Node itself. Results vary by computer and software version. Run the benchmark yourself.

Two HTML output modes

visual

Looks like the PDF

Text and graphics stay in their original positions. Use visual mode for viewers, previews, and page-by-page PDF-to-HTML conversion.

semantic

Reads like a web page

Headings, paragraphs, lists, cards, addresses, summaries, and tables follow the normal document flow and adapt to the available width. Explore semantic HTML →

Install the reader and HTML writer

npm install @boxpdf/reader @boxpdf/html-writer

Open PDFs from several sources

openPdf(httpSource(url))
openPdf(blobSource(file))
openPdf(memorySource(bytes))

HTTP range requests, browser files, in-memory bytes, or your own size + read(offset, length) source.

Stream HTML pages

for (let i = 0; i < await pdf.getPageCount(); i++) {
  await write(pageToHtml(await pdf.getPage(i), {
    profile: "visual"
  }));
}

Read, convert, and write one page at a time.