Open large files
The reader fetches requested byte ranges and keeps them within cache limits set by the application.
Open-source PDF-to-HTML for JavaScript
@boxpdf/reader and @boxpdf/html-writer convert PDFs into visual or reflowable HTML, preserving text, images, fonts, and page layout. They run in Node, edge functions, and the browser.
import { httpSource, openPdf } from "@boxpdf/reader";
import { pageToHtml } from "@boxpdf/html-writer";
const pdf = await openPdf(httpSource(url), {
maxBytes: 2 * 1024 * 1024,
maxObjectCacheBytes: 2 * 1024 * 1024
});
const page = await pdf.getPage(0);
viewer.innerHTML = await pageToHtml(page, { profile: "visual" });
Why @boxpdf/reader
Plain-text extractors lose columns, labels, images, and other parts of the page. Page renderers give you an image but not much structure. boxpdf gives you the content and where it appeared, as JavaScript data and HTML.
The reader fetches requested byte ranges and keeps them within cache limits set by the application.
Use it with an HTTP URL, a browser File, bytes already in memory, or your own data source. It is pure JavaScript, with no native binary or headless browser.
Go directly to page one, page 10,000, or any page in between. The rest of the document is read only when requested.
Visual HTML includes positioned text, embedded fonts, images, vector graphics, opacity, rotation, clipping, and paint order.
Visual mode recreates the PDF page for viewers and previews. Semantic mode turns the same content into headings, paragraphs, lists, and tables that reflow like a web page.
Live comparison
Choose a document, including a 100 MiB, 1,000-page PDF. The left side uses your browser's PDF viewer. The right side reads page one and turns it into HTML. The numbers below are measured live in your browser.
Loading the live demo…
Large PDFs
Opening the 1,000-page example reads the PDF's page index and page one. The other 999 pages stay on the server until your app asks for them.
Local Node 24 benchmark, seven warm/cold runs, synthetic 1,000-page PDF. Hardware and document complexity vary.
The large-file test opens page 10,000 while keeping the same memory limits.
Measured process memory
Each tool reads the same test document in its own Node 24 process. PDF.js and unpdf receive the entire file in memory; @boxpdf/reader reads only the ranges it needs.
The process-memory numbers include about 46 MiB used by Node itself. Results vary by computer and software version. Run the benchmark yourself.
Text and graphics stay in their original positions. Use visual mode for viewers, previews, and page-by-page PDF-to-HTML conversion.
Headings, paragraphs, lists, cards, addresses, summaries, and tables follow the normal document flow and adapt to the available width. Explore semantic HTML →
npm install @boxpdf/reader @boxpdf/html-writer
openPdf(httpSource(url))
openPdf(blobSource(file))
openPdf(memorySource(bytes))HTTP range requests, browser files, in-memory bytes, or your own size + read(offset, length) source.
for (let i = 0; i < await pdf.getPageCount(); i++) {
await write(pageToHtml(await pdf.getPage(i), {
profile: "visual"
}));
}Read, convert, and write one page at a time.