boxpdf

Semantic PDF extraction

Turn PDFs into reflowable HTML.

A PDF knows where text belongs on a page, but not whether it is a heading, paragraph, label, or table row. @boxpdf/reader uses the layout to rebuild that document structure as HTML.

Reading order Paragraphs and lines Table structure Links back to each page Reflowable HTML Multi-page documents

Live semantic comparison

The PDF on the left. Reflowable HTML on the right.

Choose a document. Your browser displays the PDF on the left while boxpdf turns the same content into headings, paragraphs, lists, cards, addresses, totals, and tables on the right.

Visual PDFbrowser renderer · scrollable
Semantic HTMLreflowable document structure
—open
—semantic HTML
—pages structured
—peak reader cache

Loading the semantic comparison…

Two ways to use the same PDF

Keep the page layout, or rebuild the document.

Visual mode keeps everything where it appeared on the page. Semantic mode uses those positions to work out the reading order and create HTML that flows from top to bottom.

visual

Looks like the PDF

Text, images, and graphics stay in their original positions on each page.

  • Viewers and previews
  • Page-by-page conversion
  • Side-by-side comparison with the PDF
semantic

Reads like a web page

Text follows the reading order and becomes headings, paragraphs, lists, sections, and tables. Repeated page headers and footers are removed.

  • Responsive reading
  • Search and indexing
  • Structured data extraction

How it works

How semantic mode builds the document

1Read the PDFtext · fonts · images · graphics
→
2Rebuild each pagepositions · sizes · drawing order
→
3Create the documentheadings · paragraphs · lists · tables

The HTML keeps links to the original text and page coordinates, so an app can highlight the source or open the matching place in the PDF.

What semantic HTML is for

Responsive reading

Turn fixed PDF pages into headings, paragraphs, lists, and tables that adapt to the screen width.

Data extraction

Work with paragraphs, fields, and tables while keeping a link back to the source page.

Search and indexing

Search paragraphs and sections instead of disconnected pieces of text. Every result still points to its page and position.

Tables and forms

Keep rows, totals, addresses, labels, and values together instead of flattening them into a block of text.

Semantic HTML

Create a reflowable document

import { blobSource, openPdf } from "@boxpdf/reader";
import { writeHtmlDocument } from "@boxpdf/html-writer";

const pdf = await openPdf(await blobSource(file));

await writeHtmlDocument(pdf.pages(), write, {
  profile: "semantic",
  semanticLookaheadPages: 4
});

This example examines four pages at a time so it can join tables and sections that continue across page breaks.