Looks like the PDF
Text, images, and graphics stay in their original positions on each page.
- Viewers and previews
- Page-by-page conversion
- Side-by-side comparison with the PDF
Semantic PDF extraction
A PDF knows where text belongs on a page, but not whether it is a heading, paragraph, label, or table row. @boxpdf/reader uses the layout to rebuild that document structure as HTML.
Live semantic comparison
Choose a document. Your browser displays the PDF on the left while boxpdf turns the same content into headings, paragraphs, lists, cards, addresses, totals, and tables on the right.
Loading the semantic comparison…
Two ways to use the same PDF
Visual mode keeps everything where it appeared on the page. Semantic mode uses those positions to work out the reading order and create HTML that flows from top to bottom.
Text, images, and graphics stay in their original positions on each page.
Text follows the reading order and becomes headings, paragraphs, lists, sections, and tables. Repeated page headers and footers are removed.
How it works
The HTML keeps links to the original text and page coordinates, so an app can highlight the source or open the matching place in the PDF.
Turn fixed PDF pages into headings, paragraphs, lists, and tables that adapt to the screen width.
Work with paragraphs, fields, and tables while keeping a link back to the source page.
Search paragraphs and sections instead of disconnected pieces of text. Every result still points to its page and position.
Keep rows, totals, addresses, labels, and values together instead of flattening them into a block of text.
Semantic HTML
import { blobSource, openPdf } from "@boxpdf/reader";
import { writeHtmlDocument } from "@boxpdf/html-writer";
const pdf = await openPdf(await blobSource(file));
await writeHtmlDocument(pdf.pages(), write, {
profile: "semantic",
semanticLookaheadPages: 4
});
This example examines four pages at a time so it can join tables and sections that continue across page breaks.