hn.today

Lightweight PDF parser with layout, tables, formulas and bounding boxes

github.com86 points7 comments
Screenshot of Lightweight PDF parser with layout, tables, formulas and bounding boxes

A lightweight PDF parsing toolkit that reconstructs document structure - columns, reading order, tables, formulas, figures and the precise bounding box of every block - without heavy ML models. It runs on CPU in Python, in the browser, or as an API, and exports to JSON, Markdown, HTML, DOCX (in the browser), CSV/XLSX for tables, PNG crops for figures and formulas, and full per-block metadata. Tables (including ruled, borderless and LaTeX booktabs) are returned as rows and columns, formulas are converted to LaTeX plus a cropped image, images and charts are cropped with captions preserved, and text styling (alignment, indents, fonts, bold runs) is retained so Word exports resemble the original. OCR via Tesseract and non-PDF formats are handled through Apache Tika.

Two engines run in parallel: a PDFium-based layout engine that extracts glyph positions and rebuilds structure with a column-aware XY-cut, and Tika for metadata, OCR and alternate formats; a JavaScript port powers the browser demo. Benchmarks show competitive speed (structured mode ~543 ms per document, fast mode ~136 ms), robust handling of dense multi-column papers with zero failures on a 54-paper test set, and median 39 ms per page. Known limits: highly complex stacked math is linearized (an image is provided), narrow borderless tables can be misread, scanned PDFs require OCR on the server path, and some exports are browser-only. The project is MIT-licensed and available with CLI, Python API, Docker and a web UI.

Read on github.com7 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in Other

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.