A commercial API that converts PDFs and images into validated, structured JSON with a per-field accuracy service-level agreement and a pay-only-for-correct-pages pricing model. It targets common document types (invoices, payslips, bank statements, closing disclosures, tax forms and custom schemas) and industries (accounts payable, mortgage, insurance, healthcare, payroll, logistics). Security and compliance are emphasized (SOC 2, GDPR, ISO 27001, EU/US, on‑prem options) alongside a zero-retention policy. The offering includes free trials, webhook delivery, Zapier integration and performance reporting so teams can measure accuracy on their own documents and contract for per-field SLAs that refund or waive charges for incorrect extractions.
The technical pipeline is described end-to-end: ingestion and pre-processing (deskew, denoise), dual-pass OCR to capture text and layout, page splitting and classification, multi-step extraction that reconstructs tables and line items, entity normalization (dates, currencies, tax codes), schema mapping, cross-field validation and provenance linking each extracted value back to page/region. Outputs in examples show richly structured JSON with invoice numbers, seller/buyer details, line item breakdowns and tax amounts, closing disclosure loan terms and payslip totals. Continuous learning, confidence scoring, agentic review and iterative model upgrades are used to improve and monitor accuracy in production.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.