I tried pdf-inspector on a 947-page clinical listing to see if it could replace my current PDF table extraction pipeline.
For my use case, readable text is not enough. I need to know exactly which table cell a value came from: page, table, row, column, and the exact bounding box. In regulated documents, that matters because a citation should highlight the actual number supporting a claim.
My current extractor uses pdfplumber/pdfminer for table structure and PyMuPDF for text and geometry. After some optimisation, the full 947-page extraction dropped from ~55 minutes to 6.2 minutes, while keeping cell geometry and row context.
pdf-inspector had one very impressive result: classify_pdf processed all 947 pages in 0.65 seconds and correctly identified the document as text-based.
But Markdown extraction was a different story. Two pages took ~275 seconds. Based on that result, a simple projection for 947 pages would be around 36 hours.
More importantly, Markdown does not preserve the cell-level provenance I need. I also saw separate table cells getting merged into one Markdown cell.
So I would definitely consider pdf-inspector for PDF classification, OCR routing, and document inspection. But for evidence-grade clinical table extraction, I am staying with the current pipeline.
Maybe in the future, will test Per-cell geometry and much faster page extraction.