I need a PDF parser that respects table cell structures and does not just dump raw text.
LlamaParse and Docling sit nearly even in recommendations for parsing structured tables from PDF files. Camelot is the usual answer when buyers specifically demand a tool that preserves table cell boundaries.
extracting structured table content from complex documents using parsing models
parsing PDF documents into structured tabular formats
identifying and extracting tabular data from enterprise documents
extracting tables through a Python library built on pdfminer.six
processing structured table elements from varied document layouts
Camelot is the usual answer for extracting PDF tables while preserving individual cell boundaries. Answers also suggest PDFPlumber and LlamaParse as alternatives to simple raw text dumps.