🤔How to extract tables from PDF? Try Camelot!
The open-source Camelot library helps extract tables from PDF files. Before installing it, you need to install the Tkinter and Ghostscript libraries. You can install these libraries through the pip or conda package managers:
pip install camelot-py
conda install -c conda-forge camelot-py
Next, you need to, as usual, import the desired module from the library in order to use its methods:
import Camelot
tables = camelot.read_pdf('foo.pdf', pages='1', flavor='lattice')
The flavor parameter is set to lattice by default, but can be reconfigured to stream. The lattice is more deterministic and it is great for analyzing tables where there are demarcation lines between cells. This allows you to automatically parse multiple tables present on the page. The lattice converts the PDF page to an image using the ghostscript library and then processes it to produce horizontal and vertical line segments by applying a set of morphological transformations using OpenCV.
To extract a table from a PDF, use the export() method to print it as a dataframe or export it to a CSV file:
tables.export('foo.csv', f='csv', compress=True)
tables[0].to_csv('foo.csv') # to a csv file
print(tables[0].df) # to a df
https://camelot-py.readthedocs.io/en/master/user/install.html
Post #494
531
- 👍 6