Mining PDFs for tabular data, part 1
Recently, I've had several use cases that lead me to use the excellent python library pdfplumber.
Quick note regarding pdfs with a very high number of pages.
By using python for document\textual analysis, you can take advantage of the language's expressiveness, conciseness and flexibility.
Extracting tables is not always straightforward...
While pdfplumber has functions for extracting tables, I found that it did not work with the default settings for all but the simplest of examples.
The library does allow you to tweak all manner of table-extraction parameters, and recommends cropping pages before you begin extracting.
... but it doesn't have to be complicated.
However, I haven't yet found a need to get buried in the details of table identification, but have used the following rules of thumb in extracting financial data from bank and other statements:
If the structure of the documents and the rows and columns is predictable and simple enough, just use regex and string manipulation.
pdfplumber's extract_text function does a pretty good job of representing many documents as a series of lines which you can then parse. To get the entire document's text in a string, do something like:
pdf = pdfplumber.open(file_path)
text = "".join(p.extract_text() for p in pdf.pages)
I have written up a recent example of a regex-based approach in part 2.
If you cannot deduce which field a piece of text belongs to from its context alone, you will have deal with the object model which gives you information about the horizontal and vertical position of every word.
But for tables with just a few well-defined columns, I still think this can be simpler and more reliable than using a dynamic, parameter-tweaked approach to table-identification.
For more information on this technique, see part 3.
For financial statements, don't make the mistake of assuming that every transaction will occupy exactly one line.
That might be true for one statement pdf, but you may later encounter another in which they wrap across several lines. In my experience, the date and amount columns were never the cause of this, always the description which sometimes gets wordy.
Wrapped lines makes delineating one transaction from the next a little trickier than just text.split('\n').
