Parsing Bank Statement PDFs with Python
Building a bank statement parser in Python: the libraries, the extraction strategies that actually work, and validating the output.
If you are building this yourself, here is an honest account of what the problem actually involves — it is considerably more than extracting text.
The libraries
| Library | Good for | Limitation |
|---|---|---|
| pdfplumber | Text with coordinates, table detection, line detection | Slower on large files |
| PyMuPDF (fitz) | Fast text extraction, rendering to images | Table detection is weaker |
| pypdf | Metadata, page manipulation, decryption | Text extraction loses layout |
| camelot / tabula | Ruled tables | Needs actual ruling lines, which most banks omit |
pdfplumber is usually the right starting point because it exposes both the text with coordinates and the drawn lines, which is what you need when a statement has partial ruling.
Why table detection is not enough
The obvious approach — call `page.extract_table()` and use the result — works on a minority of statements. Most banks do not draw a full grid, so there is nothing to detect.
What works more generally is clustering text by vertical position to find rows, then by horizontal position to find columns. Two fragments at the same y-coordinate belong to the same transaction; the x-coordinate determines which column.
The rules that matter more than the extraction
- A row with no parseable date is a continuation line — append it to the previous row
- Ten or more digits with no decimal separator is a reference, never an amount
- Rows starting with total, balance or summary keywords are not transactions
- Infer the date convention from the whole file, not row by row
- Test both decimal conventions against the printed closing balance
Validate, always
The single most valuable thing you can build is the reconciliation check. Sum the debits and credits, walk the balance chain, compare against the printed totals, and refuse to return a result you cannot verify.
Without it, a parser failure is indistinguishable from a quiet month. There is more on the specific checks in balance-chain reconciliation.
A minimal working structure
If you are starting, the shape that scales is roughly:
- An extractor that returns text fragments with coordinates, per page
- A row builder that clusters fragments by vertical position
- A column mapper that assigns fragments to fields by horizontal position and content type
- A cleaner applying the continuation-line, reference and summary-row rules
- A validator that walks the balance chain and compares against printed totals
- A per-format cache so a layout you have already solved is not re-derived every time
The validator is the component to write first, not last. Without it you cannot tell whether a change to the extractor made things better or worse — which means you are tuning blind.
Frequently asked questions
How long does this take to build?
A parser for one bank's current format is a weekend. A parser that handles many banks, format changes, scans and multi-currency is an ongoing project — which is the honest reason services like this exist.
What should I build first?
The validator. Without an automated pass/fail signal you cannot tell whether a parser change improved anything, which makes iteration guesswork.
What about using an LLM to parse?
It works well for structure discovery on unfamiliar layouts, and it still needs the same arithmetic validation afterwards. A model that hallucinates a plausible amount is exactly the failure the balance chain catches.
Related posts
- Analysing a Bank Statement in Excel with Pivot Tables
- Bank Statement Formats: MT940, BAI2, CAMT.053, OFX and QIF
- How to Get a Bank Statement into Google Sheets
Convert a statement now
Bank Statement PDF to JSON, or see every format. More in the guides.