Parsing Bank Statement PDFs with Python

Building a bank statement parser in Python: the libraries, the extraction strategies that actually work, and validating the output.

If you are building this yourself, here is an honest account of what the problem actually involves — it is considerably more than extracting text.

The libraries

LibraryGood forLimitation
pdfplumberText with coordinates, table detection, line detectionSlower on large files
PyMuPDF (fitz)Fast text extraction, rendering to imagesTable detection is weaker
pypdfMetadata, page manipulation, decryptionText extraction loses layout
camelot / tabulaRuled tablesNeeds actual ruling lines, which most banks omit

pdfplumber is usually the right starting point because it exposes both the text with coordinates and the drawn lines, which is what you need when a statement has partial ruling.

Why table detection is not enough

The obvious approach — call `page.extract_table()` and use the result — works on a minority of statements. Most banks do not draw a full grid, so there is nothing to detect.

What works more generally is clustering text by vertical position to find rows, then by horizontal position to find columns. Two fragments at the same y-coordinate belong to the same transaction; the x-coordinate determines which column.

The rules that matter more than the extraction

Validate, always

The single most valuable thing you can build is the reconciliation check. Sum the debits and credits, walk the balance chain, compare against the printed totals, and refuse to return a result you cannot verify.

Without it, a parser failure is indistinguishable from a quiet month. There is more on the specific checks in balance-chain reconciliation.

A minimal working structure

If you are starting, the shape that scales is roughly:

The validator is the component to write first, not last. Without it you cannot tell whether a change to the extractor made things better or worse — which means you are tuning blind.

Frequently asked questions

How long does this take to build?

A parser for one bank's current format is a weekend. A parser that handles many banks, format changes, scans and multi-currency is an ongoing project — which is the honest reason services like this exist.

What should I build first?

The validator. Without an automated pass/fail signal you cannot tell whether a parser change improved anything, which makes iteration guesswork.

What about using an LLM to parse?

It works well for structure discovery on unfamiliar layouts, and it still needs the same arithmetic validation afterwards. A model that hallucinates a plausible amount is exactly the failure the balance chain catches.

Related posts

Convert a statement now

Bank Statement PDF to JSON, or see every format. More in the guides.