Why we stopped building our own document parser

Eight weeks into building LEA, our lead engineer told me the document parser was ready. I asked him to run it on a batch of real client statements.

It failed on 40% of them.

Not dramatically — it didn’t crash or throw errors. It just quietly returned wrong numbers. Account balances off by a few cents. Rows dropped from multi-page statements. Tax schedule line numbers that looked right but weren’t. The kind of failures you’d only catch if you knew what to look for.

That day started a two-month detour that taught me more about building for financial services than anything else we’ve done.

The problem we underestimated

When we started LEA, I assumed document parsing was a solved problem. OCR has existed for decades. Dozens of companies have built on top of it. How hard could reading a PDF be?

Very hard, it turns out — for wealth management specifically.

RIA documents aren’t clean PDFs with simple text. They’re multi-column custodian statements with tables that span pages. Tax forms whose structure changes by preparer and year. Transfer paperwork that’s half handwritten and half typed. Account applications mixing structured fields with free-text disclosures buried in footnotes.

Generic OCR tools handle clean, single-column text well. They fail on everything else — and in wealth management, everything else is every third document.

Why “good enough” isn’t good enough here

The failure rate that matters isn’t the one you measure in testing. It’s the one that shows up in production, on documents you’ve never seen before.

Our parser handled about 60% of document types reliably. The other 40% required manual review — which meant someone on the ops team had to catch what the parser missed. That’s fine as a temporary patch. It’s not a product.

But the deeper issue is what happens when the parser fails silently. If 5% of account balances are extracted incorrectly, the ops team has to catch them — but only if someone checks every output. If 5% of transfer amounts are wrong, you have a compliance exposure. If 5% of account numbers are off, data flows silently into the wrong record.

The ops teams we work with don’t have time to spot-check every extraction. They need to trust the output completely. That meant our 60% reliable parser wasn’t a starting point we could gradually improve — it was a fundamental architecture problem.

What we tried to do about it

We spent two months pushing the parser forward. Every new document type required engineering time to handle. We added heuristics for merged cells. We wrote special cases for multi-page statements. We built a fallback that flagged low-confidence extractions for human review.

The backlog of document types we couldn’t handle reliably kept growing faster than we could clear it.

Our estimate to reach 90% reliability: another four to six months of engineering. And we still weren’t confident we’d crack the hardest cases — the ones that required understanding document structure at a level that’s genuinely difficult to encode in rules.

At some point I had to ask the question: is this the problem we’re trying to be the best in the world at?

The answer was no. The problem we’re trying to be the best in the world at is wealth management operations — the workflows, the integrations, the compliance logic, the system-of-record updates. The document parser is the input to all of that. It’s not the product.

The decision

We evaluated document intelligence APIs. We landed on Reducto.

What stood out wasn’t just their headline accuracy numbers. It was how they handled the specific failure modes we’d been fighting — merged cells, multi-page tables, mixed handwritten-and-typed documents. They’d clearly built for production use cases, not benchmarks. They process three billion pages in production. The edge cases we were discovering, they had already seen and solved.

We integrated Reducto as our extraction layer. Everything downstream — the wealth management-specific schema, the validation logic, the workflow routing, the custodian integrations — stayed ours.

Extraction accuracy went from roughly 60% to 98.7% across the document types we process. The parser backlog stopped growing. The engineering team shifted from maintaining a document parser to building the workflow automation that the parser feeds.

What I’d do differently

Do it sooner. We spent two months proving to ourselves that document parsing is genuinely hard before evaluating APIs that had already solved it. That was two months of engineering time we could have put into the layers that are actually our product.

The lesson I carry: infrastructure is only worth building in-house if it’s your moat. For LEA, the moat is the wealth management workflow layer — the logic that understands what to do with the data after it’s extracted. Not the extraction itself.

There’s a version of every fintech company where the team spends a year building plumbing that a focused API provider has already solved better. The hardest thing in early-stage product development isn’t deciding what to build — it’s deciding what not to build, even when building it feels like forward progress.

If you’re building on top of financial documents and fighting your own version of this, I’d be glad to compare notes. My DMs are open.