Cash flow underwriting header image

Why Cash Flow Underwriting Is Finally Catching On

Ask any credit officer what they look at when they evaluate a borrower, and you will hear some version of the same two ideas. Does this person have a history of paying back what they borrow. And do they have the income and financial capacity to repay right now. Willingness to pay and ability to pay. Those two pillars have anchored lending for as long as lending has existed, and no one in the industry would argue with either of them in principle.

What is strange, once you sit with it, is how lopsided the infrastructure built around those two pillars turned out to be. I have spent a lot of time in underwriting departments over the years, and I have come to appreciate just how much of that imbalance was never really a decision about what mattered. It was a decision about what was possible.

Fifty years of investment went to one side of the ledger

For the better part of five decades, the credit industry poured enormous resources into extracting every possible signal out of willingness-to-pay data. Credit bureaus built national infrastructure around it. FICO and VantageScore refined it into precise, standardized scores. An entire analytics industry grew up around inferring a borrower’s financial stress from things like the number of recent credit inquiries or the age of their oldest account. It is genuinely sophisticated work, built and rebuilt over generations.

The ability-to-pay side of the equation received almost none of that same investment. Cash flow data, meaning the actual pattern of income and expenses moving through a person’s or a business’s accounts, sat mostly untouched by comparison. That is not because it lacked predictive value. Research on cash-flow-based scoring has found it performs on par with traditional credit scores in forecasting loan performance, with area-under-the-curve accuracy in the same range as VantageScore. The two data sets are also measuring something different from each other. The correlation between a traditional credit score and a cash-flow score sits around 0.22, which tells you they are picking up on largely separate signals rather than restating the same information in a different form.

So if cash flow data has been this valuable the whole time, why did the industry build fifty years of infrastructure around the other half of the picture instead?

The 1970s made a practical choice, not a philosophical one

The answer is almost entirely about what was technically feasible at the time the infrastructure was being designed. Credit bureau data in the 1970s was structured, predictable, and low in volume. It updated once a month. It fit neatly into the batch-processing systems of the era. Cash flow data was the opposite of convenient. A single consumer generates thousands of transactions a year, arriving continuously, in no standard format, at a volume that would have overwhelmed the computing capacity available at the time.

Asking a lender in 1975 to build a scoring infrastructure around raw transaction data would have been a bit like asking an automaker in 1910 to build around electric batteries instead of gasoline. Gasoline was the obvious engineering choice given the technology of the moment, and the industrial infrastructure that got built around that choice, refineries, distribution networks, service stations, became so deeply entrenched that it persisted for a century, long after better alternatives were technically possible. Credit bureau infrastructure followed a similar path. It was the right practical choice in 1975, and it became so deeply embedded in underwriting, servicing, securitization, and regulation that it kept shaping decisions long after the original constraint that justified it had started to loosen.

Larry Rosenberger, who ran FICO for years, has said it about as plainly as it can be said. Cash flow data is gold. The only reason it was not used at the start was because it was not widely available. That is a remarkable admission from the person who built much of the willingness-to-pay infrastructure we still rely on today. It was never that the industry decided credit history mattered more than cash flow. It was that credit history was the only one you could actually build a national scoring system around with 1970s technology.

What changed, and why it matters now

The constraint that shaped fifty years of underwriting infrastructure has largely dissolved. Bank transaction data is now accessible programmatically, with consumer permission, in something close to real time. The analytical tooling to turn that transaction stream into a reliable, standardized risk signal exists and is being used in production by lenders today, not as an experiment but as a functioning underwriting input. And the regulatory environment has stopped being neutral on the question. The Consumer Financial Protection Bureau’s open banking rulemaking is deliberately designed to make consumer-permissioned financial data portable across providers, and the Office of the Comptroller of the Currency has been actively encouraging banks to use deposit account data to qualify borrowers who would otherwise be invisible to traditional scoring.

CFPB Director Rohit Chopra has framed the shift in terms that matter directly to lenders serving underserved populations. Bringing a personal financial ledger to a new provider lets that provider evaluate a borrower’s full financial picture instead of relying on a summary compiled by a credit bureau. Individuals without years of credit history, or those who had a rough patch years ago that still shows up on their file, can be evaluated on what their finances actually look like today rather than on a score shaped by events from years past.

This is not a marginal population. An estimated 26 million American adults are completely invisible to the traditional credit system, with no file at any of the three major bureaus. Another 10 million have files too thin to generate a reliable score. Add in the tens of millions more who are technically scorable but sit on thin, brittle files, and you are talking about a meaningful share of the adult population that traditional underwriting infrastructure was never built to see clearly.

Why this lands hardest at CDFIs and community lenders

I think about this population differently after spending time with CDFIs and mission-driven community lenders, because the credit-invisible and thin-file borrowers are not an edge case for them. They are frequently the core of the portfolio. Recent immigrants who have not yet built a domestic credit file. Young adults who have not had the opportunity to establish one. Small business owners whose personal credit history has nothing to do with the financial health of the business they are actually running day to day.

Traditional credit scoring was never designed with these borrowers in mind, because they simply were not part of the population the infrastructure was built to serve in 1975. Cash flow data changes the question entirely. Instead of asking what a bureau file says about a borrower’s history, it asks what is actually happening in their accounts right now. Is income arriving consistently. Are essential obligations being met. Is there a cash flow pattern that indicates capacity to take on and repay new credit, independent of whether that pattern has ever been translated into a bureau score.

The results from early adopters are hard to dismiss. Under the OCC’s Project REACh initiative, banks piloting deposit-account-based underwriting for first-time credit products had established over 110,000 accounts as of late 2023, with credit-invisible borrowers who were approved through this approach reaching an average FICO score of 680 within twelve months. In other words, borrowers who could not have been approved under a traditional model went on to become traditionally creditworthy, because someone was willing to evaluate their actual cash flow rather than the absence of a file.

What this means operationally, not just philosophically

It is one thing to accept the argument that cash flow data is predictive and underused. It is another thing to actually operationalize it inside a lending organization that has spent years, in some cases decades, building processes around credit-bureau-driven underwriting. This is where I think a lot of the conversation about cash flow underwriting stays too abstract. The data being available through open banking connections does not mean an underwriting team can use it well.

Cash flow data is high volume, unstructured relative to a bureau file, and constantly changing. Making it usable in a live underwriting workflow requires an origination system that can ingest transaction-level data, normalize it into something an underwriter or an automated rule can actually evaluate, and route it through the same approval and exception workflow as every other piece of the credit decision. Bolting a cash flow data feed onto a process built around static bureau pulls tends to produce exactly the kind of disconnected, manually reconciled workaround that most lenders are already trying to eliminate elsewhere in their operations. The lenders who get real value out of cash flow underwriting are the ones treating it as a core input to their loan origination software, not as a side lookup that a credit analyst checks manually when a file looks borderline.

This is also where the fair lending questions get real rather than theoretical. Attributes derived from spending categories, like discretionary purchases or timing of recurring payments, have not been tested extensively under the Equal Credit Opportunity Act and Regulation B. Lenders who build cash flow underwriting into their process need the same rigor and documentation discipline they apply to any other underwriting variable, with a clear rationale for why a given signal is being used and evidence that it is not operating as a proxy for a protected characteristic. That is a governance conversation as much as a data conversation, and it belongs inside the underwriting policy framework, not bolted on afterward.

The infrastructure decision in front of lenders today

What strikes me most about this history is how long a purely practical, technology-driven decision from 1975 continued to shape underwriting policy for fifty years afterward. Nobody sat down in 1975 and decided that cash flow data was less valuable than credit history. They decided that credit history was the only one they could actually build a national infrastructure around at the time. That practical constraint calcified into an assumption, and the assumption outlived the constraint by decades.

We are now at the point where the original constraint no longer holds. The data is accessible. The tooling to interpret it exists. The regulatory signal is pointing toward, not away from, its use. What is left is an infrastructure decision, the same kind of decision the industry made in 1975, except this time lenders get to make it deliberately rather than by default.

For CDFIs and community lenders in particular, this is not just a technical upgrade. It is an opportunity to serve the borrowers their mission is built around more accurately than the inherited infrastructure ever allowed. The lenders who understand this history, and who build the operational capability to evaluate ability to pay with the same rigor the industry has spent fifty years applying to willingness to pay, are the ones who will be positioned to serve those borrowers well. The rest will keep operating on infrastructure that was designed for a data problem that no longer exists.