Applications for the Rice MBA are open. Round 1 deadline: October 16.

AI | Finance

AI at the Kitchen Table

Financial technology is increasingly making the decisions households once made on their own. New Rice Business research tests three of these tools and finds that better automation does not reliably produce better outcomes.

This story is part of our special anniversary issue of Rice Business Wisdom on artificial intelligence.
In his 2006 American Finance Association presidential address, “Household Finance,” Harvard economist John Y. Campbell partly positioned the field this way: “By analogy with corporate finance, household finance asks how households use financial instruments to attain their objectives.” 

In the last 20 years, the financial instruments being used to help households attain their objectives have evolved drastically. Budgeting apps now decide what gets noticed. Automated systems raise red flags. Chatbots are becoming the go-to source for financial advice. And increasingly, machine learning, algorithms and fintech are shaping decisions that families once worked through on the back of envelopes at the kitchen table. One of the most urgent tasks for household finance researchers is to test those new instruments at scale, and to gauge how well the machines that shape more and more of our financial lives are actually working. 

Image
Line-drawn headshots of Bruce Carlin, Stephanie Johnson, David Zhang, and Benedict Guttman-Kenney

The studies that follow examine this new terrain in three corners of household finance: investing, mortgage lending and credit-card debt. Each takes a single technical intervention and measures what changes when it reaches real people. And each turns up a version of the distance between what technology appears to do and what people experience. 

What happens when you ask LLMs to pick your stocks? 

Ask five of the most powerful AI models to invest your money and beat the stock market, and they mostly agree on one move: buy Nvidia. A lot of Nvidia. 

That’s what Bruce Carlin, the George R. Brown Professor of Finance, found when he set out to study how AI handles an ordinary person’s money. For nearly a year, the Rice Business professor and two colleagues at Michigan State University and the University of California, Berkeley asked the leading large language models (LLMs) the same question almost every morning — what should I buy to beat the market? — and tracked every answer as it played out in real time. 

For most of the last century, having a portfolio built for you meant hiring someone to do it — a broker, an adviser, a manager with a license and a fee. But today the same request costs nothing and takes seconds. You ask, and the LLM answers: here’s what to buy, here’s how much, here’s why. Now that nearly everyone has a free financial adviser in their pocket, the real question is whether it’s any good. That’s what Carlin and his co-authors of a new National Bureau of Economic Research working paper, “AI Managed Household Portfolios: A Preliminary Report,” wanted to find out. 

 

The risk isn’t that the AI gives obviously bad advice,” he says. “It’s that the advice arrives instantly, costs nothing and sounds reasonable, which is exactly what makes an ordinary investor likely to act on it.”

 

Putting the same demand to every system they tested — ChatGPT 5.0 and 5.2, Claude Sonnet 4.5, Gemini 2.5 Flash and Grok 4.1 Fast — Carlin and his co-authors made the LLMs show their work. Each pick had to come with a reason and a source. Why this company? And what were the models reading to get there? The answers to those questions became the real research data. 

It turned out the logic for LLM recommendations had less to do with a firm’s actual figures and more with their media mentions. The companies that made the cut drew nearly 10 times the news coverage of the average public company. What the AI recommended, ultimately, was whatever was being talked about most online. 

The cost of that news-chasing is that you get a narrow portfolio with a handful of headline names, mostly crowded into the same corner of the market. Semiconductors alone made up around 40% of the AI-managed holdings — enough that a single bad year for chipmakers could drag down the whole account. Banks, meanwhile — the backbone of most retirement accounts — barely appeared. Gemini, at one point, treated four or five stocks as a finished portfolio, a job most advisers would spread across dozens of names. “For the household investor, the protection that matters most — diversification, the spreading of risk — is the very thing LLMs left out,” Carlin says. 

On paper, the AI advice looked like a winner. The portfolios beat the S&P 500, the one thing they had been asked to do. But when Carlin applied a harder test, measuring each pick against other stocks of the same size, growth and recent momentum, the edge all but vanished. The LLMs hadn’t picked smarter stocks, only riskier ones. The gains came from taking a concentrated risk that happened to pay off in a rising market but would be considered reckless in a falling one. 

Carlin is careful to call the findings preliminary. The data are still coming in. But his early read is sobering. “The risk isn’t that the AI gives obviously bad advice,” he says. “It’s that the advice arrives instantly, costs nothing and sounds reasonable, which is exactly what makes an ordinary investor likely to act on it.” 

As more of the money decisions families used to work out together get handed to chatbots — what to buy, what to borrow, how to save for retirement — the same questions will increasingly follow from one budgeting conversation to the next: Does AI actually give good financial advice? Or does it only sound that way? 

Has automated underwriting kept its promises? 

Three decades after automated underwriting systems (AUS) began reshaping the mortgage market, Stephanie Johnson and David Zhang are asking whether the technology has delivered on its original promises. Both assistant professors of finance at Rice Business, they study AUS at a moment when the tools are now integral to the basic machinery of mortgage origination, used to evaluate millions of borrowers and determine who can buy a home and for how much. 

Writing in an October 1994 article titled “Mortgages in Minutes,” Peter Maselli, then vice president of automated underwriting at Freddie Mac, captured the promises of a new era: “Artificial intelligence is the key to automating the underwriting process.” With AI and machine learning, he wrote, mortgage decisions could become faster, more consistent and less dependent on the judgment and potential biases of individual underwriters. 

For lenders and real estate agents in the 1990s, the new technology had strong appeal. Human underwriting was slow, not to mention inconsistent and susceptible to prejudice. Promising greater efficiency, accuracy and consistency, Fannie Mae and Freddie Mac helped push the technology into the mainstream, and today roughly two-thirds of new American mortgages are underwritten by a small handful of programs.

 

The question, then, is not simply whether AI and machine learning can make better lending decisions — it’s what kind of borrower those systems are built to see.

 

That long history gives Johnson and Zhang the data to test how well AUS has lived up to its billing. In separate studies, they find a mixed record: efficiency gains have been marginal; accuracy has improved in some important ways; and fairness remains the hardest promise to assess.

  • Efficiency: Johnson’s research finds that adoption of AUS trimmed only about four days from the average processing window, with larger reductions for those that were denied than for those approved. Four days is not nothing, but it’s also not the revolution the industry was selling. 

    Working with a colleague at the Dallas Fed, Johnson found that the real change following AUS implementation was in the lending standards. Freddie Mac’s system introduced statistically derived rules that relaxed traditional limits on how much a household could borrow against its income, allowing borrowers to qualify for larger loans on the same paycheck. 

    Over time, that extra borrowing capacity spread across millions of households, helping drive up both home prices and household debt.
     

  • Accuracy: The most ambitious promise of AUS was that a machine could judge a borrower’s risk more reliably than a human — lending to those who would repay and turn away those who would not. It’s also the promise the evidence most clearly supports. 

    For years there was no way to test it because the riskiest borrowers stayed in the hands of human underwriters. But a 2016 Federal Housing Administration (FHA) policy changed that, letting algorithms approve loans to applicants with low credit scores and high debt. 

    In a forthcoming Journal of Finance article, Zhang and his co-authors, including Rice Business alumna Hanyi Livia Yi (Ph.D. ’21), treat this policy change as a natural experiment. They found that the machines judged well. Lending to these borrowers surged, more than doubling in the most affected group, while the share who fell behind on their payments barely changed. “The algorithms approved people that underwriters had been turning away, and delinquency rates did not rise,” Zhang says. “The AUS lent more freely than humans had, without making worse lending decisions.” 
     

  • Fairness: The most consequential promise of AUS technology, and the hardest to keep, was that a race-blind formula would widen the door for borrowers that a biased system had shut out. 

    The same 2016 policy that improved accuracy offered Zhang a direct test on the question of inequity. The FHA serves a disproportionately Black and lower-income population, exactly the groups the automation was supposed to help. 

    Zhang and his co-authors found the credit expansion reached non-Black borrowers at about 7.3% — but it was essentially zero for Black borrowers. The gap held even after matching borrowers by income and credit score, so it was not simply a question of income in disguise. “The algorithm never sees race,” Zhang says, “but it rewards a clean, documented financial history, and that’s exactly what disadvantaged borrowers are least likely to have.”

Taken together, these findings complicate the founding promises of automated underwriting. AUS did not make mortgage lending dramatically faster, but it did change who could borrow, how much they could borrow and how well lenders could predict repayment. It also shows why fairness is the hardest part to automate. A race-blind system can remove some forms of individual discretion, but it cannot erase the unequal financial histories already present in the data. 

The question, then, is not simply whether AI and machine learning can make better lending decisions — it’s what kind of borrower those systems are built to see.

Can a nudge help people get out of credit-card debt?

According to Chicago Booth Review, credit-card balances in the U.S. topped $1 trillion in 2023, much of it carried by people who pay only the monthly minimum at interest rates north of 20%. For years, behavioral economists thought they had an inexpensive way to help consumers reduce their debt: a “nudge” to hide the minimum payment option. The thinking — based on the psychology theory of anchoring whereby the salience of seemingly irrelevant numbers can distort choices — was that hiding the minimum payment number would encourage borrowers to pay their debt more quickly. 

Two recent big experiments by Benedict Guttman-Kenney undermine this idea. Guttman-Kenney, an assistant professor of finance at Rice Business, has tested the nudge at a scale most behavioral research never reaches: a British trial of more than 40,000 credit card holders, and a second study of nearly 7,000 cardholders in the United States. Both turned on the same kind of intervention fintech apps deploy constantly to shape how households handle their money. 

 

Put the same design into a real app, attached to real money, and the effect disappeared. 

 

Paying only the minimum required amount on a credit card keeps a balance alive for years, each month’s interest compounding on the last. In Britain, the one in five cardholders enrolled in minimum-only autopay generate nearly half of all the interest and fees the industry collects. And three-quarters of those in long-term, persistent debt pay this way. These are the households the British trial was built to rescue.

In that trial, the nudge meant removing the “pay the minimum” option from the autopay menu that new cardholders saw at sign-up. By every immediate measure, it worked. Enrollment in minimum-only autopay fell from 37% to under 10%, and the share of borrowers paying exactly the minimum dropped by 23%. Yet six months later, their credit-card debt had not moved — and neither had their spending, payments or borrowing costs. 

The reason, Guttman-Kenney found, was that minimum payments were not simply a psychological anchor pulling borrowers toward a lower number. For many cardholders, minimum payments were close to the limit of what they could afford. Half had no readily available cash or had been overdrawn at some point in the three months before applying for the card. 

A second experiment, run with Arro, a U.S. fintech credit card lender, emphasized the point. Among nearly 7,000 borrowers, most with below-prime credit scores, Guttman-Kenney tested two versions of the intervention: one that hid the minimum payment and another that added a button to pay half the balance. Neither changed how much people paid. That result ran against earlier online experiments, where hiding the minimum payment pushed people to say they would pay about 12 percentage points more. But those were hypothetical choices. Put the same design into a real app, attached to real money, and the effect disappeared. 

Together, Guttman-Kenney’s two studies suggest that the minimum payment was doing something more practical than behavioral theory had assumed. Rather than pulling borrowers toward a lower number, cash-strapped households used it as a benchmark they could work toward. 

Like the studies on AI-managed portfolios and automated underwriting, the credit-card nudge research tests whether a financial technology improves the outcome it was built to address, not just whether it works as designed. The difference is that here, the technology worked exactly as designed. It changed behavior on contact, but even so, debt, spending and borrowing costs didn’t move. Of the three studies, it’s the clearest proof that technologies can perform as designed and still not improve a household’s finances.

The limits of smarter systems 

Taken together, the findings raise a central question: Can the financial instruments shaping household decisions help families attain their objectives, given the money, time, risk and documentation available to them? It’s not an easy standard to design for. A model can only optimize what it can see, after all, and what it sees is often a thin proxy for how families actually live.

The next challenge is to identify where those proxies break down, which households they misrepresent and how to account for their limits. As more and more decisions move from the kitchen table to the user interface, designers and researchers will need to build and test financial technologies against household conditions that are often difficult to observe and slow to change.

Written by Scott Pett