Incentives Matter

The Hubris of Choice: Can We Actually Diagnose Bad Decisions?

March 31, 202617:59Incentives Matter

This episode critically examines the foundational assumptions of behavioral economics, specifically how experts identify "bad" or "improvable" decisions. It highlights a new NBER paper which reveals that the three primary diagnostic methods used by behavioral economists to identify mistakes are inconsistent and contradict each other, challenging the validity of widespread "nudge" policies and choice architecture. Listeners will learn how this fundamental flaw undermines the rationale for paternalistic interventions, with the first diagnostic method, the Characterization Assessment, being introduced.

Key Takeaways

Detailed Report

The Shaky Foundations of 'Bad Decisions'

For decades, behavioral economics has influenced public policy and product design, operating on the premise that experts can identify human cognitive biases and design environments to 'nudge' people toward better choices. However, groundbreaking new research challenges this very foundation, revealing that the primary methods used to diagnose a 'bad' decision fundamentally contradict each other.

Unmasking the Hubris of Choice Architects

The core idea behind 'nudging'—popularized by Richard Thaler and Cass Sunstein's book *Nudge* and implemented by government 'nudge units' worldwide—is that people make irrational choices due to cognitive biases (like procrastination, loss aversion, or present bias). Experts then design 'choice architecture' to gently steer individuals toward 'optimal' decisions. This approach is seen in auto-enrollment for 401(k)s, opt-out organ donation systems, and even diet-tracking apps or auto-renewing subscriptions, all justified by the belief that experts know what's best for the individual, sometimes even better than the individual themselves.

This implicit claim represents a significant leap of faith. But what if the foundational assumption—that we can actually *diagnose* a bad decision—is itself flawed? This is the central question tackled by B. Douglas Bernheim, Aldo Lucia, Kirby Nielsen, and Charles D. Sprenger in their NBER working paper, "When Are Decisions Improvable? An Evaluation of Diagnostic Methods," published in *Nature Human Behaviour*.

The Three Conflicting Diagnostic Methods

The paper identifies three main ways behavioral economists have historically tried to determine if someone has made an "improvable" decision:

1. Characterization Assessment Method

This method flags a choice as a mistake if the decision-maker holds specific, factual misconceptions about it. Researchers ask subjects to describe the parameters of their decision; if they cannot accurately calculate expected value, misunderstand probabilities, or fail a basic comprehension test, their decision is deemed 'improvable.' For example, an employee choosing a high-deductible health plan but unable to explain what a deductible is might be seen as having made a mistake. The critique here is that an inability to mathematically articulate a rational choice doesn't necessarily mean the choice itself is irrational; it might simply mean the individual isn't an economist.

2. Decision Confidence Method

Also known as measuring "Cognitive Uncertainty" (CU), this method gauges how confident a person is in the choice they just made. If a subject reports low confidence—say, 50%—their decision is flagged as suboptimal. The assumption is that if you're not confident, you probably made a mistake. An example would be a retail investor buying a volatile stock and admitting they're just guessing if it was the optimal financial move. Critics argue this relies heavily on self-reporting and assumes subjective doubt perfectly correlates with objective error, ignoring inherent uncertainties in the world or individual cautiousness.

3. Pattern Matching Method

This method identifies a behavior as 'improvable' if it mirrors a specific behavioral pattern known to be mathematically and objectively suboptimal in another domain, often violating Expected Utility Theory. For instance, if someone exhibits extreme risk aversion in a small-stakes game, or buys an extended warranty on a cheap item despite terrible expected value, this method would flag it as a mistake because the math doesn't add up, aligning with a known cognitive bias.

The Shocking Contradiction

The Bernheim paper's critical finding emerged when researchers applied all three diagnostic methods to the *same set of decisions* made by participants in an experiment involving binary lottery choices. The results were stark: the three methods completely contradicted each other, implying that entirely different choices were the 'improvable' ones. One method might deem a safe option a mistake due to misunderstanding, while another might flag a risky option due to low confidence, and a third might point to a completely different choice because it violated a known economic pattern.

This means that whether a human being's decision is labeled 'rational' or 'irrational' depends entirely on which academic tool a researcher chooses to use. The diagnostic tools, in essence, argue among themselves over what constitutes a 'wrong' decision, exposing a fatal flaw in distinguishing between a cognitive error and a valid, legitimate appetite for risk or preference.

The Academic Skirmish and Its Implications

The paper's findings sparked immediate academic debate. Weeks before its official publication, Benjamin Enke and Thomas Graeber published a rebuttal, specifically pushing back on the interpretation of the Decision Confidence method. They argued that subjects aren't necessarily conflating 'best decision' with 'best outcome,' but rather that the very presence of *outcome uncertainty* makes the choice subjectively difficult. Their 'interpretational middle ground' suggests that low confidence might reflect imperfect information or the inherent difficulty of decision-making in uncertain environments, rather than a diagnosable 'mistake.'

Regardless of the nuances in interpretation, the core conclusion remains: the field struggles to objectively identify genuine human error. Researchers, often trained to maximize Expected Value, may project their own risk aversion onto the public, labeling any deviation as an 'error,' even if it reflects a legitimate, different utility function (e.g., valuing peace of mind over maximizing returns).

This research calls for a reevaluation of decades of behavioral insights and the policies built upon them, as they may be founded on flawed diagnoses of human error.

Navigating a Nudged World: Key Takeaways for the Reader

This scientific reckoning has massive implications for individuals living in a 'nudged' world:

  • Trust Your Preferences: What an economist labels as a 'cognitive bias' or 'error' might actually be your perfectly legitimate risk preference. If you prefer a safe, low-yield savings account because it helps you sleep at night, that's a valid choice, regardless of what 'Pattern Matching' might suggest.
  • Question the Defaults: Recognize that every default option you face—from your 401(k) contribution rate to your data privacy settings—was designed by someone who assumed they knew what was best for you. This paper proves they often don't. These defaults are powerful, but they are not neutral or objectively superior.
  • Demand Better Science: Policymakers need to clarify the assumptions underlying their behavioral interventions. Until the field can establish a unified, reliable way to objectively diagnose a 'bad' decision, institutions should default to maximizing freedom of choice rather than paternalistic nudging. The 'God Complex' of behavioral economics is being dismantled from the inside out, which is a positive development for individual autonomy.

This research should make us extremely skeptical of 'nudge-washing,' where corporations or policymakers use the veneer of behavioral science to justify paternalistic interventions that primarily serve their institutional bottom lines rather than the individual's true preferences. If the scientific justification for identifying a 'mistake' is built on such shaky ground, then claims of 'optimizing' individual behaviors through nudges warrant rigorous scrutiny.

Show Notes

Works Referenced

  • When Are Decisions Improvable? An Evaluation of Diagnostic Methods: Groundbreaking research by B. Douglas Bernheim, Aldo Lucia, Kirby Nielsen, and Charles D. Sprenger that critically examines and finds contradictions among the primary methods behavioral economists use to diagnose 'bad' decisions.
  • Nudge: Improving Decisions About Health, Wealth, and Happiness: An influential book by Richard Thaler and Cass Sunstein that popularized the concept of 'nudge' and choice architecture, advocating for gentle interventions to guide people toward better choices.
  • Pension Protection Act of 2006: A U.S. federal law that, among other provisions, significantly impacted retirement savings by facilitating the widespread adoption of automatic enrollment in 401(k) plans.
  • Robinhood: A financial services company known for its commission-free trading platform, which has become popular among retail investors.
  • A Note on 'When Are Decisions Improvable? An Evaluation of Diagnostic Methods': A working paper by Benjamin Enke and Thomas Graeber that provides a rebuttal and alternative interpretation of the findings in the Bernheim et al. paper, particularly regarding the Decision Confidence method.
  • National Science Foundation (NSF): An independent agency of the United States government that supports fundamental research and education in non-medical fields of science and engineering, including funding for the Bernheim et al. paper.
  • Behavioral Insights Team (BIT): A social purpose company, originally part of the UK government, that applies behavioral science to improve public services and policy outcomes.
  • Office of Evaluation Sciences (OES): A team within the U.S. General Services Administration that helps federal agencies use evidence and rigorous evaluation to improve programs and policies.

Glossary

  • Nudge: A concept from behavioral economics referring to subtle interventions in choice architecture that guide people toward certain decisions without restricting their freedom of choice.
  • Behavioral Economics: A field that integrates insights from psychology and economics to understand how psychological factors influence human economic decision-making.
  • Choice Architecture: The design of how choices are presented to individuals, and the impact of that presentation on their decision-making.
  • Cognitive Biases: Systematic patterns of deviation from rationality in judgment, often leading to decisions that are not objectively optimal.
  • Loss Averse: A cognitive bias where the psychological impact of a loss is felt more intensely than the pleasure of an equivalent gain.
  • Present Bias: A cognitive bias where individuals tend to overvalue immediate rewards and undervalue future rewards, often leading to procrastination or impatience.
  • 401(k): A retirement savings and investment plan offered by many U.S. employers, allowing employees to contribute a portion of their salary before taxes.
  • Opt-out organ donation: A system where individuals are presumed to consent to organ donation unless they explicitly register their refusal.
  • Characterization Assessment Method: A diagnostic method that identifies a decision as 'improvable' if the decision-maker holds specific factual misconceptions about the choice.
  • Decision Confidence Method: A diagnostic method that flags a decision as suboptimal if the decision-maker expresses low confidence in having made the best choice for themselves; also known as Cognitive Uncertainty (CU).
  • Pattern Matching Method: A diagnostic method that identifies a decision as 'improvable' if it mirrors a specific behavioral pattern known to be mathematically suboptimal, often violating Expected Utility Theory.
  • Expected Utility Theory: A theory in economics and decision science that describes how rational individuals should make decisions when faced with uncertain outcomes, aiming to maximize their expected utility.
  • Common Ratio Effect: A phenomenon in decision-making where preferences between gambles change when the probabilities of winning are scaled by a common ratio, violating Expected Utility Theory.
  • Small-stakes risk aversion: A behavioral pattern where individuals exhibit an aversion to risk even for very small potential losses, which can appear irrational when scaled up.
  • Nudge-washing: The practice of using the appearance of behavioral science to justify paternalistic interventions that primarily serve institutional interests rather than the individual's true preferences.
  • Dark Pattern: A user interface design that intentionally misleads or tricks users into doing things they might not want to do, such as signing up for unwanted subscriptions or sharing more data than intended.

Sources / References

Full Transcript

HostI have to say, when I first read the abstract for this NBER paper, my jaw dropped. We've spent decades talking about how to `nudge` people to make better choices, but this new research asks a much more fundamental question: how do we even know what a "bad" decision *is*?
ExpertRight? It's like behavioral economics has been operating with a kind of God complex, assuming it can diagnose human error, then prescribe solutions. But what if the very tools we use to make that diagnosis are completely at odds with each other?
HostAnd that's exactly what this paper finds. The three primary methods behavioral economists use to identify what they call "improvable" decisions—you know, mistakes—they just flat-out contradict one another.
ExpertThey flag entirely different choices as errors. It’s like three doctors examining the same patient and giving three completely different, mutually exclusive diagnoses. It's wild.
HostIt really is. And it makes you wonder, if the foundations are this shaky, what does that mean for all the nudges and choice architecture we've built on top of them?
ExpertExactly. For the last couple of decades, behavioral economics has been riding high, influencing everything from public policy to product design. You look at Richard Thaler and Cass Sunstein’s book *Nudge*, and then all the "nudge units" that popped up in governments worldwide. The core idea was always: people have cognitive biases, they make irrational choices, and we—the experts—can design environments to gently steer them towards optimal decisions.
HostAnd it seemed so elegant, so logical at the time. We *know* people procrastinate, we *know* they're loss averse, they have present bias. So, let's just make it easier for them to do the "right" thing.
ExpertAnd you see this everywhere. Think about your 401(k). After the Pension Protection Act of 2006, auto-enrollment became the standard. The assumption there is, if you're not saving enough, it's a cognitive error, maybe inertia or present bias, so we'll just opt you in by default.
HostOr healthcare. Countries like Spain and Wales use opt-out organ donation systems. The idea being that most people *want* to donate, but the friction of opting in prevents them. So, failing to donate is a flaw of friction, not a genuine preference.
ExpertEven on your phone! Diet-tracking apps, those smartwatch reminders to stand up, auto-renewing subscriptions that are a pain to cancel. They're all justified by this same premise: we know you're flawed, and we're here to help you overcome your own bad decision-making.
HostIt's this incredible leap of faith. The "choice architect" implicitly claims they know what's best for the individual, sometimes even better than the individual themselves.
ExpertAnd that's where the hubris really comes in. But what if that foundational assumption—that we can actually *diagnose* a bad decision—is itself flawed? This is what B. Douglas Bernheim, Aldo Lucia, Kirby Nielsen, and Charles D. Sprenger tackle in their new NBER working paper, "When Are Decisions Improvable? An Evaluation of Diagnostic Methods," published just this March.
HostSo, instead of asking *how* to nudge people, they're asking: *how do we even know if someone made a mistake in the first place?* It’s a meta-critique. An audit of the auditors.
ExpertExactly. And it's backed by the National Science Foundation, so this isn't some fringe critique. It's a rigorous, academic interrogation of the very instruments used to justify paternalism. If their diagnostic tools are broken, then a lot of what we've called "choice architecture" is just paternalism dressed up as math.
HostThat's a pretty strong claim. Let's dig into these diagnostic methods. The paper identifies three main ways behavioral economists have historically tried to determine if someone's made a "bad" or "improvable" decision. What’s the first one?
ExpertThe first is what they call the **Characterization Assessment Method**. This one is pretty straightforward. It says you've made a bad choice if you hold specific, factual misconceptions about that choice.
HostSo, if I can't explain my own decision, it's considered a mistake?
ExpertPrecisely. Researchers ask subjects to describe the parameters of their decision. If you can't accurately calculate the expected value, or you misunderstand the probabilities, or you fail a basic comprehension test about the rules of the choice, then the researcher concludes your decision was "improvable." It's deemed a mistake because you didn't understand what you were doing.
HostGive me a real-world example of that.
ExpertThink about a company's HR department trying to get employees to pick a health insurance plan. If they offer a high-deductible health plan, and an employee chooses it, but then can't explain what a deductible is, or what their out-of-pocket maximum is, the HR department might conclude, "Ah, that employee made a mistake. They chose the HDHP because they didn't understand it, not because it was genuinely their preference."
HostBut is that fair? Just because I can't *mathematically articulate* the ins and outs of a decision doesn't mean my gut preference is wrong. I might not know the exact formula for compound interest, but I still know I prefer cash today over cash tomorrow, or that saving for retirement feels prudent.
ExpertAnd that's the academic critique of this method. Inability to articulate a rational choice doesn't necessarily mean the choice itself is irrational. It might just mean you're not an economist.
HostOkay, so that’s Method 1. What’s the second way they diagnose a "bad" decision?
ExpertThat's the **Decision Confidence Method**. This one gauges how confident a person is in the choice they just made. It's often referred to as measuring "Cognitive Uncertainty," or CU.
HostSo, if I'm not super sure about my decision, it's flagged as potentially suboptimal?
ExpertExactly. After you make a choice, they'll ask you, "What is the percentage chance that you made the *best* decision for yourself?" If you report low confidence—say, 50%, essentially a coin toss—they flag that decision as suboptimal or improv-able. The idea is, if you're not confident, you probably made a mistake.
HostLike that old joke, "I'm not uncertain, I just don't know."
ExpertPretty much. Imagine a retail investor buying a volatile meme stock on Robinhood. If you ask them, "Are you sure this was the optimal financial move for your portfolio?" and they say, "Honestly, I have no idea, I'm just guessing," a behavioral economist using this method would log that trade as a mistake.
HostAnd this relies heavily on self-reporting, doesn't it? It assumes that my subjective doubt perfectly correlates with an objective error. What if I'm just a naturally cautious person, or I understand the inherent uncertainty of, say, the stock market, so I'm never *100%* confident in any financial decision?
ExpertGood point. It's a leap to assume that internal feeling of uncertainty equals a diagnosable error that needs correcting.
HostAlright, third method?
ExpertThe third is the **Pattern Matching Method**. This one is a bit more abstract. It flags a behavior as "improvable" if it mirrors a specific behavioral pattern that's known to be mathematically and objectively suboptimal in another domain.
HostSo, if my decision looks like a known cognitive bias, it's a mistake?
ExpertPrecisely. Economists look for violations of Expected Utility Theory—things like the "Common Ratio Effect" or "small-stakes risk aversion." For example, if you exhibit a pattern of being overly risk-averse in a small-stakes game, like flipping a coin for ten dollars, the pattern matching method might say that if you scaled that up, it would imply an absurd, irrational fear of losing money.
HostSo, if I buy an extended warranty on a $20 toaster, even though the expected value is terrible, Pattern Matching would say I'm making a mistake.
ExpertExactly. Because the math just doesn't add up, and that pattern of over-valuing small probabilities of loss is a known cognitive bias. Therefore, your choice is a mistake.
HostOkay, so we have Characterization Assessment—did you understand it? Decision Confidence—are you sure about it? And Pattern Matching—does it look like a known error? These all sound reasonable in isolation. But the Bernheim paper applied all three to the same set of decisions. What happened? This is the juicy part, right?
ExpertThis is the climax. They set up an experiment with risky choices, specifically binary lottery choices. Participants chose between different lotteries—like a safe guaranteed payout versus a risky gamble with a higher potential payout. Nothing crazy. But the key was, they then applied *all three* diagnostic methods to the exact same decisions these people made.
HostSo, for each choice, they asked: "Did you understand the lottery parameters?" "How confident are you in your choice?" and then they looked at "Does this choice align with rational expected utility, or does it look like a known cognitive error?"
ExpertAnd the results, as you said, were a mess—but in the most scientifically revealing way possible. They found that the three methods completely contradicted each other. They implied that totally different choices were the "improvable" ones.
HostWait, so Method 1 might say choosing the *safe* option was a mistake because the subject misunderstood the math. But Method 2 might say choosing the *risky* option was a mistake because the subject expressed low confidence. And Method 3 might say a completely *different* choice was a mistake because it violated, say, the Common Ratio Effect?
ExpertExactly that! The authors explicitly state that these methods "have conflicting implications regarding legitimate risk preferences." It's like the diagnostic tools are arguing with each other over what constitutes a "wrong" decision.
HostThat's terrifying for public policy! If I take the exact same human being, making the exact same decision, the diagnosis of whether they are "rational" or "irrational," whether they made a "mistake" or a "valid choice," depends entirely on which academic tool the researcher randomly decided to use. That's not science; that's interpretive dance!
ExpertIt fundamentally exposes a fatal flaw. The inability to distinguish between a cognitive error and a valid appetite for risk. If someone chooses a mathematically suboptimal lottery because they simply enjoy the thrill of gambling, is that a mistake? Pattern Matching would probably say yes. But if they fully understand the odds—they pass the Characterization Assessment—and they're highly confident in their thrill-seeking—they pass Decision Confidence—then two out of three methods say it's a perfectly valid choice!
HostSo the foundational tools of behavioral science cannot agree on what a bad decision looks like. That's a huge problem for a field that prides itself on diagnosing and fixing human irrationality.
ExpertIt's an earthquake. And it's already sparked a huge academic debate. In fact, just weeks before the Bernheim paper was officially published, there was a rebuttal.
HostA rebuttal? Like, other researchers had seen an early draft and immediately pushed back?
ExpertPrecisely. In February 2026, Benjamin Enke from Harvard and Thomas Graeber from the University of Zurich published their own working paper, "A Note on 'When Are Decisions Improvable? An Evaluation of Diagnostic Methods'." They specifically pushed back on the findings regarding Method 2, the Decision Confidence method.
HostWhat was their argument?
ExpertBernheim et al. had suggested that subjects might be misunderstanding the "Cognitive Uncertainty" question. Like, when asked "Did you make the *best decision* for yourself?", subjects might be confusing it with "Will you get the *best outcome*?" Which is a different thing, right? One is about the process, the other about the result.
HostLike, I make a good decision to invest, but the market crashes. The *decision* was good, the *outcome* was bad.
ExpertExactly. But Enke and Graeber replicated the survey and amended it, and they argued that subjects aren't stupidly conflating the two. They say the very presence of *outcome uncertainty* makes the choice *subjectively difficult*. As Enke and Graeber write: "Subjects exhibit uncertainty about their best ex-ante decision... precisely because there is outcome uncertainty."
HostSo, it's not that I made a bad decision, it's that the world is inherently uncertain, and I'm reflecting that uncertainty in my confidence level?
ExpertThat's their "interpretational middle ground." They argue that when a subject shows low confidence and exhibits a behavioral anomaly, like small-stakes risk aversion, it might not be a "mistake" or a "non-standard preference." Instead, it simply reflects *imperfect information* or the inherent difficulty of making decisions in uncertain environments.
HostOkay, so let me translate this academic skirmish. The Stanford and Caltech guys are saying, "The tools are broken and contradict each other, so we can't tell what a mistake is." And the Harvard and Zurich guys are saying, "Hold on, the tools aren't entirely broken, people are just naturally confused when outcomes are uncertain, and that confusion isn't necessarily a 'mistake' in the way we've been thinking about it."
ExpertEither way you slice it, the conclusion is pretty much the same: we have no idea if these people are actually making mistakes. We're just projecting our own definitions onto their choices.
HostAnd that's a key point the authors of the NBER paper make, isn't it? That researchers often project their own risk aversion onto the public. Because economists are trained to maximize Expected Value, they view any deviation from it as an "error."
ExpertAbsolutely. But a retail investor, a patient, or a consumer might have a completely different, entirely legitimate utility function. Maybe they value peace of mind over maximizing returns. Or they value the thrill of a gamble over a safe bet. The paper essentially calls for a reevaluation of decades of behavioral insights—and all the policies built upon them—because they might be built on flawed diagnoses of human error.
HostThis has massive implications for us, the listeners, who are living in this "nudged" world. What does this mean for how we should think about the choices we make and the choices others make for us?
ExpertWell, first, it should make us extremely skeptical. If the very science of "mistakes" is this flimsy, it opens the door to what I'd call "nudge-washing."
HostNudge-washing? Tell me more.
ExpertNudge-washing is when corporations or policymakers use the veneer of behavioral science to justify paternalistic interventions that actually serve *their* institutional bottom lines, rather than the individual's true preferences.
HostAh, I see. Like a bank auto-enrolling you into a higher-fee target-date fund, claiming "behavioral science shows retail investors make mistakes when picking their own allocations."
ExpertExactly. Or a software company making it incredibly difficult to cancel a subscription—a dark pattern—internally justifying it by saying, "consumers suffer from present bias and will regret canceling later." The defense for these actions is built on the idea that they know what's best for you, and you're too biased to choose correctly.
HostBut if this NBER paper is correct, that scientific justification is built on sand. If the three main diagnostic methods can't even agree on what a mistake is in a simple binary lottery, how can a bank definitively claim that my financial choice is a "mistake" that needs correcting?
ExpertThis research, funded by the National Science Foundation, isn't some fringe conspiracy theory. It's a mainstream, rigorously peer-reviewed reckoning within the field. So, listeners should be highly skeptical of corporate or government nudges that claim to "optimize" their health or financial behaviors.
HostSo, if I'm listening, what are the key takeaways I should really walk away with?
ExpertFirst, **trust your preferences.** What an economist labels as a "cognitive bias" or "error" might actually be your perfectly legitimate risk preference. If you prefer a safe, low-yield savings account because it helps you sleep at night, that's a valid choice, regardless of what "Pattern Matching" says.
HostAnd second?
Expert**Question the defaults.** Recognize that every default option you face—from your 401(k) contribution rate to your data privacy settings—was designed by someone who assumed they knew what was best for you. This paper proves they often don't. These defaults are powerful, but they are not neutral or objectively superior.
HostAnd finally?
Expert**Demand better science.** Policymakers need to clarify the assumptions underlying their behavioral interventions. Until the field can establish a unified, reliable way to objectively diagnose a "bad" decision, institutions should default to maximizing freedom of choice rather than paternalistic nudging. The "God Complex" of behavioral economics is being dismantled from the inside out, and that's a good thing for individual autonomy.
HostThat's a powerful thought to end on. It truly feels like behavioral economics is entering a period of necessary humility. If we discard these three diagnostic methods, what replaces them? How *do* we identify a genuine mistake—like forgetting to take life-saving medication—versus a legitimate preference for, say, a riskier investment?
ExpertAnd will government "Nudge Units" like the UK's Behavioral Insights Team or the US Office of Evaluation Sciences update their frameworks in light of this 2026 NBER paper? The debate between researchers like Bernheim and Sprenger, and Enke and Graeber, will undoubtedly shape the next decade of behavioral public policy.
HostIt's going to be fascinating to watch.