Incentives Matter

The Paradox of Warning Signals: Do We Actually Want to Know?

April 10, 202616:34Incentives Matter

This episode delves into a new NBER paper that uncovers a fundamental human flaw: our distorted perception of warning signals. It explores how people irrationally value the accuracy of alarms, often overvaluing inaccurate signals, particularly when threats are rare, leading to potentially catastrophic decisions. Listeners will learn about the cognitive biases influencing our response to warnings and the critical trade-offs between false positives and false negatives in various real-world scenarios.

Key Takeaways

Detailed Report

A catastrophic event like the Deepwater Horizon explosion in 2010, where a critical safety alarm was disabled by the crew, highlights a baffling human tendency: our distorted perception of warning signals. A new NBER working paper, "Preferences for Warning Signal Quality: Experimental Evidence" by Ugarov, Gaduh, and McGee, delves into this fundamental flaw, revealing how we inherently misvalue information meant to keep us safe.

The Flawed Logic of Warning Signals

The paper investigates a seemingly simple question: How much do people truly value the accuracy of a warning signal? The implications, however, are profound, impacting everything from medical tests to cybersecurity alerts. Every warning system balances the risk of false positives (an alarm when there's no danger) against false negatives (silence when there *is* danger). This trade-off is often poorly understood by users.

To isolate the core cognitive issue, researchers conducted a rigorous experiment with 105 subjects at the Behavioral Business Research Lab. Participants were given a baseline probability of an adverse event and then offered warning signals with varying, clearly stated rates of false positives and false negatives. Crucially, the study controlled for subjects' baseline risk preferences and their Bayesian updating skills, ensuring that poor decision-making wasn't simply due to a lack of mathematical ability or general fear.

Shattering the Rational Model

Traditional economic theory posits that a rational actor would perfectly calculate the costs of a false alarm versus a missed threat. Their willingness-to-pay for a signal should scale perfectly with its quality. For rare events, a rational person would heavily penalize systems with high false-positive rates, knowing most alarms would be false. Conversely, for common events, they would heavily penalize systems with high false-negative rates.

However, the NBER paper's findings completely contradict this rational model. The researchers observed "asymmetric under-responsiveness by prior." When a threat was rare, subjects did not sufficiently lower their willingness-to-pay, overvaluing inaccurate signals despite the mathematical reality that rare events combined with even moderate false-positive rates guarantee a barrage of useless alarms. Similarly, for common threats, subjects failed to adequately penalize systems with high false-negative rates, not reacting proportionally to the degradation of information.

The "Error is Error" Heuristic

The core cognitive glitch identified by the paper is a deeply ingrained, flawed shortcut: the "Error is Error" heuristic. The human brain, it turns out, does not distinguish between false-positive and false-negative errors. It treats the probabilistic cost of a false alarm as cognitively equivalent to the probabilistic cost of a missed threat, completely ignoring the base rate of the event and the vastly different real-world consequences of these failures.

This means that, to the human mind, the annoyance of a car alarm triggered by a cat might feel as costly as a smoke detector failing during a house fire. This shocking simplification leads us to wildly misprice the value of warnings, demanding highly sensitive systems for rare threats, inadvertently signing ourselves up for an overwhelming number of false positives.

Real-World Consequences: From Hospitals to Finance

This "Error is Error" heuristic has severe real-world implications, particularly evident in healthcare:

The False Positive Paradox in Medical Screening

Patients often demand screenings for extremely rare diseases, oblivious to the high likelihood and emotional cost of a false positive. Consider a rare disease affecting 1 in 100,000 people and a new test that is 99.9% accurate, with a 0% false-negative rate and only a 0.1% false-positive rate. While seemingly excellent, if 100,000 people are tested, the test will correctly identify the one person with the disease but also incorrectly flag 100 healthy people as having it. This means if you test positive, your actual chance of having the disease is less than 1%. This statistical trap leads to immense patient anxiety, unnecessary biopsies, and significant medical waste.

The Crisis of Alarm Fatigue

In hospitals, the demand for systems with low false-negative rates (to catch every possible threat) leads engineers to lower alarm thresholds. The inevitable result is alarm fatigue. A UCSF study found nurses in ICUs were exposed to nearly 190 audible alarms per patient per day, with up to 90% being false positives triggered by minor issues like sweating or loose electrodes. Responding to these non-actionable alarms consumes 10% or more of nursing time, leading to burnout and, tragically, missed true life-threatening events when nurses tune out or disable alarms. This problem extends to cybersecurity (alert overwhelm) and finance (ignored stock market alerts).

Designing for Imperfect Humans

The paper's most critical takeaway is that simply providing transparent error rates to users does not fix the problem. Due to the "Error is Error" cognitive blind spot, users fail to multiply the error rate by the base rate of the event, fundamentally misjudging the value of the information.

Paternalistic Design and Concrete Costs

Choice architects and product designers must adopt a more paternalistic approach. Instead of relying on flawed user preferences to set alarm thresholds, systems should be mathematically optimized based on the *actual* real-world consequences of errors. When communicating error rates, designers must translate abstract probabilities into natural frequencies and specific, concrete costs.

For example, instead of saying, "This fraud alert system has a 99% accuracy rate and a 1% false positive rate," a better design would state: "Because fraud is very rare, if you turn on this maximum-security setting, you will receive roughly 50 alerts this year. 49 of them will be false alarms that will temporarily freeze your credit card. Are you sure you want this setting?" This forces the user's brain to weigh the asymmetric cost of the false positive, bypassing the flawed heuristic.

Systemic Solutions

In high-stakes environments like hospitals, solutions also involve systemic workflow redesign: adjusting device sensitivity to trigger alarms only for highly specific dangers, consolidating redundant alerts, and implementing slight delays before sounding an alarm to allow for self-correction. The biggest barrier to these improvements often lies in legal liability, as missed threats (false negatives) lead to lawsuits, while thousands of false positives merely lead to an annoyed workforce.

The "Error is Error" heuristic is a powerful insight into why we consistently mismanage warning signals. As AI takes over signal generation in various fields, the challenge will be to train these systems to act as paternalistic choice architects, filtering noise and presenting humans with only those signals that truly have a high expected utility, rather than exacerbating the False Positive Paradox. We must design warning systems for the easily fatigued, statistically blind human beings we actually are, not for perfectly rational calculators.

Show Notes

Works Referenced

  • Preferences for Warning Signal Quality: Experimental Evidence: The foundational NBER working paper by Ugarov, Gaduh, and McGee, which explores how human cognitive biases lead to a distorted perception and misvaluation of warning signal accuracy.
  • Deepwater Horizon Oil Spill: Information about the catastrophic 2010 oil rig explosion, cited as a real-world example of the deadly consequences of disabling critical safety alarms due to a high rate of false positives.
  • ECRI Institute: A non-profit organization that consistently identifies alarm fatigue as a top healthcare technology hazard, highlighting the significant risks associated with poorly designed or managed warning systems.

Glossary

  • False Positive: An alarm or test result that indicates a threat or condition when none actually exists (e.g., a car alarm triggered by a cat, a medical test incorrectly indicating a disease).
  • False Negative: A failure of an alarm or test to indicate a threat or condition when one does exist (e.g., a smoke detector failing to activate during a fire, a medical test missing an existing disease).
  • Bayesian Updating: The process of revising an estimate of a probability or belief when new information becomes available, often used in statistics to calculate how probabilities change with evidence.
  • Heuristic: A mental shortcut or rule of thumb that allows people to solve problems and make judgments quickly, but can sometimes lead to systematic errors or biases.
  • Alarm Fatigue: A phenomenon where excessive alarms or notifications lead to desensitization, delayed responses, or even the disabling of warning systems, often resulting in missed critical events.
  • False Positive Paradox: A statistical phenomenon where, for very rare events, a test with a low false-positive rate can still yield a high proportion of false positives among all positive results, making a positive test result less indicative of the actual condition than expected.
  • Choice Architect: An individual or entity responsible for designing the environment in which people make decisions, influencing their choices without restricting options, often by structuring information or defaults.

Sources / References

Full Transcript

HostImagine a scenario: it's 3 AM, an oil rig in the middle of the ocean. A critical safety alarm is blaring. But instead of responding, the crew disables it. Not because the danger is gone, but because they didn't want people woken up by *false* alarms.
ExpertIt sounds utterly insane, doesn't it? But that's exactly what was reported on the Deepwater Horizon rig before that catastrophic explosion in 2010. And it perfectly encapsulates this fundamental, almost baffling human flaw that a new NBER paper brings to light: our completely distorted perception of warning signals.
HostSo, it's not just about ignoring an alarm because you're tired, but something deeper, something in how we inherently value—or *misvalue*—the very information meant to keep us safe?
ExpertPrecisely. It turns out we're hardwired to make some incredibly poor choices about what makes a 'good' warning, and the consequences range from annoying to truly deadly.
HostOkay, this is fascinating. Let's dig into this NBER working paper, "Preferences for Warning Signal Quality: Experimental Evidence," by Ugarov, Gaduh, and McGee. They're asking a question that seems so simple on the surface: How much do people actually value the accuracy of a warning signal? But the implications, as you just hinted, are anything but.
ExpertRight, and it's a question that defines modern safety. Think about it: every alarm, every notification, every medical test – it's all about balancing false positives against false negatives. How often does the alarm go off when there's no danger, versus how often does it stay silent when there *is* danger? These aren't abstract concepts; they're the difference between a minor annoyance and, well, a burning oil rig.
HostThe researchers really tried to isolate the core cognitive issue here, setting up a rigorous experiment at the Behavioral Business Research Lab. They gave 105 subjects a baseline probability of an adverse event occurring – a known prior – and then offered them warning signals with varying, clearly stated rates of false positives and false negatives.
ExpertAnd importantly, they weren't just measuring how good people were at math, or how generally fearful they were. They controlled for that. They used separate tasks to measure each subject's baseline risk preference and their Bayesian updating skill – essentially, their raw ability to calculate how probabilities change when new information comes in. So, we can't just blame "bad math."
HostSo, they're really trying to get to the heart of what our brain *does* when confronted with information that has potential errors. And for me, as someone who looks at how this plays out in the wild, a "signal" isn't just a physical siren. It's that push notification from your banking app about possible fraud, it's the check engine light on your dashboard, it's a mammogram, or a PSA test for prostate cancer. Every single one of those requires someone to set a threshold for when that alarm triggers.
ExpertExactly. And that threshold is a behavioral design choice. If you set it low, meaning the system is super sensitive, you reduce missed threats – false negatives – but you unleash a flood of false alarms. If you set it high, you get fewer false alarms, but you risk missing actual dangers. It’s a constant trade-off, and these researchers found we're terrible at understanding that trade-off.
HostTo truly appreciate how wild these findings are, we first need to understand what a perfectly rational actor *should* do. In traditional economics, the "Value of Information" is all about how much an alert improves the expected outcome. A rational, risk-neutral person should perfectly calculate the costs of a false alarm versus a missed threat, and their willingness-to-pay for that signal should scale perfectly with its quality.
ExpertWhich means, if an event is incredibly rare, a rational person knows that almost every alarm is going to be a false positive. So, they should heavily penalize—pay much, much less for—a system with a high false-positive rate. Conversely, if an event is super common, the real danger is missing it. So, you should heavily penalize a system with a high false-negative rate. That's the textbook, rational approach.
HostAnd the NBER paper’s data completely shatters that rational model. They found something they call "asymmetric under-responsiveness by prior." That's a mouthful, but it means our reaction to signal quality is all over the map depending on how common or rare the threat is.
ExpertIt's astonishing. When the threat was rare, subjects did *not* lower their willingness-to-pay enough to account for the high cost of false positives. They just overvalued inaccurate signals. They ignored the mathematical reality that a rare event, paired with even a moderate false-positive rate, guarantees a barrage of useless, costly false alarms.
HostSo even if the math was right in front of them, they just couldn't internalize it?
ExpertRight! And it went the other way too. When the threat was common, subjects *did* reduce their demand for signals with high false-negative rates for rare events, but they *failed* to properly penalize the system for false negatives when the event was frequent. They just didn't react proportionally to the degradation of the information. As the paper points out, this aligns with other behavioral economics research that shows we consistently overvalue inaccurate signals.
HostOkay, so if it's not simply "people are bad at math"—because the researchers controlled for Bayesian updating skills—and it's not just that people are super risk-averse—because they controlled for that too—then what *is* going on? What's the underlying cognitive glitch?
ExpertThat’s the real gem of this paper. They conclude that our brains default to a deeply ingrained, flawed cognitive shortcut. A decision-making heuristic where we simply do not distinguish between false-positive and false-negative errors.
HostYou mean... an error is just an error?
ExpertExactly. To the human brain, an "error is an error." The mind treats the probabilistic cost of a false alarm as cognitively equivalent to the probabilistic cost of a missed threat. It completely ignores the base rate of the event and the vastly different real-world consequences of those two types of failures.
HostThat's... absurd, when you think about it. So, my brain is basically telling me that the annoyance of my car alarm going off when a cat brushes against it is as bad as my smoke detector failing to go off when my house is on fire? The *cost* feels the same to my brain?
ExpertCognitively, yes. It's a shocking simplification. And because we weigh these errors equally, we wildly misprice the value of warnings. We demand systems that are highly sensitive to rare threats, completely blind to the fact that we're signing ourselves up for a barrage of false positives that will eventually render the system useless.
HostThis isn't just an academic curiosity, then. This "Error is Error" heuristic has massive real-world fallout. Like in healthcare, for example, which is rife with these kinds of paradoxes.
ExpertAbsolutely. The False Positive Paradox, specifically. It perfectly explains why patients constantly demand medical screenings for extremely rare diseases, completely ignoring the high likelihood and emotional cost of a false positive. Let's walk through it. Imagine a rare disease that affects, say, 1 in 100,000 people.
HostOkay, very rare.
ExpertNow, a new medical test comes out. It's 99.9% accurate. It has a 0% false-negative rate—it catches every single case. And it only has a 0.1% false-positive rate. To the average patient relying on the "Error is Error" heuristic, that sounds like an incredible, highly valuable signal. They will pay a high premium for it.
Host99.9% accurate, 0.1% false positive. That sounds like a dream test.
ExpertRight? But here's the reality. If you test 100,000 random people, the test *will* correctly identify the one person who has the disease. But because of that 0.1% false-positive rate, it will also incorrectly flag *100* healthy people as having the disease.
HostWait. So out of 100,000 people, it finds one sick person, and 100 healthy people it says are sick?
ExpertPrecisely. So, if you get a positive result on that test, you're one of 101 people who tested positive. And only one of you actually has the disease. Your actual chance of having the disease, if you get a positive result, is less than 1%. It's a statistical trap.
HostThat's mind-blowing. People get told they have a rare, potentially deadly disease, go through biopsies, anxiety, further tests... all because of a mathematically high likelihood of it being a false alarm?
ExpertExactly. And because we cannot intuitively process this False Positive Paradox, and because we view "errors as errors" regardless of the base rate, we over-demand these low-probability medical screenings. This leads to millions of dollars wasted, profound patient anxiety, and systemic medical waste.
HostThis leads us directly to the crisis of "alarm fatigue," which I know is a huge problem in hospitals. When we demand systems with low false-negative rates—we want to catch *every* possible threat—engineers respond by lowering the threshold for the alarm. And then, inevitably...
ExpertThe alarms start going off constantly. In hospital ICUs, patients are hooked up to cardiac monitors that are incredibly sensitive but have low specificity. A study at UCSF found nurses were exposed to nearly *190* audible alarms per day, per patient. In some units, nurses hear over a hundred alarms an hour.
Host190 alarms per patient per day? That's just noise.
ExpertAnd the research shows that the vast majority of these rhythm alarms—sometimes up to 90%—are false positives. They're triggered by patients sweating, or squirming, or an ECG electrode losing contact with the skin. It's not a real danger.
HostSo nurses are spending a huge chunk of their day just responding to non-threats?
ExpertA massive amount. Responding to non-actionable alarms consumes 10% or more of nursing time. That's essentially 10% of nursing wages spent chasing false positives. And the human cost is fatal. The ECRI Institute consistently identifies alarm fatigue as a top healthcare technology hazard. When the cacophony becomes background noise, nurses inevitably tune out, turn down, or even turn off the alarms, leading to tragic cases where a true, life-threatening event is missed.
HostAnd this isn't just a hospital problem, right? We see this in cybersecurity, in finance...
ExpertEverywhere. In cybersecurity, software is notoriously tuned to flag any anomalous behavior, which leads to IT professionals drowning in "alert overwhelm." In finance, retail investors set up stock market alerts, only to ignore them when the market becomes volatile and the alerts become ubiquitous. And the critical takeaway here is that *transparency doesn't fix the problem*.
HostWait, really? You're saying if a company tells me, "Hey, this fraud alert system has a 1% false positive rate," that doesn't help me?
ExpertIt doesn't, because of the "Error is Error" cognitive blind spot. You read "1% error" and you fail to multiply it by the base rate of the event. You fundamentally misjudge the value of that transparent information. So, the transparency, while well-intentioned, ends up being useless from a decision-making perspective.
HostSo if humans are predictably irrational and mathematically blind to these asymmetric costs, how do we fix the systems we rely on? What's the actionable takeaway for product designers, for policymakers, for choice architects?
ExpertThe NBER paper's findings strongly suggest that choice architects need to be far more paternalistic when designing warning systems. Relying on "user preference" to set alarm thresholds is dangerous because user preference is demonstrably flawed. If you ask a user what they want, they will demand a system that catches *everything*, inadvertently signing themselves up for alarm fatigue.
HostSo we shouldn't ask users what they want? That feels counter-intuitive to good design.
ExpertNot entirely, but designers *must* take the burden of Bayesian updating away from the user. Instead of expecting users to weigh the costs of false positives against false negatives, the system designer has to mathematically optimize that threshold based on the *actual* real-world consequences of those errors. Even if that means the system occasionally misses a low-level threat.
HostThat means, if designers *do* need to communicate error rates to users, they can't use raw percentages like we just discussed. They have to translate those statistical likelihoods into something tangible.
ExpertExactly. To bypass the "Error is Error" heuristic, you need to use natural frequencies and specific, concrete costs. So, a bad design would be: "This fraud alert system has a 99% accuracy rate and a 1% false positive rate." Your brain thinks: "Great! I'll turn that on!"
HostAnd the good design?
ExpertA good design would be: "Because fraud is very rare, if you turn on this maximum-security setting, you will receive roughly 50 alerts this year. 49 of them will be false alarms that will temporarily freeze your credit card. Are you sure you want this setting?"
HostWhoa. That's a completely different proposition. My brain now immediately understands the actual, painful cost of those false positives.
ExpertIt forces your brain to actually weigh the asymmetric cost of the false positive, bypassing that flawed heuristic entirely. It's about translating abstract probability into concrete, experienced consequence.
HostAnd in high-stakes environments like hospitals, this also means systemic workflow redesign. It's not just about the numbers, but how the alerts are delivered.
ExpertYes. Solutions include adjusting the sensitivity of devices so alarms only trigger when a *highly* specific danger threshold is reached, drastically reducing nuisance alarms. Consolidating redundant alerts is crucial – instead of multiple devices screaming about the same thing, systems should escalate only when multiple data points corroborate a threat. And even implementing slight delays before sounding an alarm – giving the patient's vitals a moment to self-correct before screaming for a nurse.
HostThis NBER paper truly is a perfect encapsulation of behavioral economics in the 2020s. It goes beyond just saying "humans are biased" and gives us specific, mathematically modeled evidence of *how* our cognitive machinery misfires in the modern information age.
ExpertIt highlights that gap between the rational economic models we often build our systems upon, and the messy, predictably irrational brains we actually possess. The "Error is Error" heuristic is a powerful insight into why we get so much wrong.
HostSo, thinking about the future, especially with AI coming into play – as AI takes over signal generation, reading medical scans, predicting supply chain disruptions – will it just exacerbate this False Positive Paradox by detecting even smaller anomalies? Or can AI be trained to be that "paternalistic choice architect," filtering out the noise and only presenting humans with signals that truly have a high expected utility?
ExpertThat's the million-dollar question. And then there's the legal side. The biggest barrier to fixing alarm fatigue, for instance, is legal liability. Hospitals and tech companies err on the side of false positives because a missed threat—a false negative—results in a lawsuit, whereas a thousand false positives just result in an annoyed, burned-out workforce. How do we align those legal incentives with human cognitive limitations?
HostRational economic models wish we were calculators. But as the "Error is Error" heuristic proves, we are not. We are biological creatures easily overwhelmed by noise.
ExpertIf we want to build a safer, more efficient world, we simply have to stop designing warning signals for the perfectly rational actor, and start designing them for the easily fatigued, statistically blind human beings we actually are.