Daniel Kahneman, Olivier Sibony, and Cass R. Sunstein published Noise: A Flaw in Human Judgment with Little, Brown Spark in 2021. The book examines a neglected form of error: people or systems making different judgments in cases that should be treated alike.
Bias and noise are not synonyms. Bias is a systematic tendency to err in a direction—for example, consistently estimating too high. Noise is unwanted variability around the judgment. A group can be consistent and biased, variable but correct on average, or both biased and noisy.
This distinction matters wherever judgment affects sentencing, diagnosis, hiring, insurance, forecasting, performance reviews, child protection, or strategy. It also creates a danger: reducing variability can make an unjust standard more consistently unjust. Decision quality therefore needs tests of both consistency and validity.
Noise appears only across comparable judgments
One decision cannot display variability by itself. Noise becomes visible when judges evaluate the same case, when one judge evaluates matched cases, or when outcomes are compared against a defensible standard.
Not every difference is unwanted. Two medical cases may look similar while differing in a relevant symptom. Two employees may receive different evaluations because their roles or constraints differ. A noise analysis begins by defining which cases should receive similar treatment and why.
For a repeated decision, write:
- the unit being judged;
- the outcome or scale;
- features that should legitimately change the judgment;
- features that should not;
- the acceptable range of disagreement;
- the evidence available for later validation.
Without that contract, variation can be mislabeled according to whoever controls the process.
Level, pattern, and occasion noise
The authors distinguish several sources. Level noise occurs when some judges are generally more severe, optimistic, or generous than others. Pattern noise occurs when judges react differently to particular features of a case. Occasion noise occurs when the same person's judgment changes with transient conditions.
Suppose three managers score identical work samples. One manager scores everyone lower: level noise. Another strongly rewards technical novelty while a colleague prioritizes clarity: pattern noise. A manager gives different scores on different days despite the same standard: occasion noise.
These categories guide remedies. Anchored scales may reduce level differences. Explicit weighting can expose pattern differences. Delaying a fatigued judgment or using a second assessment may reduce occasion effects. The categories are analytical, not diagnoses of the people involved.
Run a noise audit before prescribing a cure
A noise audit gives multiple qualified judges the same realistic cases and measures dispersion. The judges should work independently under comparable conditions. If they discuss cases first, conformity and hierarchy can conceal the original variation.
A small organizational audit can follow six steps:
- Select several representative cases with confidential details protected.
- Define the judgment scale and relevant information.
- Obtain independent assessments before discussion.
- Compare spread, average, reasons, and subgroup patterns.
- Ask which differences reflect legitimate interpretation.
- Decide whether the cost of reduction justifies intervention.
An audit of five cases cannot estimate an entire system with precision, but it can reveal whether assumed agreement exists. High-stakes use requires appropriate statistical design, legal review, privacy controls, and affected-party input.
Decision hygiene structures the process
The book calls its remedies decision hygiene. Like physical hygiene, the methods target classes of error without predicting exactly which mistake they prevent.
Important practices include:
- collecting independent judgments before group discussion;
- breaking a complex evaluation into separate dimensions;
- delaying the overall impression until components are assessed;
- using common, behaviorally anchored scales;
- comparing with relevant reference classes;
- aggregating several estimates where appropriate;
- creating guidelines or rules when the tradeoff favors consistency.
The sequence matters. If a hiring panel shares an overall impression before members assess evidence, later scores may rationalize the first speaker. Independent component ratings preserve information long enough to compare it.
A decision journal can perform a related function for personal choices by recording criteria and forecasts before the outcome invites revision.
Structured judgment is not mechanical judgment
Breaking a decision into components does not require pretending that every quality is measurable with precision. A structured process can include narrative evidence and professional judgment while making criteria visible.
Consider evaluating a grant proposal. Separate feasibility, expected benefit, risk, team capacity, and ethical safeguards. Require reasons for each score before an overall recommendation. The final decision may still involve a tradeoff, but disagreement becomes easier to locate and review.
Structure can also create false confidence. A numbered scale may conceal ambiguous definitions, unreliable inputs, or inappropriate weights. The design should be tested against outcomes and revised when it produces systematic harm.
Rules and algorithms involve value choices
The authors argue that simple rules or mechanical models can sometimes outperform unaided judgment in consistency. A rule never becomes tired or changes severity after a difficult morning. That advantage is real but incomplete.
Algorithms inherit the target, training data, variables, and thresholds chosen by people. Historical data may encode discrimination. A highly consistent model may optimize cost while ignoring dignity, rare conditions, or unequal access to documentation.
Before replacing judgment with a rule, ask:
- What outcome is the model designed to predict?
- Is that outcome itself fair and valid?
- Which groups are underrepresented or differently measured?
- What happens in unusual cases?
- Can a person understand and contest the decision?
- Who monitors drift and disparate impact?
Human review is not automatically safer, since reviewers can reintroduce noise and bias. Appeals need reasons, authority, and evidence rather than symbolic human presence.
Accuracy cannot be inferred from agreement
If all judges agree, noise is low. They may still share a stereotype, defective guideline, or missing fact. Consistency should therefore be paired with calibration or outcome evidence wherever possible.
For predictions, compare forecast probabilities with observed frequencies. For evaluations, examine whether scores predict the intended performance and whether irrelevant characteristics influence results. For legal or ethical judgments, empirical prediction alone cannot determine legitimacy; rights and due process constrain the objective.
The distinction also corrects a common use of cognitive-bias language. Naming bias or noise does not prove which decision is right. It identifies a reason to improve the process and evidence.
Some discretion is legitimate
Professionals often work with incomplete, contextual, or morally complex information. Eliminating all variation may erase compassion, local knowledge, or meaningful differences not captured by a form. Medicine, social work, education, and law contain cases where exceptions are necessary.
The relevant choice is not “discretion or rules” in the abstract. It is which parts of judgment should be standardized, which should remain open, and how exceptions are documented and reviewed.
A useful architecture has a default standard, explicit exception criteria, written reasons, periodic pattern review, and an appeal route. If exceptions consistently favor high-status people, the system has learned selective flexibility.
Evidence quality varies across examples
Noise synthesizes research from many fields, but vivid examples do not all carry equal weight. Some popular behavioral findings have faced replication or interpretation challenges. The often-repeated story that judges become dramatically harsher before meal breaks, for example, has attracted substantial methodological criticism and should not bear the general case for occasion noise.
The broader phenomenon of inconsistent professional judgment does not depend on one dramatic study. Still, each intervention should be supported by evidence close to its domain. A finding from forecasting does not automatically establish how to structure diagnosis or sentencing.
An independent 2024 review of the book's evidence base argues that some examples rest on fragile research. That criticism supports a source-by-source approach rather than dismissal of the entire concept.
High-stakes safeguards
Reducing noise in health, employment, credit, child protection, or criminal justice is not a personal productivity exercise. Errors can remove liberty, income, treatment, housing, or family contact. Changes require professional governance, privacy protection, legal compliance, and analysis of unequal impact.
Do not use a simplified scorecard to make a medical, legal, or financial decision outside appropriate expertise. Do not collect sensitive employee or client cases for an informal workshop without authority and protection. An audit can itself create harm if data are exposed or consequences follow from experimental ratings.
A low-stakes application
Choose a repeated personal decision such as prioritizing maintenance requests or reviewing draft proposals. Define three criteria before seeing the next case. Score them separately, write an overall judgment, and note confidence. After five cases, compare whether the same evidence received similar treatment.
Then test usefulness, not merely lower variance. Did important context disappear? Did the criteria improve outcomes? Was the process worth its cost? Revise one element rather than turning the first scorecard into permanent policy.
Noise gives organizations a vocabulary for error that average results can hide. Its strongest lesson is procedural: independent judgment, explicit standards, and measurement can expose inconsistency. Its necessary companion is justice. A system should become more consistent only in applying a defensible standard, with room for reasons, correction, and appeal.