explainer
what is a Brier score?
The Brier score measures the quality of a probability forecast. It asks a simple question: when you said you were 80% sure, how close was that number to what actually happened? It was introduced by the meteorologist Glenn W. Brier in 1950 to grade weather forecasters, and it has since become the standard yardstick for forecasters of every kind, from intelligence analysts to superforecasting tournaments.
The formula
For a single yes-or-no forecast, take your stated probability, subtract what happened (1 if the event occurred, 0 if it did not), and square the difference. Your overall score is the average across all your forecasts:
Say you were 90% sure a product launch would ship on time, and it did: (0.9 − 1)² = 0.01, a near-perfect score. If it slipped, the same forecast costs (0.9 − 0)² = 0.81, one of the worst scores possible. High confidence is a bet with real stakes.
How to read your score
0 is perfect foresight. 1 is perfect wrongness. The most useful landmark is 0.25: that is what you get by shrugging and saying 50% to everything. A score below 0.25 means your confidence is carrying real information. A score above it means your confidence is actively misleading you, and you would do better by admitting you do not know.
Try it: calculate a Brier score
Why it beats “how often was I right”
Counting correct answers treats a timid 51% and a brash 99% as the same claim. The Brier score does not. It rewards forecasters who are bold exactly when boldness is justified, and humble exactly when it is not. That combination has a name: calibration. Research on expert judgment, most famously Philip Tetlock’s twenty-year study of political forecasts, found that most experts are poorly calibrated, and that keeping score is what separates the ones who improve.
What is a good Brier score?
Lower is better. 0 is perfect. Answering 50% to everything scores 0.25, so anything meaningfully below 0.25 shows real forecasting skill. Superforecasters in the Good Judgment Project averaged roughly 0.15 to 0.2 on hard geopolitical questions, and top performers went lower.
Who invented the Brier score?
Glenn W. Brier, an American meteorologist, proposed it in 1950 to grade weather forecasts. It remains the standard way to score probabilistic predictions in meteorology, intelligence analysis, and forecasting tournaments.
Why not just count how often someone is right?
Accuracy ignores confidence. Someone who says 51% and someone who says 99% are treated the same when the event happens, yet the second claim was far stronger. The Brier score rewards saying 99% only when you can back it up, and punishes it hard when you cannot.
How is it different from calibration?
Calibration asks whether your 80% claims come true about 80% of the time. The Brier score wraps calibration and sharpness into one number: to score well you must be both honest about your uncertainty and bold enough for your forecasts to be informative.
Measure your own
Two ways to find out where you stand. The fast one: a two-minute calibration test on questions with known answers. The real one: log your actual decisions and predictions with a confidence number, wait, and grade them when reality reports back. Brier is a journal built for exactly that, with reminders that arrive when each decision is due for review and a calibration curve that grows as you go.