Free lesson · Prediction foundations
Brier score: measure the whole probability
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A forecast should be assessed by how much probability it assigns, not only by whether its favorite outcome wins.
Symbols, units & horizon
- BS: Brier score in squared probability units
- n: number of evaluated forecasts
- i: forecast/event index
- p_i: forecast probability between 0 and 1
- y_i: resolved binary outcome 0 or 1
- common forecast horizon is specified by the evaluation
When and why to use this
Compare probabilistic forecasts on matched chronological event sets.
A forecast should be assessed by how much probability it assigns, not only by whether its favorite outcome wins.
The Brier score averages squared errors between probabilities and binary outcomes. Lower is better under the stated event weighting. A constant .5 forecast scores .25 on every binary event, providing a simple reference.
Choose one forecast at a fixed horizon per event or explicitly weight repeated forecasts. Thousands of trades on one election are not thousands of independent outcomes. A lower Brier score does not include the cost of purchasing shares.
Brier score: measure the whole probability
- For each event subtract its outcome from its forecast.
- Square the error so errors do not cancel.
- Average with the declared event weighting and compare the same events/horizons for a baseline.
p=[.8,.3], y=[1,0] give squared errors .04 and .09. Brier=(.04+.09)/2=.065, versus .25 for a .5 baseline.
Apply it in a strategy
- Compare probabilistic forecasts on matched chronological event sets.
- Record the input timestamp, executable quantity, currency and horizon. Reconcile the result with a cash-flow or state table.
- Stress this failure condition: A scoring improvement can be economically irrelevant if the market ask and costs already exceed the model value.
Research deliverable
Build and explain a brier score: measure the whole probability worksheet. Compare probabilistic forecasts on matched chronological event sets.
Evidence boundary: Synthetic arithmetic and scenarios illustrate mechanics. They are not historical returns, a paper replication, or evidence of an executable edge. Research sources and their access limitations are recorded at the end of this module.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
# Python 3.10+; standard library unless NumPy is imported below.
# Inputs and outputs use the units defined in this lesson. Synthetic teaching example.
def brier(probabilities,outcomes):
if not probabilities or len(probabilities)!=len(outcomes): raise ValueError("Aligned nonempty forecasts required")
if any(not 0<=p<=1 for p in probabilities) or any(y not in (0,1) for y in outcomes): raise ValueError("Invalid probability/outcome")
return sum((p-y)**2 for p,y in zip(probabilities,outcomes))/len(outcomes)
print(brier([.8,.3],[1,0]))Continue learning
Prediction Markets: Contracts, Probability and Evidence — all lessons- A dollar claim is not a news headline
- From probability to a decision price
- Conditional probabilities and contract dependence
- Brier score: measure the whole probability
- Log loss and overconfident mistakes
- Calibration bins and their uncertainty
- Resolution delay and capital lock-up
- A causal forecast research ledger
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations