Free lesson · Prediction foundations
Calibration bins and their uncertainty
Open interactive lessonPractice calculationsExplore labs
Start with the idea
Among comparable events assigned about 70%, roughly 70% should occur if the forecasts are calibrated. Small bins can look misleadingly far from that target.
Symbols, units & horizon
- k: YES outcomes in a bin
- n: independent event count for this approximation
- p-hat: observed YES frequency
- SE: approximate standard error in probability units
- fixed forecast horizon and bin rules required
When and why to use this
Inspect forecast reliability by domain and horizon with explicit sample-size uncertainty.
Among comparable events assigned about 70%, roughly 70% should occur if the forecasts are calibrated. Small bins can look misleadingly far from that target.
Group forecasts by probability and fixed time-to-resolution. In each bin report mean forecast, event count and observed frequency. Compare domains only when weighting and horizons are compatible.
For a rough independent Bernoulli approximation, frequency standard error is square root of frequency times its complement divided by sample count. Related events and repeated snapshots break this independence assumption; use event clusters in a real study.
Calibration bins and their uncertainty
- Divide successes by event count.
- Use estimated Bernoulli variance p-hat(1−p-hat).
- Divide by n and take the square root; do not treat this plug-in SE as reliable at tiny samples or boundary frequencies.
70 YES outcomes among 100 events give frequency .70 and SE=sqrt(.7×.3/100)≈.045826. The apparent precision is much lower than “70.000%.”
Apply it in a strategy
- Inspect forecast reliability by domain and horizon with explicit sample-size uncertainty.
- Record the input timestamp, executable quantity, currency and horizon. Reconcile the result with a cash-flow or state table.
- Stress this failure condition: Correlated outcomes, selected bins and repeated forecasts can make naive standard errors too small.
Research deliverable
Build and explain a calibration bins and their uncertainty worksheet. Inspect forecast reliability by domain and horizon with explicit sample-size uncertainty.
Evidence boundary: Synthetic arithmetic and scenarios illustrate mechanics. They are not historical returns, a paper replication, or evidence of an executable edge. Research sources and their access limitations are recorded at the end of this module.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
# Python 3.10+; standard library unless NumPy is imported below.
# Inputs and outputs use the units defined in this lesson. Synthetic teaching example.
from math import sqrt
def calibration_bin(successes,count):
if count<=0 or not 0<=successes<=count: raise ValueError("Valid bin counts required")
frequency=successes/count
return frequency,sqrt(frequency*(1-frequency)/count)
print(calibration_bin(70,100))Continue learning
Prediction Markets: Contracts, Probability and Evidence — all lessons- A dollar claim is not a news headline
- From probability to a decision price
- Conditional probabilities and contract dependence
- Brier score: measure the whole probability
- Log loss and overconfident mistakes
- Calibration bins and their uncertainty
- Resolution delay and capital lock-up
- A causal forecast research ledger
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations