Free lesson · Statistics
Estimation uncertainty and Bayesian updating
Open interactive lessonPractice calculationsExplore labs
Start with the idea
An interval makes the uncertainty in an estimate visible. Frequentist confidence intervals and Bayesian posterior intervals answer different questions; neither corrects an unrepresentative sample automatically. Priors are especially influential when data are sparse.
Symbols, units & horizon
- x̄: sample mean
- s: sample SD
- n: observation count
- SE=s/√n: IID standard error
- μ: population mean
- T: t-distributed statistic
- α: probability outside the confidence interval, e.g. .05 for 95% confidence
- t_(n−1,1−α/2), also written t*: positive critical value for the chosen confidence and n−1 degrees of freedom
- a, b: positive Beta prior shape parameters
- W, L: win and loss counts
- p: unknown stable win probability
- ∝: proportional to, omitting the normalising constant
- ∼: distributed according to
- Beta: probability distribution on [0,1]
When and why to use this
Use intervals when deciding whether to continue research, and posterior sensitivity when small samples tempt you to size aggressively. Report the assumptions alongside the estimate.
Invert a t-statistic to get an interval
- Under IID normal sampling, has a t distribution with n−1 degrees of freedom.
- Start from . Multiply by positive SE and isolate μ: .
n=25, mean .20%, s=1%, t*=2.064 at 95%: SE=.20%; interval=.20%±.4128%=[−.2128%, .6128%].
This confidence interval for an IID normal sample mean uses sample standard deviation s and a Student-t critical value. Its repeated-sampling coverage is about the procedure; it is not a posterior probability that a fixed mean lies in this realised interval. Dependence and selection require a different uncertainty calculation.
Multiply the prior by the likelihood
- A Beta prior has kernel . W wins and L losses have likelihood proportional to .
- Add exponents: posterior kernel is , identifying Beta(a+W,b+L). Its mean is first shape parameter divided by their sum.
Beta(1,1) plus 12 wins and 8 losses → Beta(13,9), mean 13/22=59.09%. The raw win fraction was 60%; the prior pulls it toward 50%.
A Beta(a,b) prior with conditionally independent Bernoulli wins updates by adding wins W and losses L. Starting from Beta(1,1), 12 wins and 8 losses produce Beta(13,9), with posterior mean 13/22 ≈ 59.1%. Payoff magnitudes and changing market conditions need additional modelling; win probability alone does not establish an edge.
Research sources, review dates and limitations
Extend the research question
Use a binary contract to separate a forecast probability, its sampling uncertainty and the price you could pay. Assess calibration on events withheld from model selection.
Continue with the connected research module →
Connect the ideas: Information and decision time
Retrieve: Use only information available when the decision is made.
Check the change: Observation dates, release delays, revisions and label maturity require different availability checks.
Research & backtests → Point-in-time data → Time series → Research & robust tuning → Financial machine learning → Putting it all together
Self-assessed. Write your explanation before opening this comparison. No. A revised value or delayed release may still contain unavailable information. Audit actual availability timestamps and fit preprocessing inside each training window.Explain it yourself: Does shifting a feature by one row guarantee that it was available?
Connect the ideas: Dependence and diversification
Retrieve: Joint behavior matters when combining uncertain outcomes.
Check the change: Correlation, cointegration, covariance and event dependence answer different questions.
Linear algebra → Time series → Portfolio construction → Portfolio management theory → Prediction foundations
Self-assessed. Write your explanation before opening this comparison. Correlation measures co-movement under a chosen sample and horizon. It does not establish a stationary combination or contractual convergence; those require separate definitions, tests and implementation checks.Explain it yourself: Why does a highly correlated pair not automatically provide a converging spread?
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
def mean_interval(mean_value, sample_sd, n, t_critical):
"""Supply the t-table critical value for confidence and n-1 df."""
radius = t_critical*sample_sd/n**.5
return mean_value-radius, mean_value+radius
def beta_update(a, b, wins, losses):
if min(a,b) <= 0 or min(wins,losses) < 0:
raise ValueError("Positive shapes and nonnegative counts required")
posterior_a, posterior_b = a+wins, b+losses
return posterior_a, posterior_b, posterior_a/(posterior_a+posterior_b)
print(mean_interval(.002, .01, 25, 2.064))
print(beta_update(1, 1, 12, 8))Continue learning
Statistics & Probability — all lessons- Start with counts, probabilities and averages
- Random outcomes, sample averages and the limits of the bell curve
- Conditional probability: update a belief with evidence
- Prediction markets: probability, price and net expected value
- Expected value: measure the payoff before choosing the risk
- Correlation: how many strategies do you really have?
- Conditional probability and Bayes: where the edge actually lives
- The central limit theorem: sampling means under explicit assumptions
- Estimation uncertainty and Bayesian updating
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations