Free lesson · Prediction foundations
A causal forecast research ledger
Open interactive lessonPractice calculationsExplore labs
Start with the idea
A valid forecast comparison uses only information available when the forecast was issued and outcomes that eventually matured.
Symbols, units & horizon
- BSS: dimensionless Brier skill score
- BS_model: model Brier score
- BS_baseline: positive baseline Brier on identical events and weights
- evaluation horizon: predeclared and common
When and why to use this
Produce a forecast research memo with frozen chronology, matched baselines and a distinct trading-cost analysis.
A valid forecast comparison uses only information available when the forecast was issued and outcomes that eventually matured.
Save event ID, contract text/version, decision timestamp, horizon, model probability, contemporaneous bid/ask, and resolution source. Keep an append-only original forecast even when later data corrections are necessary.
Use a chronological split by event, not trade row. Lock model and cost rules before the final evaluation. Report both forecast scores and separate implementable outcomes, including unfilled orders and unsettled events.
A causal forecast research ledger
- Compute model and baseline Brier scores on the same eligible outcomes.
- Divide model score by the positive baseline score.
- Subtract from one; positive means improvement relative to this baseline, not guaranteed trading value.
Model BS=.18 and baseline BS=.24 give BSS=1−.18/.24=.25, a 25% reduction in this squared-error metric.
Apply it in a strategy
- Produce a forecast research memo with frozen chronology, matched baselines and a distinct trading-cost analysis.
- Record the input timestamp, executable quantity, currency and horizon. Reconcile the result with a cash-flow or state table.
- Stress this failure condition: Selecting domains or horizons after inspecting test scores turns the test set into a tuning set.
Research deliverable
Build and explain a a causal forecast research ledger worksheet. Produce a forecast research memo with frozen chronology, matched baselines and a distinct trading-cost analysis.
Evidence boundary: Synthetic arithmetic and scenarios illustrate mechanics. They are not historical returns, a paper replication, or evidence of an executable edge. Research sources and their access limitations are recorded at the end of this module.
Primary research and operational references · reviewed 12 September 2026
Further reading: Saguillo et al. — Unravelling the Probabilistic Forest ↗
arXiv preprint v1, 5 August 2025 · Review: Abstract, data-section excerpts and full-text structure inspected 2026-09-12. Markets/data: Polymarket markets resolved 1 April 2024–1 April 2025; data-section endpoint dates checked in the primary manuscript. Interpretation: Distinguishes within-market complete sets from cross-market logical relations. Limits: Historical detected opportunities depend on contract classification, timestamps and execution assumptions; not a current opportunity list. No independent replication performed.
Further reading: Le — Decomposing Crowd Wisdom ↗
arXiv preprint v2, revised 4 August 2026 · Review: Abstract, data-section excerpts and full-text structure inspected 2026-09-12. Markets/data: Kalshi and Polymarket; the primary data tables identify a 31 December 2025 cutoff; inspect exact extraction rules before reproducing. Interpretation: Motivates calibration by domain/horizon and explicit sampling uncertainty. Limits: Descriptive calibration and in-sample decompositions do not establish net trading profits; clustered events matter. No independent replication performed.
Further reading: Polymarket — How positions work ↗
Living official product documentation · Review: Documentation inspected 2026-09-12. Markets/data: Conditional outcome-token mechanics; no empirical sample. Interpretation: Operational starting point for split, merge and redemption workflows. Limits: Record the exact market rules and collateral/version before applying a generic binary model; platform behavior may change. No independent replication performed.
Research sources, review dates and limitations
Connect the ideas: Dependence and diversification
Retrieve: Joint behavior matters when combining uncertain outcomes.
Check the change: Correlation, cointegration, covariance and event dependence answer different questions.
Statistics → Linear algebra → Time series → Portfolio construction → Portfolio management theory
Self-assessed. Write your explanation before opening this comparison. Correlation measures co-movement under a chosen sample and horizon. It does not establish a stationary combination or contractual convergence; those require separate definitions, tests and implementation checks.Explain it yourself: Why does a highly correlated pair not automatically provide a converging spread?
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
# Python 3.10+; standard library unless NumPy is imported below.
# Inputs and outputs use the units defined in this lesson. Synthetic teaching example.
def brier_skill(model_score,baseline_score):
if model_score<0 or baseline_score<=0: raise ValueError("Nonnegative model and positive baseline scores required")
return 1-model_score/baseline_score
print(brier_skill(.18,.24))Continue learning
Prediction Markets: Contracts, Probability and Evidence — all lessons- A dollar claim is not a news headline
- From probability to a decision price
- Conditional probabilities and contract dependence
- Brier score: measure the whole probability
- Log loss and overconfident mistakes
- Calibration bins and their uncertainty
- Resolution delay and capital lock-up
- A causal forecast research ledger
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations