Trading Dev AcademyFree quant education

Free lesson · Prediction foundations

Log loss and overconfident mistakes

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Log loss heavily penalizes assigning almost no chance to an event that happens. This rewards honest uncertainty.

Symbols, units & horizon
  • L: mean log loss in nats
  • n: forecast count
  • p_i: probability strictly between 0 and 1
  • y_i: binary resolved outcome
  • ln: natural logarithm
  • i: event at the stated evaluation horizon

When and why to use this

Detect overconfident forecast systems and compare full predictive distributions.

Log loss heavily penalizes assigning almost no chance to an event that happens. This rewards honest uncertainty.

Use natural logarithms and report log loss in nats per forecast. Probabilities exactly zero or one create infinite loss when contradicted.

Numerical clipping prevents undefined floating-point operations but changes the score. Record the clipping threshold and never use it to conceal unjustified certainty. Compare forecasts on the same outcome set and horizon.

L=−1n∑i[yiln⁡pi+(1−yi)ln⁡(1−pi)]
Model assumptions, derivation and arithmetic

Log loss and overconfident mistakes

  1. If y=1 retain −ln(p); if y=0 retain −ln(1−p).
  2. Compute each event contribution without rounding probabilities first.
  3. Average contributions and report any clipping policy separately.
Work it by hand

A forecast p=.8 followed by YES loses −ln(.8)≈.223144 nats. Assigning .01 to the same YES gives about 4.60517 nats.

Apply it in a strategy

  • Detect overconfident forecast systems and compare full predictive distributions.
  • Record the input timestamp, executable quantity, currency and horizon. Reconcile the result with a cash-flow or state table.
  • Stress this failure condition: A rare mislabeled or voided event can dominate log loss; resolve data quality before interpreting model failure.

Research deliverable

Build and explain a log loss and overconfident mistakes worksheet. Detect overconfident forecast systems and compare full predictive distributions.

Evidence boundary: Synthetic arithmetic and scenarios illustrate mechanics. They are not historical returns, a paper replication, or evidence of an executable edge. Research sources and their access limitations are recorded at the end of this module.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

# Python 3.10+; standard library unless NumPy is imported below.
# Inputs and outputs use the units defined in this lesson. Synthetic teaching example.
from math import log
def log_loss(probabilities,outcomes):
    if not probabilities or len(probabilities)!=len(outcomes): raise ValueError("Aligned forecasts required")
    if any(not 0<p<1 for p in probabilities) or any(y not in (0,1) for y in outcomes): raise ValueError("Use probabilities strictly inside (0,1)")
    return -sum(y*log(p)+(1-y)*log(1-p) for p,y in zip(probabilities,outcomes))/len(outcomes)

print(log_loss([.8],[1]))

Continue learning

Prediction Markets: Contracts, Probability and Evidence — all lessons
  1. A dollar claim is not a news headline
  2. From probability to a decision price
  3. Conditional probabilities and contract dependence
  4. Brier score: measure the whole probability
  5. Log loss and overconfident mistakes
  6. Calibration bins and their uncertainty
  7. Resolution delay and capital lock-up
  8. A causal forecast research ledger

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations