Trading Dev AcademyFree quant education

Free lesson · Alternative data

A reproducible dictionary score for text

Open interactive lessonPractice calculationsExplore labs

Start with the idea

Before using a complex language model, make a small text feature whose exact counting rule can be inspected.

Symbols, units & horizon
  • n₊,n₋: counts of tokens exactly matching disjoint positive/negative dictionaries
  • n_tokens>0: all tokens under the declared lowercase alphabetic tokenizer
  • s: dimensionless score per document at publication
  • repeated matching words count repeatedly

When and why to use this

Create an interpretable text baseline and a debugging fixture before training a learned classifier.

Before using a complex language model, make a small text feature whose exact counting rule can be inspected.

Define the dictionary, tokenization, treatment of case, punctuation, negation and duplicate documents. A naive positive-minus-negative score may classify “not improving” incorrectly. The toy function intentionally ignores negation so its limitation can be tested explicitly.

Financial meanings depend on context: “liability” can be a normal accounting term rather than negative sentiment. Fit dictionaries and learned text models without seeing evaluation outcomes; freeze versions and preserve original publication and revision timestamps.

s=n+−n−ntokens
Deterministic feature definition with fixed tokenization

A reproducible dictionary score for text

  1. Lowercase text and extract alphabetic tokens with the frozen tokenizer.
  2. Count dictionary matches in each polarity set; verify the sets do not overlap.
  3. Subtract negative count from positive count and divide by total token count, including neutral tokens.
Work it by hand

“profit improves but risk rises” has five tokens. Positive dictionary {profit,improves} gives two; negative {risk} gives one. Score=(2−1)/5=.2.

Apply it in a strategy

  • Freeze inputs at the stated decision time and record their units.
  • Create an interpretable text baseline and a debugging fixture before training a learned classifier.
  • Recompute the example, then change the material assumption and explain the difference.

Research deliverable

A reproducible dictionary score for text: produce the worked calculation, a timestamped input record and a written decision addressing this limitation: Dictionary counts omit negation, context, quotation and semantic shifts; a score is not automatically a tradable forecast.

These are synthetic mechanics examples, not historical performance or paper replications. Module evidence and research boundaries record the 12 September 2026 review.

Python implementation

Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.

# Python 3.10+; standard library and NumPy only.
# Synthetic teaching inputs; conventions and units are defined in the notation above.
import re

def dictionary_score(document,positive,negative):
    pos=set(positive); neg=set(negative)
    if pos&neg: raise ValueError('Polarity dictionaries must be disjoint')
    tokens=re.findall(r'[a-z]+',document.lower())
    if not tokens: raise ValueError('No tokens')
    return (sum(t in pos for t in tokens)-sum(t in neg for t in tokens))/len(tokens)

assert dictionary_score('profit improves but risk rises',{'profit','improves'},{'risk'})==.2
print(dictionary_score('profit improves but risk rises',{'profit','improves'},{'risk'}))

Continue learning

Alternative Data: Measurement, Text & Incremental Value — all lessons
  1. From a sampled panel to a population estimate
  2. A reproducible dictionary score for text
  3. An event return needs a predeclared benchmark
  4. Availability delays and signal decay
  5. Noisy proxies and attenuation
  6. Measure improvement against a frozen baseline
  7. From forecast accuracy to a costed decision
  8. A data investment includes coverage, access and ongoing costs

Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations