Free lesson · Alternative data
A reproducible dictionary score for text
Open interactive lessonPractice calculationsExplore labs
Start with the idea
Before using a complex language model, make a small text feature whose exact counting rule can be inspected.
Symbols, units & horizon
- n₊,n₋: counts of tokens exactly matching disjoint positive/negative dictionaries
- n_tokens>0: all tokens under the declared lowercase alphabetic tokenizer
- s: dimensionless score per document at publication
- repeated matching words count repeatedly
When and why to use this
Create an interpretable text baseline and a debugging fixture before training a learned classifier.
Before using a complex language model, make a small text feature whose exact counting rule can be inspected.
Define the dictionary, tokenization, treatment of case, punctuation, negation and duplicate documents. A naive positive-minus-negative score may classify “not improving” incorrectly. The toy function intentionally ignores negation so its limitation can be tested explicitly.
Financial meanings depend on context: “liability” can be a normal accounting term rather than negative sentiment. Fit dictionaries and learned text models without seeing evaluation outcomes; freeze versions and preserve original publication and revision timestamps.
A reproducible dictionary score for text
- Lowercase text and extract alphabetic tokens with the frozen tokenizer.
- Count dictionary matches in each polarity set; verify the sets do not overlap.
- Subtract negative count from positive count and divide by total token count, including neutral tokens.
“profit improves but risk rises” has five tokens. Positive dictionary {profit,improves} gives two; negative {risk} gives one. Score=(2−1)/5=.2.
Apply it in a strategy
- Freeze inputs at the stated decision time and record their units.
- Create an interpretable text baseline and a debugging fixture before training a learned classifier.
- Recompute the example, then change the material assumption and explain the difference.
Research deliverable
A reproducible dictionary score for text: produce the worked calculation, a timestamped input record and a written decision addressing this limitation: Dictionary counts omit negation, context, quotation and semantic shifts; a score is not automatically a tradable forecast.
These are synthetic mechanics examples, not historical performance or paper replications. Module evidence and research boundaries record the 12 September 2026 review.
Python implementation
Self-contained teaching example. Python 3.10+; dependencies and input conventions are shown in the code and notation. Run in your own Python environment.
# Python 3.10+; standard library and NumPy only.
# Synthetic teaching inputs; conventions and units are defined in the notation above.
import re
def dictionary_score(document,positive,negative):
pos=set(positive); neg=set(negative)
if pos&neg: raise ValueError('Polarity dictionaries must be disjoint')
tokens=re.findall(r'[a-z]+',document.lower())
if not tokens: raise ValueError('No tokens')
return (sum(t in pos for t in tokens)-sum(t in neg for t in tokens))/len(tokens)
assert dictionary_score('profit improves but risk rises',{'profit','improves'},{'risk'})==.2
print(dictionary_score('profit improves but risk rises',{'profit','improves'},{'risk'}))Continue learning
Alternative Data: Measurement, Text & Incremental Value — all lessons- From a sampled panel to a population estimate
- A reproducible dictionary score for text
- An event return needs a predeclared benchmark
- Availability delays and signal decay
- Noisy proxies and attenuation
- Measure improvement against a frozen baseline
- From forecast accuracy to a costed decision
- A data investment includes coverage, access and ongoing costs
Quantitative finance and development glossary · Python resources and libraries · Research sources and limitations