Preregistered primary prediction
Ridge · realized_volatility_1_20
Raw p-value 0.00067; supports exploratory volatility-ranking evidence.
Every claim must be linked to a timestamped, testable artifact.
Filings, labels, splits, features, models, predictions, inference, portfolio diagnostics, and audit checks form one evidence chain.
50_company_public_fmp_alpha_2016_2025_v4
Leakage-safe research design
Every text feature and model decision must be available before the prediction timestamp.
Key controls
| Control | Implementation |
|---|---|
| Event-time alignment | Filing timestamps define information availability. |
| Rolling OOS design | Train / validation / test windows roll through time. |
| Leakage control | Embargo purge and split-leakage logs. |
| Model selection | Validation-only Rank IC. |
| TF-IDF control | Train-window-only vocabulary fitting. |
| Incremental text diagnostic | Industry-neutral Rank IC and feature ablation. |
| Statistical uncertainty | Newey-West and clustered bootstrap confidence intervals. |
| Data snooping | Specification registry and multiple-testing report. |
| Parser quality | Manual review appendix for short or malformed sections. |
Feature construction
Loughran-McDonald dictionary tone and TF-IDF/SVD are built over full filing, Business, Risk Factors, Legal Proceedings, and MD&A scopes.
| Feature set | Meaning |
|---|---|
industry_only | Training-window industry-mean baseline. |
dictionary_only | Dictionary-tone text features. |
tfidf_svd_only | TF-IDF/SVD text representations. |
combined_text | Combined dictionary and text representation. |
industry_plus_text | Industry features plus text features. |
Industry-neutral Rank IC is a descriptive diagnostic, not a causal decomposition.
Bootstrap inference
Inconclusive for v4 because there are only four OOS split clusters.
Supports a positive raw primary Rank IC interval.
Supports a positive raw primary Rank IC interval.
Positive point estimate, but not bootstrap-robust.
Parser quality review
Item 1A and Item 7 below 100 words are excluded from section-level features. Core sections from 100 to 499 words remain included but carry a warning.
Evidence boundary
Preregistered primary prediction
realized_volatility_1_20Raw p-value 0.00067; supports exploratory volatility-ranking evidence.
Preregistered primary portfolio
Raw p-value 0.1147; does not establish tradable alpha.
Formal empirical-finance claims are blocked by data-boundary issues, not pipeline failures: mixed FMP/Yahoo data, applied-grade market-cap estimates, fixed 50-company panel, parser-quality limitations, and a small number of missing diagnostic model-label pairs.
The project should be interpreted as an applied-grade, auditable financial NLP workflow for exploratory volatility ranking.