FIN 550 — Course Map (Box sync)

Canonical source: Course Map.xlsx (Box file 2228694659133) Last synced: 2026-09-28 (auto, scripts/box-autosync.py) · Box modified 2026-09-27 18:24 CDT · sha1 d6a1a810e4f7 Editing policy: do not edit this file directly; edit the Excel in Box. It is regenerated whenever the Box file changes.

Owned by the learning-design team and faculty. For grading weights, schedule and assessment details, the course page is authoritative; this map records outcomes, topics and module-level alignment.

Course description

FIN 550: Big Data Analytics in Finance (ML I). Fall 2026, eight weeks, online. Students extract signals from structured and unstructured data, build predictive models, evaluate them out of sample, and communicate results with appropriate caveats. Financial markets provide the main laboratory, with applications to credit and real estate. Google Colab, Google Drive, and GitHub support the coursework. AI assistance is expected, with student understanding and verification required. Assessment combines five group exercises (E1-E5), one team project with a midpoint feasibility checkpoint and final submission, and peer review.

Course-level learning outcomes (CLOs)

ID Outcome
CLO 1 Extract signals from structured data (accounting ratios, trading records) and unstructured data (corporate filings, sentiment analysis)
CLO 2 Build predictive models using logistic regression, Lasso, random forests, and XGBoost
CLO 3 Evaluate models honestly using time-series cross-validation, out-of-sample testing, and subsample robustness checks
CLO 4 Detect common traps: look-ahead bias, survivorship bias, overfitting, data mining, and selection effects
CLO 5 Communicate results with appropriate caveats: what works, what doesn’t, and why

Assessment strategy (as entered in the Course Map)

Assessment type Aligned CLOs % of total grade
Group exercises E1-E5 (8% each) CLO 1-5 0.4
Team project: midpoint feasibility checkpoint CLO 1, 3, 4, 5 0.15
Team project: final submission CLO 1-5 0.35
Peer reviews given (5% + 5%) CLO 3, 4, 5 0.1
Participation (not graded) CLO 1-5 —
Canvas quizzes (tentative; no weight set) CLO 1-5, as offered —
E1 diversification-curve extension (optional) CLO 1, 3 —

Course topics map

Module Title Topic Aligned CLOs
1 Returns and market models M1.1: Stock Returns: Compounding, Beta, and t-Tests CLO 1, CLO 3
1 Returns and market models M1.2: Portfolios: Weights, Benchmarks, and Alpha CLO 1, CLO 3
2 Event studies: short-run reaction, then long-run performance M1.3: Event Studies: Short-Run Reactions CLO 1, CLO 3, CLO 4
2 Event studies: short-run reaction, then long-run performance M1.4: Event Studies: Long-Run Performance CLO 3, CLO 5
3 Momentum, from construction through evaluation M2.1: Momentum: Building the Portfolios CLO 1, CLO 4
3 Momentum, from construction through evaluation M2.2: Momentum: Evaluating the Payoff CLO 3, CLO 5
3 Momentum, from construction through evaluation M2.3: Momentum: Pressure-Testing the Payoff CLO 4, CLO 3
4 Text measurement M3.1: Text as Data: Building a Reproducible Measure CLO 1, CLO 4
4 Text measurement M3.2: Text as Data: Measuring Exposure in Context CLO 1, CLO 5
4 Text measurement M3.3: Text as Data: Predicting Forward Returns CLO 1, CLO 3, CLO 4

MOOC 2

Module Title Topic Aligned CLOs
1 Classification and next-month ranking M4.1: Classification: Probability Models for Ranking CLO 2, CLO 4
1 Classification and next-month ranking M4.2: Classification: Out-of-Sample Portfolio Returns CLO 3, CLO 2
1 Classification and next-month ranking M4.3: Classification: Rare Events, Error Costs, and Timing CLO 3, CLO 4
2 Time-safe validation, regularization, and trees M5.1: Validation: Time-Ordered Folds and Data Cleaning CLO 3, CLO 4
2 Time-safe validation, regularization, and trees M5.2: Lasso: Choosing Among Many Signals CLO 2, CLO 4
2 Time-safe validation, regularization, and trees M5.3: Lasso: Forecasting Returns Out of Sample CLO 3, CLO 5
2 Time-safe validation, regularization, and trees M6.1: Tree Models: Splits, Bagging, and Random Forests CLO 2, CLO 4
3 Boosting, neural networks, and why one benchmark is not enough M6.2: Tree Models: Boosting and a Fair Comparison CLO 2, CLO 3
3 Boosting, neural networks, and why one benchmark is not enough M6.3: Neural Networks: Building and Comparing Forecasts CLO 3, CLO 2
3 Boosting, neural networks, and why one benchmark is not enough M7.1: Factor Models: Why CAPM Falls Short CLO 3, CLO 5
4 Factor models, robustness, and the capstone M7.2: Factor Models: Loadings, Alpha, and Benchmarks CLO 3, CLO 5
4 Factor models, robustness, and the capstone M8.2: Robustness: Later Data, Delays, and Many Tests CLO 3, CLO 4
4 Factor models, robustness, and the capstone M8.3: Tail Risk: Drawdowns, Crashes, and a Recommendation CLO 5, CLO 3

Module 1 map — Returns and market models

Objective 1.1 (CLO 1): Compute simple, log, and cumulative returns, estimate a stock’s market-model beta, and use a t-test to judge whether an estimate differs from a benchmark value. Also supports: CLO 3

Objective 1.2 (CLO 1): Compare equal- and prior-period value-weighted portfolio returns, and interpret a raw-return market-model alpha, explaining why a significant estimate cannot distinguish luck, mispricing, and risk the model omits. Also supports: CLO 3

Module 2 map — Event studies: short-run reaction, then long-run performance

Objective 2.1 (CLO 1): Frame an event study around market efficiency, construct market-adjusted abnormal and cumulative abnormal returns, and test whether announcement-window evidence survives dependence checks and equal-length pre-announcement comparisons. Also supports: CLO 3, CLO 4

Objective 2.2 (CLO 3): Construct 12-month buy-and-hold abnormal returns and calendar-time portfolio returns for the same events, account for skewness when testing mean BHAR, and explain why the two long-horizon estimates can disagree. Also supports: CLO 5

Module 3 map — Momentum, from construction through evaluation

Objective 3.1 (CLO 1): Build a point-in-time 12-2 momentum strategy (signal timeline, eligibility, quintile membership, one-month returns) and explain what diversification does and does not remove. Also supports: CLO 4

Objective 3.2 (CLO 3): Evaluate momentum portfolios using the quintile profile, economic magnitude and statistical precision, and CAPM beta and alpha; explain what benchmark-adjusted performance does and does not establish. Also supports: CLO 5

Objective 3.3 (CLO 4): Compare stock-weighting rules, test whether momentum survives a size split, and identify how look-ahead, future-availability, and survivorship errors distort a backtest. Also supports: CLO 3

Module 4 map — Text measurement

Objective 4.1 (CLO 1): Build an auditable text measure from a dated 10-K passage with a finance-specific dictionary, and determine when that measure becomes usable as a signal. Also supports: CLO 4

Objective 4.2 (CLO 1): Construct keyword and dictionary exposure measures and an evidence-grounded language-model extraction, and distinguish textual prominence from economic exposure. Also supports: CLO 5

Objective 4.3 (CLO 1): Merge dated text measures to subsequent returns through an audited identifier link, test whether the signal predicts returns, and distinguish that association from shock-window validation and from an investable-strategy claim. Also supports: CLO 3, CLO 4

Module 5 map — Classification and next-month ranking

Objective 5.1 (CLO 2): Define a next-month classification target using only information available at the decision date, and contrast linear probability and logistic models for ranking stocks. Also supports: CLO 4

Objective 5.2 (CLO 3): Rank stocks by out-of-sample predicted probability, evaluate the classifier on a temporal holdout, and judge whether the ranking produces a portfolio payoff. Also supports: CLO 2

Objective 5.3 (CLO 3): Choose a classification threshold from explicit error costs for rare events, and verify that a text feature was available before it enters a model. Also supports: CLO 4

Module 6 map — Time-safe validation, regularization, and trees

Objective 6.1 (CLO 3): Explain why random cross-validation leaks future information, and build expanding-window validation with data cleaning fitted inside each training fold. Also supports: CLO 4

Objective 6.2 (CLO 2): Fit Lasso with time-safe validation to select among many candidate signals, interpret the coefficient path, and distinguish regularized prediction from inference. Also supports: CLO 4

Objective 6.3 (CLO 3): Interpret date-blocked out-of-sample return forecasts and prediction-sorted portfolio returns, identify the benchmark comparison still needed for an accuracy claim, and explain why selected coefficients do not establish causality. Also supports: CLO 5

Objective 6.4 (CLO 2): Explain how decision trees split and overfit, and how bagging and random forests reduce prediction variance. Also supports: CLO 4

Module 7 map — Boosting, neural networks, and why one benchmark is not enough

Objective 7.1 (CLO 2): Explain gradient boosting as sequential error correction, distinguish time-safe tuning from a fixed-specification teaching example, and compare boosted-tree and Lasso forecasts on the same held-out observations. Also supports: CLO 3

Objective 7.2 (CLO 3): Describe what a shallow neural network can represent, and compare linear, regularized, tree-based, and neural forecasts in a common horse race. Also supports: CLO 2

Objective 7.3 (CLO 3): Explain why the CAPM can be incomplete, construct the size and value factors, and separate factor decomposition from a risk-versus-mispricing interpretation. Also supports: CLO 5

Module 8 map — Factor models, robustness, and the capstone

Objective 8.1 (CLO 3): Run CAPM, Fama-French three-factor, and Carhart regressions on a common sample, and interpret loadings and alpha relative to each benchmark. Also supports: CLO 5

Objective 8.2 (CLO 3): Stress-test a strategy across later data, execution delay, and multiple testing, and report what each robustness check does and does not establish. Also supports: CLO 4

Objective 8.3 (CLO 5): Measure drawdowns and tail losses, and turn the course’s evidence into a conditional recommendation with explicit limits and reversal conditions. Also supports: CLO 3


Auto-generated by scripts/box-autosync.py from the canonical Box Course Map. To change content, edit the Excel in Box; this file refreshes on the next sync.