FIN 550 — Course Map (Box sync)
Canonical source: Course Map.xlsx (Box file 2228694659133)
Last synced: 2026-09-28 (auto, scripts/box-autosync.py) · Box modified 2026-09-27 18:24 CDT · sha1 d6a1a810e4f7
Editing policy: do not edit this file directly; edit the Excel in Box. It is regenerated whenever the Box file changes.
Owned by the learning-design team and faculty. For grading weights, schedule and assessment details, the course page is authoritative; this map records outcomes, topics and module-level alignment.
Course description
FIN 550: Big Data Analytics in Finance (ML I). Fall 2026, eight weeks, online. Students extract signals from structured and unstructured data, build predictive models, evaluate them out of sample, and communicate results with appropriate caveats. Financial markets provide the main laboratory, with applications to credit and real estate. Google Colab, Google Drive, and GitHub support the coursework. AI assistance is expected, with student understanding and verification required. Assessment combines five group exercises (E1-E5), one team project with a midpoint feasibility checkpoint and final submission, and peer review.
Course-level learning outcomes (CLOs)
| ID | Outcome |
|---|---|
| CLO 1 | Extract signals from structured data (accounting ratios, trading records) and unstructured data (corporate filings, sentiment analysis) |
| CLO 2 | Build predictive models using logistic regression, Lasso, random forests, and XGBoost |
| CLO 3 | Evaluate models honestly using time-series cross-validation, out-of-sample testing, and subsample robustness checks |
| CLO 4 | Detect common traps: look-ahead bias, survivorship bias, overfitting, data mining, and selection effects |
| CLO 5 | Communicate results with appropriate caveats: what works, what doesn’t, and why |
Assessment strategy (as entered in the Course Map)
| Assessment type | Aligned CLOs | % of total grade |
|---|---|---|
| Group exercises E1-E5 (8% each) | CLO 1-5 | 0.4 |
| Team project: midpoint feasibility checkpoint | CLO 1, 3, 4, 5 | 0.15 |
| Team project: final submission | CLO 1-5 | 0.35 |
| Peer reviews given (5% + 5%) | CLO 3, 4, 5 | 0.1 |
| Participation (not graded) | CLO 1-5 | — |
| Canvas quizzes (tentative; no weight set) | CLO 1-5, as offered | — |
| E1 diversification-curve extension (optional) | CLO 1, 3 | — |
Course topics map
| Module | Title | Topic | Aligned CLOs |
|---|---|---|---|
| 1 | Returns and market models | M1.1: Stock Returns: Compounding, Beta, and t-Tests | CLO 1, CLO 3 |
| 1 | Returns and market models | M1.2: Portfolios: Weights, Benchmarks, and Alpha | CLO 1, CLO 3 |
| 2 | Event studies: short-run reaction, then long-run performance | M1.3: Event Studies: Short-Run Reactions | CLO 1, CLO 3, CLO 4 |
| 2 | Event studies: short-run reaction, then long-run performance | M1.4: Event Studies: Long-Run Performance | CLO 3, CLO 5 |
| 3 | Momentum, from construction through evaluation | M2.1: Momentum: Building the Portfolios | CLO 1, CLO 4 |
| 3 | Momentum, from construction through evaluation | M2.2: Momentum: Evaluating the Payoff | CLO 3, CLO 5 |
| 3 | Momentum, from construction through evaluation | M2.3: Momentum: Pressure-Testing the Payoff | CLO 4, CLO 3 |
| 4 | Text measurement | M3.1: Text as Data: Building a Reproducible Measure | CLO 1, CLO 4 |
| 4 | Text measurement | M3.2: Text as Data: Measuring Exposure in Context | CLO 1, CLO 5 |
| 4 | Text measurement | M3.3: Text as Data: Predicting Forward Returns | CLO 1, CLO 3, CLO 4 |
MOOC 2
| Module | Title | Topic | Aligned CLOs |
|---|---|---|---|
| 1 | Classification and next-month ranking | M4.1: Classification: Probability Models for Ranking | CLO 2, CLO 4 |
| 1 | Classification and next-month ranking | M4.2: Classification: Out-of-Sample Portfolio Returns | CLO 3, CLO 2 |
| 1 | Classification and next-month ranking | M4.3: Classification: Rare Events, Error Costs, and Timing | CLO 3, CLO 4 |
| 2 | Time-safe validation, regularization, and trees | M5.1: Validation: Time-Ordered Folds and Data Cleaning | CLO 3, CLO 4 |
| 2 | Time-safe validation, regularization, and trees | M5.2: Lasso: Choosing Among Many Signals | CLO 2, CLO 4 |
| 2 | Time-safe validation, regularization, and trees | M5.3: Lasso: Forecasting Returns Out of Sample | CLO 3, CLO 5 |
| 2 | Time-safe validation, regularization, and trees | M6.1: Tree Models: Splits, Bagging, and Random Forests | CLO 2, CLO 4 |
| 3 | Boosting, neural networks, and why one benchmark is not enough | M6.2: Tree Models: Boosting and a Fair Comparison | CLO 2, CLO 3 |
| 3 | Boosting, neural networks, and why one benchmark is not enough | M6.3: Neural Networks: Building and Comparing Forecasts | CLO 3, CLO 2 |
| 3 | Boosting, neural networks, and why one benchmark is not enough | M7.1: Factor Models: Why CAPM Falls Short | CLO 3, CLO 5 |
| 4 | Factor models, robustness, and the capstone | M7.2: Factor Models: Loadings, Alpha, and Benchmarks | CLO 3, CLO 5 |
| 4 | Factor models, robustness, and the capstone | M8.2: Robustness: Later Data, Delays, and Many Tests | CLO 3, CLO 4 |
| 4 | Factor models, robustness, and the capstone | M8.3: Tail Risk: Drawdowns, Crashes, and a Recommendation | CLO 5, CLO 3 |
Module 1 map — Returns and market models
Objective 1.1 (CLO 1): Compute simple, log, and cumulative returns, estimate a stock’s market-model beta, and use a t-test to judge whether an estimate differs from a benchmark value. Also supports: CLO 3
- Guiding questions: Does a 10% Gain Undo a 10% Loss?; Does a Bumpier Path Cost You Anything?; Does Amazon Move With the Market, More Than One-for-One?; When Is a Difference Big Enough to Believe?
- Assessments and supporting activities: W1: team formation and signal ideas (ungraded). E1 due before W3.; E1: Diversification and the Market Model (graded group exercise); Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M1L1.1; M1L1.2; M1L1.3; M1L1.4
Objective 1.2 (CLO 1): Compare equal- and prior-period value-weighted portfolio returns, and interpret a raw-return market-model alpha, explaining why a significant estimate cannot distinguish luck, mispricing, and risk the model omits. Also supports: CLO 3
- Guiding questions: Should Bigger Companies Count for More?; What Counts as Beating the Market?; Does a Positive Alpha Prove Skill?
- Assessments and supporting activities: E1: Diversification and the Market Model (graded group exercise); Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M1L2.1; M1L2.2; M1L2.3
Module 2 map — Event studies: short-run reaction, then long-run performance
Objective 2.1 (CLO 1): Frame an event study around market efficiency, construct market-adjusted abnormal and cumulative abnormal returns, and test whether announcement-window evidence survives dependence checks and equal-length pre-announcement comparisons. Also supports: CLO 3, CLO 4
- Guiding questions: Is There Anything Left to Trade On?; What Do You Fix Before You Look?; How Much of the Move Was Exxon?; What Would No Effect Even Look Like?; Is It Really the Dividend?
- Assessments and supporting activities: W2 tutorial: event returns and information timing. E2 due before W4.; E2: Event Study on Stock-Split Announcements; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M1L3.1; M1L3.2; M1L3.3; M1L3.4; M1L3.5
- Suggested readings: Fama (1970), Efficient Capital Markets: A Review of Theory and Empirical Work, Journal of Finance; Huberman and Regev (2001), Contagious Speculation and a Cure for Cancer: A Nonevent That Made Stock Prices Soar, Journal of Finance
Objective 2.2 (CLO 3): Construct 12-month buy-and-hold abnormal returns and calendar-time portfolio returns for the same events, account for skewness when testing mean BHAR, and explain why the two long-horizon estimates can disagree. Also supports: CLO 5
- Guiding questions: Does the Stock Keep Outperforming After the News?; What Was the Portfolio Holding Each Month?; Do the Two Long-Run Methods Agree?
- Assessments and supporting activities: No exercise yet practises 12-month BHAR or the BHAR vs calendar-time comparison; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M1L4.1; M1L4.2; M1L4.3
Module 3 map — Momentum, from construction through evaluation
Objective 3.1 (CLO 1): Build a point-in-time 12-2 momentum strategy (signal timeline, eligibility, quintile membership, one-month returns) and explain what diversification does and does not remove. Also supports: CLO 4
- Guiding questions: Why Hold Portfolios?; How Do We Build a Momentum Strategy?; What Did the Sorted Portfolios Earn?
- Assessments and supporting activities: W3 tutorial: equal- and value-weight portfolio formation. E1 due before W3.; E5: Momentum Strategies and Factor Models; Optional E1 extension: diversification curve (after M2.1); Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M2L1.1; M2L1.2; M2L1.3
- Suggested readings: Jegadeesh and Titman (1993), Returns to Buying Winners and Selling Losers: Implications for Stock Market Efficiency, Journal of Finance
Objective 3.2 (CLO 3): Evaluate momentum portfolios using the quintile profile, economic magnitude and statistical precision, and CAPM beta and alpha; explain what benchmark-adjusted performance does and does not establish. Also supports: CLO 5
- Guiding questions: How Much Does Momentum Earn?; How Much Risk Comes With the Payoff?; How Much Is Compensation for Market Exposure?; Momentum’s CAPM Alpha
- Assessments and supporting activities: E5: Momentum Strategies and Factor Models; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M2L2.1; M2L2.2; M2L2.3; M2L2.4
Objective 3.3 (CLO 4): Compare stock-weighting rules, test whether momentum survives a size split, and identify how look-ahead, future-availability, and survivorship errors distort a backtest. Also supports: CLO 3
- Guiding questions: How Should We Weight Each Stock?; Does Momentum Depend on Weighting and Size?; Could We Have Run This Strategy in Real Time?
- Assessments and supporting activities: E5: Momentum Strategies and Factor Models; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M2L3.1; M2L3.2; M2L3.3
Module 4 map — Text measurement
Objective 4.1 (CLO 1): Build an auditable text measure from a dated 10-K passage with a finance-specific dictionary, and determine when that measure becomes usable as a signal. Also supports: CLO 4
- Assessments and supporting activities: Midpoint feasibility checkpoint: data access, inspected sample and evaluation plan. Slides due W4; 15%.; E3: Short Interest and Text Classification; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M3L1.1; M3L1.2; M3L1.3
- Suggested readings: Loughran and McDonald (2011), When Is a Liability Not a Liability? Textual Analysis, Dictionaries, and 10-Ks, Journal of Finance
Objective 4.2 (CLO 1): Construct keyword and dictionary exposure measures and an evidence-grounded language-model extraction, and distinguish textual prominence from economic exposure. Also supports: CLO 5
- Assessments and supporting activities: E3: Short Interest and Text Classification; No exercise yet practises keyword measures or language-model extraction; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M3L2.1; M3L2.2; M3L2.3
- Suggested readings: Loughran and McDonald (2011), When Is a Liability Not a Liability? Textual Analysis, Dictionaries, and 10-Ks, Journal of Finance; Glasserman and Lin (2023)
Objective 4.3 (CLO 1): Merge dated text measures to subsequent returns through an audited identifier link, test whether the signal predicts returns, and distinguish that association from shock-window validation and from an investable-strategy claim. Also supports: CLO 3, CLO 4
- Assessments and supporting activities: E3: Short Interest and Text Classification; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M3L3.1; M3L3.2; M3L3.3
Module 5 map — Classification and next-month ranking
Objective 5.1 (CLO 2): Define a next-month classification target using only information available at the decision date, and contrast linear probability and logistic models for ranking stocks. Also supports: CLO 4
- Assessments and supporting activities: W5: proposal overflow if needed; two proposal peer reviews due (5%). E3 due before W6.; E3: Short Interest and Text Classification; W5-W6 tutorial: compare rule, logit, L1 and forest predictions; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M4L1.1; M4L1.2; M4L1.3; M4L1.4
Objective 5.2 (CLO 3): Rank stocks by out-of-sample predicted probability, evaluate the classifier on a temporal holdout, and judge whether the ranking produces a portfolio payoff. Also supports: CLO 2
- Assessments and supporting activities: E3: Short Interest and Text Classification; W5-W6 tutorial: compare rule, logit, L1 and forest predictions; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M4L2.1; M4L2.2; M4L2.3
Objective 5.3 (CLO 3): Choose a classification threshold from explicit error costs for rare events, and verify that a text feature was available before it enters a model. Also supports: CLO 4
- Assessments and supporting activities: E3: Short Interest and Text Classification; W5-W6 tutorial: compare rule, logit, L1 and forest predictions; No exercise yet practises cost-based thresholds for rare events; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M4L3.1; M4L3.2; M4L3.3
Module 6 map — Time-safe validation, regularization, and trees
Objective 6.1 (CLO 3): Explain why random cross-validation leaks future information, and build expanding-window validation with data cleaning fitted inside each training fold. Also supports: CLO 4
- Assessments and supporting activities: W6 tutorial: fair model comparison. E4 due before W7.; E4: Accounting Anomalies with Lasso and Random Forests; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M5L1.1; M5L1.2; M5L1.3
Objective 6.2 (CLO 2): Fit Lasso with time-safe validation to select among many candidate signals, interpret the coefficient path, and distinguish regularized prediction from inference. Also supports: CLO 4
- Assessments and supporting activities: E4: Accounting Anomalies with Lasso and Random Forests; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M5L2.1; M5L2.2; M5L2.3
- Suggested readings: Jensen, Kelly, and Pedersen (2023), Is There a Replication Crisis in Finance?, Journal of Finance
Objective 6.3 (CLO 3): Interpret date-blocked out-of-sample return forecasts and prediction-sorted portfolio returns, identify the benchmark comparison still needed for an accuracy claim, and explain why selected coefficients do not establish causality. Also supports: CLO 5
- Assessments and supporting activities: E4: Accounting Anomalies with Lasso and Random Forests; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M5L3.1; M5L3.2
Objective 6.4 (CLO 2): Explain how decision trees split and overfit, and how bagging and random forests reduce prediction variance. Also supports: CLO 4
- Assessments and supporting activities: E4: Accounting Anomalies with Lasso and Random Forests; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M6L1.1; M6L1.2; M6L1.3
Module 7 map — Boosting, neural networks, and why one benchmark is not enough
Objective 7.1 (CLO 2): Explain gradient boosting as sequential error correction, distinguish time-safe tuning from a fixed-specification teaching example, and compare boosted-tree and Lasso forecasts on the same held-out observations. Also supports: CLO 3
- Assessments and supporting activities: Final presentations W7-W8. All teams may revise analysis and slides through the common W8 deadline.; Optional E4 extension: XGBoost comparison; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M6L2.1; M6L2.2; M6L2.3
Objective 7.2 (CLO 3): Describe what a shallow neural network can represent, and compare linear, regularized, tree-based, and neural forecasts in a common horse race. Also supports: CLO 2
- Assessments and supporting activities: No exercise yet practises neural networks or the model horse race; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M6L3.1; M6L3.2; M6L3.3
- Suggested readings: Gu, Kelly, and Xiu (2020), Empirical Asset Pricing via Machine Learning, Review of Financial Studies
Objective 7.3 (CLO 3): Explain why the CAPM can be incomplete, construct the size and value factors, and separate factor decomposition from a risk-versus-mispricing interpretation. Also supports: CLO 5
- Assessments and supporting activities: E5: Momentum Strategies and Factor Models; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M7L1.1; M7L1.2
- Suggested readings: Fama and French (1993), Common Risk Factors in the Returns on Stocks and Bonds, Journal of Financial Economics
Module 8 map — Factor models, robustness, and the capstone
Objective 8.1 (CLO 3): Run CAPM, Fama-French three-factor, and Carhart regressions on a common sample, and interpret loadings and alpha relative to each benchmark. Also supports: CLO 5
- Assessments and supporting activities: Final revised slides + notes due W8 (35%); final peer reviews end W8 (5%). E5 due before W8.; E5: Momentum Strategies and Factor Models; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M7L2.1; M7L2.2
- Instructional materials: Fama and French (1993), Common Risk Factors in the Returns on Stocks and Bonds, Journal of Financial Economics; Carhart (1997), On Persistence in Mutual Fund Performance, Journal of Finance
Objective 8.2 (CLO 3): Stress-test a strategy across later data, execution delay, and multiple testing, and report what each robustness check does and does not establish. Also supports: CLO 4
- Assessments and supporting activities: Final project: robustness evidence; Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M8L2.1; M8L2.2; M8L2.3
- Instructional materials: McLean and Pontiff (2016), Does Academic Research Destroy Stock Return Predictability?, Journal of Finance; Martineau (2022), Rest in Peace Post-Earnings Announcement Drift, Critical Finance Review; Harvey, Liu, and Zhu (2016), … and the Cross-Section of Expected Returns, Review of Financial Studies Jensen, Kelly, and Pedersen (2023), Is There a Replication Crisis in Finance?, Journal of Finance
Objective 8.3 (CLO 5): Measure drawdowns and tail losses, and turn the course’s evidence into a conditional recommendation with explicit limits and reversal conditions. Also supports: CLO 3
- Assessments and supporting activities: Final project: presentation, slides + speaker notes (no separate writeup); Possible Canvas quiz (not scheduled; no grade weight set)
- Multimedia: M8L3.1; M8L3.2; M8L3.3
- Instructional materials: Daniel and Moskowitz (2016), Momentum Crashes, Journal of Financial Economics
Auto-generated by scripts/box-autosync.py from the canonical Box Course Map. To change content, edit the Excel in Box; this file refreshes on the next sync.