FIN 550 - Big Data Analytics in Finance (ML I)
Program-level details: See program/curriculum.md
Instructor recording tool: Xing Gao and Mathias Kronlund plan to use Quarto for lecture recording production. Quarto is an instructor/production tool only — not part of the student tool stack.
Live LD team data: tracked in
fin550/sync/— auto-synced from Box ~4x/day (scripts/box-autosync.py). Canonical files:Course Map.xlsx+Instructional Activity Roster.xlsx(roster formerlyInstructional Material Roster.xlsx). As of 2026-05-29 both are still the blank LD template — no content snapshot generated yet; the sync auto-detects when Gao/Kronlund populate them.
| Credits: 4 | Term: Fall 2026 (Weeks 9-16) | Instructor: Xing Gao / Mathias Kronlund |
Course Vision
Students learn core machine learning methods — regression, classification, regularization, tree-based models, and neural networks — applied to real-world prediction problems drawn primarily from financial markets and institutions. Finance provides the running context because it offers the richest, most granular, and most publicly available datasets for learning these skills. The methods are universal; the applications are financial.
This is ML I in the MSBAi sequence. FIN 550 focuses on supervised learning: building, evaluating, and selecting models for prediction. BADM 576 (ML II) extends to unsupervised learning, NLP, time series, deep learning, and deployment.
Domain perspective: Each MSBAi core course brings a distinct business lens. FIN 550 brings the finance and accounting perspective — students work with stock returns, firm financial statements, corporate bonds, mutual funds, and market text data. No prior finance knowledge is assumed; financial concepts are introduced as needed.
Prerequisites
- Python programming (from BADM 554 or equivalent)
- Statistics foundation: FIN 550 assumes working knowledge of descriptive statistics, probability, sampling, hypothesis testing, and regression. The program offers two Gies-on-Coursera preparatory courses as self-diagnostic stats prep, sequenced to be completed before FIN 550 begins (Fall 2026 Week 9):
- Exploring and Producing Data for Business Decision Making (University of Illinois)
- Inferential and Predictive Statistics for Business (University of Illinois)
- Per the program-wide staggered prep model (2026-05-18), these courses are self-diagnostic — students use the quiz to identify whether they should work through the prep content. Scores are not evaluated by the program. See program/curriculum.md for the full prep stack.
Learning Outcomes (L-C-E Framework)
Literacy (Foundational Awareness)
- L1: Understand supervised learning and explain when regression vs. classification applies
- L2: Explain how financial data (returns, firm characteristics, corporate filings) is structured and used for prediction
- L3: Recognize overfitting, describe train/test/validation splits, and explain why cross-validation matters
Competency (Applied Skills)
- C1: Build and evaluate regression models (linear, Fama-MacBeth, Lasso) using financial and accounting data
- C2: Build and evaluate classification models (logistic regression, decision trees) for business outcomes
- C3: Extract predictive features from structured data (firm characteristics) and unstructured text (corporate filings)
- C4: Apply cross-validation, regularization, and hyperparameter tuning to improve model performance
Expertise (Advanced Application)
- E1: Design and execute a multi-step empirical analysis from data to portfolio strategy to performance evaluation
- E2: Apply ensemble methods (random forest, gradient boosting) and neural networks, choosing appropriately among model types
- E3: Translate predictive model results into actionable business recommendations through team presentations with live Q&A
Module-by-Module Breakdown
| Week | Module | Theme | Key methods |
|---|---|---|---|
| W1 | M1: Measuring Outcomes | Does a signal work? | Returns, OLS, benchmarking (CAPM alpha), event studies, market efficiency |
| W2 | M2: From Signal to Strategy | Turn prediction into action | Portfolio sorts, calendar-time evaluation, Sharpe ratio, long-short strategies |
| W3 | M3: Text Analytics | Signals from unstructured data | Dictionary sentiment, keyword exposure, FinBERT comparison against frozen outputs, data linkage |
| W4 | M4: Classification | Predict binary outcomes | Logistic regression, AUC, class imbalance, feature engineering |
| Working Pipeline Checkpoint | |||
| W5 | M5: Model Evaluation + Feature Selection | Honest assessment | Time-series cross-validation (backtesting), Lasso, variable selection, data cleaning |
| W6 | M6: Trees, Ensembles, Boosting | Nonlinear models | Decision trees, random forests, XGBoost, feedforward neural networks, hyperparameter tuning |
| W7 | M7: Benchmarking + Attribution | Proper risk adjustment | CAPM, FF3, Carhart, FF5, Fama-MacBeth |
| W8 | M8: From Model to Decision | Real-world constraints | Transaction costs, out-of-sample decay, robustness, downside risk, data mining |
By the end of W4 students have a working pipeline — signal → features → classification → portfolio → performance evaluation. W5–W8 improve it: honest evaluation, nonlinear models, risk adjustment, and real-world constraints.
NLP boundary note: M3 introduces contextual financial-language classification using FinBERT — but students do not run the model. Outputs are precomputed and frozen, so students compare frozen FinBERT labels against dictionary-based labels. Running transformer models (BERT/embeddings as active skills) is covered in BADM 576 (ML II).
Project (Team, runs across all eight modules)
Teams of three run one project across all eight modules, in one of two forms:
Type A — firm-specific signal. A measurable characteristic that predicts outcomes: insider trading patterns, analyst revisions, patent filings, ESG ratings, governance changes, attention-driven retail trading, institutional fire sales.
Type B — differential exposure signal. Entities differentially exposed to a macro theme — tariffs, climate regulation, AI disruption, interest rates, supply chain shocks — identified using text analysis, then tested for differential effects when relevant events occur.
Projects are not limited to equities. Credit risk, real estate, ESG, and other domains are welcome where the data supports the pipeline. Where an equity project uses “portfolio sorts → alpha,” a credit project uses “risk buckets → default prediction accuracy” and a real estate project uses “neighborhood segments → pricing error.” The requirements — honest out-of-sample evaluation, appropriate cross-validation, robustness checks — are identical across domains.
Teams present progress during studio sessions and submit the final project deliverable in M8.
Assessment
One major team project (the Trading Signal System) runs alongside the course, with four exercises (E1–E4) building foundational skills and project milestones scaffolding toward the final deliverable. No traditional exam.
Exercises
Decision-point notebooks: students are told what to do, not how, and each exercise ends with interpretation questions. Designed for approximately 2 hours with AI assistance, which is assumed rather than tolerated. Work is submitted through Canvas; GitHub distributes course materials to students.
| Exercise | Released after | Due before | Topic |
|---|---|---|---|
| E0: Diversification Diagnostic | M1 | W2 | Returns → portfolios → market model. Required, ungraded. |
| E1: Stock Splits | M2 | W4 | Stock-split events → CARs → window sensitivity and cross-sectional tests |
| E2: Short Interest + Text | M4 | W6 | Short interest and text features → classification, train/test, AUC |
| E3: Accounting Anomalies | M6 | W7 | 15 accounting signals → Lasso vs. XGBoost, time-series CV |
| E4: Momentum + Factors | M7 | W8 | Momentum → CAPM vs. FF3 alpha, factor loadings |
Rubric (per exercise, 4 dimensions):
| Dimension | Excellent (A) | Proficient (B) | Developing (C) |
|---|---|---|---|
| Methodology | Correct application of module methods, justified choices | Reasonable approach with minor gaps | Flawed or missing methodology |
| Model Evaluation | Rigorous evaluation, explains metrics clearly | Evaluation applied but basic | Missing or weak evaluation |
| Code Quality | Clean, documented Jupyter notebook, reproducible | Adequate code, some comments | Messy or undocumented |
| Written Analysis | Connects methodology to findings; explains business implications | Adequate explanation | Minimal or unclear |
Project Proposal (20%)
Team project proposal presented in Week 4. Deliverable: presentation + slideshow with notes. One holistic grade. Instructor questions during the presentation are graded within the presentation grade.
Content:
- Signal hypothesis and rationale
- Data source identification and access plan
- Team roles and preliminary EDA
- Initial pipeline sketch
| Milestone checkpoint | Due | Purpose |
|---|---|---|
| M2: Signal Construction + Regression Baseline | End of Module 4 | Constructed signal, baseline regression model, initial cross-validation results — formative feedback only |
| M3: Model Expansion + Text Integration | End of Module 6 | Classification/text features added to pipeline — formative feedback only |
Rubric (3 dimensions):
| Dimension | Excellent (A) | Proficient (B) | Developing (C) |
|---|---|---|---|
| Progress | Substantial, on-track work building on prior modules | Adequate progress with some gaps | Behind schedule or superficial |
| Technical Quality | Methods applied correctly, evaluation included | Functional but basic analysis | Errors or missing components |
| Team Collaboration | Clear evidence of shared work, complementary contributions | Adequate collaboration | Uneven contribution |
Final Project (35%)
Complete the Trading Signal System and present in Weeks 7–8. Deliverable: presentation + slideshow with notes. One holistic grade. Instructor questions during the presentation are graded within the presentation grade.
Deliverables:
- Complete trading signal pipeline: signal discovery → prediction → portfolio → abnormal returns
- Apply tree-based/ensemble methods and compare with earlier regression approaches
- Presentation + slideshow with notes covering:
- Signal rationale and data source
- Prediction methodology and model comparison
- Portfolio strategy and performance analysis
- Limitations, risks, and out-of-sample considerations
- GitHub repo with all code + documentation
Rubric (4 dimensions):
| Dimension | Excellent (A) | Proficient (B) | Developing (C) |
|---|---|---|---|
| Signal Design | Creative, well-justified signal from interesting data source | Reasonable signal choice | Generic or unjustified |
| Model Pipeline | Multiple methods compared systematically, strong evaluation | Functional pipeline | Single method or weak evaluation |
| Portfolio Analysis | Rigorous abnormal return analysis, addresses look-ahead bias | Adequate portfolio construction | Flawed methodology |
| Presentation & Q&A | Confident delivery, articulates trade-offs, handles questions well | Adequate delivery, answers most questions | Unclear delivery or struggles with Q&A |
Studio Sessions (ungraded — output folds into milestones)
Weekly engagement in live studio sessions. Students work through guided exercises, discuss approaches, and share progress on project milestones. Studio output is captured in project milestone grades.
AI Tools Integration
Modules 1-3 (Regression and Evaluation):
- Use Claude/ChatGPT to:
- Explain finance concepts as encountered (returns, abnormal returns, factor models)
- Debug scikit-learn and statsmodels errors
- Interpret regression outputs and cross-validation results
- Suggest feature engineering approaches for financial data
Modules 4-6 (Classification and Text):
- Use AI to:
- Explain variable selection trade-offs (Lasso vs. stepwise)
- Generate text processing code (tokenization, sentiment)
- Debug logistic regression issues
- Review model comparison methodology
Modules 7-8 (Ensembles and Synthesis):
- Use AI to:
- Explain tree-based model hyperparameters
- Draft business case structure for trading signal report
- Review portfolio analysis methodology
- Practice Q&A scenarios for oral defense
AI Attribution: Students maintain a short project-level log of tools used. The AI policy grades the defensibility of the work; prompt-level disclosure is not required. See design/assessment_strategy.md for the attribution framework.
Assessment Summary
| Component | Weight | Notes |
|---|---|---|
| Exercises (E1–E4, graded holistically, 8.75% each) | 35% | Individual, submitted via Canvas. E0 (Diversification Diagnostic) is required/ungraded. |
| Project Proposal (presentation + slideshow with notes, week 4) | 20% | Team of 3, one holistic grade; Q&A graded within presentation |
| Final Project (presentation + slideshow with notes, weeks 7–8) | 35% | Team of 3, one holistic grade; Q&A graded within presentation |
| Peer Review (quality of reviews given) | 10% | 5% on proposals, 5% on final slideshows |
No traditional exam. Four exercises build skills; the team project integrates them. Live Q&A with the instructor is embedded in both presentations.
AI Usage Levels (AIAS)
| Assessment | AIAS Level | AI Permitted |
|---|---|---|
| Exercises (E1–E4) | 2 | AI for debugging, interpreting outputs, finance concept explanation — with attribution |
| Project Proposal | 2 | AI for code assistance, data exploration — with attribution |
| Final Project | 3 | AI as collaborator for business case drafting and model comparison — with full disclosure |
| Studio sessions (ungraded) | 1 | AI for concept exploration during exercises |
Technology Stack
- ML / statistics: scikit-learn, statsmodels, XGBoost, LightGBM (optional)
- Data: pandas, numpy, scipy, pyarrow, DuckDB (SQL-first for fetching, joins, filters, aggregation)
- Text: scikit-learn vectorizers and dictionary methods; FinBERT comparison against precomputed, frozen model outputs
- Visualization: matplotlib only
- Data sources: WRDS / Compustat / CRSP, SEC EDGAR, Kenneth French data library, yfinance
- Environment: Google Colab (primary — all coding, no local setup required); VS Code with an AI coding extension (supported for students who prefer it)
- Version control: GitHub, for distributing materials to students
Pedagogical Notes for Faculty
Design suggestions grounded in program research — not requirements. Adapt to your course and teaching style. Full references in reference/articles/.
The scenic route (cognitive friction) FIN 550 is where the tension between AI efficiency and learning depth is sharpest. Students can use Copilot to generate a random forest in seconds — but if they haven’t first built a baseline regression by hand (Modules 1-3), the ensemble result has no prediction to surprise them. The dopamine gap research shows the brain learns through prediction errors: the gap between what you expected and what happened. A student who struggles with linear regression before seeing how gradient boosting improves on it learns more than one who skips straight to XGBoost. The AIAS progression (0→1→2→3 across modules) already scaffolds this; the key is framing the early manual work as the investment that makes the later AI-assisted work register. → Machulla (2026), Schultz et al. (1997)
The IKEA effect (completion matters) The Trading Signal System runs across all 8 modules — this is the longest sustained project in the program. The IKEA effect research shows that labor leads to love only when it leads to completion. Each milestone (M1→M4) should feel like a working thing: a testable hypothesis, a running model, a pipeline that produces output. The oral defense in Module 8 is the ultimate completion signal. For career pivoters with no finance background, the moment they can explain their trading signal system to a panel is transformative — it’s when “I’m not a finance person” becomes “I built this.” → Norton, Mochon & Ariely (2012)
Variable uncertainty and calibrated difficulty The dopamine system is most engaged at ~50% uncertainty — when the student genuinely doesn’t know if they’ll succeed. Assignments that are too easy (certain success) or too hard (certain failure) produce flat responses. For career pivoters, Module 5 (Factor Models) is a known difficulty spike — the finance theory is unfamiliar. Consider front-loading the finance concepts students need (the “For career pivoters” notes are already in the right spirit) so the challenge is the ML application, not the domain vocabulary. The goal: students should feel “I might be able to do this” — not “I definitely can” or “I definitely can’t.” → Machulla (2026), Fiorillo, Tobler & Schultz (2003)
Three AI iterations before milestone submission For milestones M2-M4 (AIAS 2-3), consider requiring students to iterate with AI at least 3 times before submitting. This builds the habit of using AI as a thinking partner: first attempt → AI critique → revised attempt → AI alternative → final version with documented rationale. Produces richer AI Attribution Logs and prevents the “paste Copilot output, submit” pattern. → Means (2026, “Practice Gap”)
The Push-Back Protocol for model interpretation When students use AI to interpret regression outputs or explain cross-validation results (Modules 2-4), there’s a risk they accept AI’s explanation without verifying it against the data. The Push-Back Protocol (demand evidence → surface assumptions → request alternatives → stress-test → synthesize) is especially valuable here: ML model interpretation is exactly the kind of task where AI sounds confident but can be wrong. Consider making one exercise a structured push-back exercise. → Means (2025, “Push-Back Protocol”)
Peer review as milestone M4 The cross-team review at M4 is a strong design choice. The IKEA effect research suggests students learn as much from evaluating others’ pipelines as from building their own — exposure to different approaches calibrates their sense of quality. Consider structuring the peer review with the same rubric dimensions used for the final deliverable (signal design, model pipeline, portfolio analysis, written analysis) so students internalize the evaluation criteria before their own defense. → Norton et al. (2012), Li et al. (2020)
Attack your assessments Before the semester, have a confident AI user attempt each exercise and the trading signal project using current AI tools. AI is already strong at generating scikit-learn pipelines and interpreting regression output. Where can AI complete the task without genuine understanding of the finance context or model assumptions? Those are the spots to add pre-AI phases or shift weight toward the oral defense. → Furze (2026)
| Course Sequence: ← BDI 513 — Data Storytelling | Program sequence: BADM 558 — Big Data Infrastructures → | ML track: BADM 576 — Data Science and ML → |