FIN 550 - Big Data Analytics in Finance (ML I)

Program-level details: See program/curriculum.md

Instructor recording tool: Xing Gao and Mathias Kronlund plan to use Quarto for lecture recording production. Quarto is an instructor/production tool only — not part of the student tool stack.

Live LD team data: tracked in fin550/sync/ — auto-synced from Box ~4x/day (scripts/box-autosync.py). Canonical files: Course Map.xlsx + Instructional Activity Roster.xlsx (roster formerly Instructional Material Roster.xlsx). As of 2026-05-29 both are still the blank LD template — no content snapshot generated yet; the sync auto-detects when Gao/Kronlund populate them.

Credits: 4 Term: Fall 2026 (Weeks 9-16) Instructor: Xing Gao / Mathias Kronlund

Course Vision

Students learn core machine learning methods — regression, classification, regularization, tree-based models, and neural networks — applied to real-world prediction problems drawn primarily from financial markets and institutions. Finance provides the running context because it offers the richest, most granular, and most publicly available datasets for learning these skills. The methods are universal; the applications are financial.

This is ML I in the MSBAi sequence. FIN 550 focuses on supervised learning: building, evaluating, and selecting models for prediction. BADM 576 (ML II) extends to unsupervised learning, NLP, time series, deep learning, and deployment.

Domain perspective: Each MSBAi core course brings a distinct business lens. FIN 550 brings the finance and accounting perspective — students work with stock returns, firm financial statements, corporate bonds, mutual funds, and market text data. No prior finance knowledge is assumed; financial concepts are introduced as needed.

Prerequisites

Learning Outcomes (L-C-E Framework)

Literacy (Foundational Awareness)

Competency (Applied Skills)

Expertise (Advanced Application)

Module-by-Module Breakdown

Week Module Theme Key methods
W1 M1: Measuring Outcomes Does a signal work? Returns, OLS, benchmarking (CAPM alpha), event studies, market efficiency
W2 M2: From Signal to Strategy Turn prediction into action Portfolio sorts, calendar-time evaluation, Sharpe ratio, long-short strategies
W3 M3: Text Analytics Signals from unstructured data Dictionary sentiment, keyword exposure, FinBERT comparison against frozen outputs, data linkage
W4 M4: Classification Predict binary outcomes Logistic regression, AUC, class imbalance, feature engineering
    Working Pipeline Checkpoint  
W5 M5: Model Evaluation + Feature Selection Honest assessment Time-series cross-validation (backtesting), Lasso, variable selection, data cleaning
W6 M6: Trees, Ensembles, Boosting Nonlinear models Decision trees, random forests, XGBoost, feedforward neural networks, hyperparameter tuning
W7 M7: Benchmarking + Attribution Proper risk adjustment CAPM, FF3, Carhart, FF5, Fama-MacBeth
W8 M8: From Model to Decision Real-world constraints Transaction costs, out-of-sample decay, robustness, downside risk, data mining

By the end of W4 students have a working pipeline — signal → features → classification → portfolio → performance evaluation. W5–W8 improve it: honest evaluation, nonlinear models, risk adjustment, and real-world constraints.

NLP boundary note: M3 introduces contextual financial-language classification using FinBERT — but students do not run the model. Outputs are precomputed and frozen, so students compare frozen FinBERT labels against dictionary-based labels. Running transformer models (BERT/embeddings as active skills) is covered in BADM 576 (ML II).


Project (Team, runs across all eight modules)

Teams of three run one project across all eight modules, in one of two forms:

Type A — firm-specific signal. A measurable characteristic that predicts outcomes: insider trading patterns, analyst revisions, patent filings, ESG ratings, governance changes, attention-driven retail trading, institutional fire sales.

Type B — differential exposure signal. Entities differentially exposed to a macro theme — tariffs, climate regulation, AI disruption, interest rates, supply chain shocks — identified using text analysis, then tested for differential effects when relevant events occur.

Projects are not limited to equities. Credit risk, real estate, ESG, and other domains are welcome where the data supports the pipeline. Where an equity project uses “portfolio sorts → alpha,” a credit project uses “risk buckets → default prediction accuracy” and a real estate project uses “neighborhood segments → pricing error.” The requirements — honest out-of-sample evaluation, appropriate cross-validation, robustness checks — are identical across domains.

Teams present progress during studio sessions and submit the final project deliverable in M8.


Assessment

One major team project (the Trading Signal System) runs alongside the course, with four exercises (E1–E4) building foundational skills and project milestones scaffolding toward the final deliverable. No traditional exam.

Exercises

Decision-point notebooks: students are told what to do, not how, and each exercise ends with interpretation questions. Designed for approximately 2 hours with AI assistance, which is assumed rather than tolerated. Work is submitted through Canvas; GitHub distributes course materials to students.

Exercise Released after Due before Topic
E0: Diversification Diagnostic M1 W2 Returns → portfolios → market model. Required, ungraded.
E1: Stock Splits M2 W4 Stock-split events → CARs → window sensitivity and cross-sectional tests
E2: Short Interest + Text M4 W6 Short interest and text features → classification, train/test, AUC
E3: Accounting Anomalies M6 W7 15 accounting signals → Lasso vs. XGBoost, time-series CV
E4: Momentum + Factors M7 W8 Momentum → CAPM vs. FF3 alpha, factor loadings

Rubric (per exercise, 4 dimensions):

Dimension Excellent (A) Proficient (B) Developing (C)
Methodology Correct application of module methods, justified choices Reasonable approach with minor gaps Flawed or missing methodology
Model Evaluation Rigorous evaluation, explains metrics clearly Evaluation applied but basic Missing or weak evaluation
Code Quality Clean, documented Jupyter notebook, reproducible Adequate code, some comments Messy or undocumented
Written Analysis Connects methodology to findings; explains business implications Adequate explanation Minimal or unclear

Project Proposal (20%)

Team project proposal presented in Week 4. Deliverable: presentation + slideshow with notes. One holistic grade. Instructor questions during the presentation are graded within the presentation grade.

Content:

Milestone checkpoint Due Purpose
M2: Signal Construction + Regression Baseline End of Module 4 Constructed signal, baseline regression model, initial cross-validation results — formative feedback only
M3: Model Expansion + Text Integration End of Module 6 Classification/text features added to pipeline — formative feedback only

Rubric (3 dimensions):

Dimension Excellent (A) Proficient (B) Developing (C)
Progress Substantial, on-track work building on prior modules Adequate progress with some gaps Behind schedule or superficial
Technical Quality Methods applied correctly, evaluation included Functional but basic analysis Errors or missing components
Team Collaboration Clear evidence of shared work, complementary contributions Adequate collaboration Uneven contribution

Final Project (35%)

Complete the Trading Signal System and present in Weeks 7–8. Deliverable: presentation + slideshow with notes. One holistic grade. Instructor questions during the presentation are graded within the presentation grade.

Deliverables:

Rubric (4 dimensions):

Dimension Excellent (A) Proficient (B) Developing (C)
Signal Design Creative, well-justified signal from interesting data source Reasonable signal choice Generic or unjustified
Model Pipeline Multiple methods compared systematically, strong evaluation Functional pipeline Single method or weak evaluation
Portfolio Analysis Rigorous abnormal return analysis, addresses look-ahead bias Adequate portfolio construction Flawed methodology
Presentation & Q&A Confident delivery, articulates trade-offs, handles questions well Adequate delivery, answers most questions Unclear delivery or struggles with Q&A

Studio Sessions (ungraded — output folds into milestones)

Weekly engagement in live studio sessions. Students work through guided exercises, discuss approaches, and share progress on project milestones. Studio output is captured in project milestone grades.

AI Tools Integration

Modules 1-3 (Regression and Evaluation):

Modules 4-6 (Classification and Text):

Modules 7-8 (Ensembles and Synthesis):

AI Attribution: Students maintain a short project-level log of tools used. The AI policy grades the defensibility of the work; prompt-level disclosure is not required. See design/assessment_strategy.md for the attribution framework.

Assessment Summary

Component Weight Notes
Exercises (E1–E4, graded holistically, 8.75% each) 35% Individual, submitted via Canvas. E0 (Diversification Diagnostic) is required/ungraded.
Project Proposal (presentation + slideshow with notes, week 4) 20% Team of 3, one holistic grade; Q&A graded within presentation
Final Project (presentation + slideshow with notes, weeks 7–8) 35% Team of 3, one holistic grade; Q&A graded within presentation
Peer Review (quality of reviews given) 10% 5% on proposals, 5% on final slideshows

No traditional exam. Four exercises build skills; the team project integrates them. Live Q&A with the instructor is embedded in both presentations.

AI Usage Levels (AIAS)

Assessment AIAS Level AI Permitted
Exercises (E1–E4) 2 AI for debugging, interpreting outputs, finance concept explanation — with attribution
Project Proposal 2 AI for code assistance, data exploration — with attribution
Final Project 3 AI as collaborator for business case drafting and model comparison — with full disclosure
Studio sessions (ungraded) 1 AI for concept exploration during exercises

Technology Stack


Pedagogical Notes for Faculty

Design suggestions grounded in program research — not requirements. Adapt to your course and teaching style. Full references in reference/articles/.

The scenic route (cognitive friction) FIN 550 is where the tension between AI efficiency and learning depth is sharpest. Students can use Copilot to generate a random forest in seconds — but if they haven’t first built a baseline regression by hand (Modules 1-3), the ensemble result has no prediction to surprise them. The dopamine gap research shows the brain learns through prediction errors: the gap between what you expected and what happened. A student who struggles with linear regression before seeing how gradient boosting improves on it learns more than one who skips straight to XGBoost. The AIAS progression (0→1→2→3 across modules) already scaffolds this; the key is framing the early manual work as the investment that makes the later AI-assisted work register. → Machulla (2026), Schultz et al. (1997)

The IKEA effect (completion matters) The Trading Signal System runs across all 8 modules — this is the longest sustained project in the program. The IKEA effect research shows that labor leads to love only when it leads to completion. Each milestone (M1→M4) should feel like a working thing: a testable hypothesis, a running model, a pipeline that produces output. The oral defense in Module 8 is the ultimate completion signal. For career pivoters with no finance background, the moment they can explain their trading signal system to a panel is transformative — it’s when “I’m not a finance person” becomes “I built this.” → Norton, Mochon & Ariely (2012)

Variable uncertainty and calibrated difficulty The dopamine system is most engaged at ~50% uncertainty — when the student genuinely doesn’t know if they’ll succeed. Assignments that are too easy (certain success) or too hard (certain failure) produce flat responses. For career pivoters, Module 5 (Factor Models) is a known difficulty spike — the finance theory is unfamiliar. Consider front-loading the finance concepts students need (the “For career pivoters” notes are already in the right spirit) so the challenge is the ML application, not the domain vocabulary. The goal: students should feel “I might be able to do this” — not “I definitely can” or “I definitely can’t.” → Machulla (2026), Fiorillo, Tobler & Schultz (2003)

Three AI iterations before milestone submission For milestones M2-M4 (AIAS 2-3), consider requiring students to iterate with AI at least 3 times before submitting. This builds the habit of using AI as a thinking partner: first attempt → AI critique → revised attempt → AI alternative → final version with documented rationale. Produces richer AI Attribution Logs and prevents the “paste Copilot output, submit” pattern. → Means (2026, “Practice Gap”)

The Push-Back Protocol for model interpretation When students use AI to interpret regression outputs or explain cross-validation results (Modules 2-4), there’s a risk they accept AI’s explanation without verifying it against the data. The Push-Back Protocol (demand evidence → surface assumptions → request alternatives → stress-test → synthesize) is especially valuable here: ML model interpretation is exactly the kind of task where AI sounds confident but can be wrong. Consider making one exercise a structured push-back exercise. → Means (2025, “Push-Back Protocol”)

Peer review as milestone M4 The cross-team review at M4 is a strong design choice. The IKEA effect research suggests students learn as much from evaluating others’ pipelines as from building their own — exposure to different approaches calibrates their sense of quality. Consider structuring the peer review with the same rubric dimensions used for the final deliverable (signal design, model pipeline, portfolio analysis, written analysis) so students internalize the evaluation criteria before their own defense. → Norton et al. (2012), Li et al. (2020)

Attack your assessments Before the semester, have a confident AI user attempt each exercise and the trading signal project using current AI tools. AI is already strong at generating scikit-learn pipelines and interpreting regression output. Where can AI complete the task without genuine understanding of the finance context or model assumptions? Those are the spots to add pre-AI phases or shift weight toward the oral defense. → Furze (2026)


Course Sequence: ← BDI 513 — Data Storytelling Program sequence: BADM 558 — Big Data Infrastructures → ML track: BADM 576 — Data Science and ML →