BADM 554 - Enterprise Database Management

Program-level details: See program/curriculum.md

Soft launch 2026-08-17 (Week 0 + Module 1); Modules 2–8 unlock on Day 1, 2026-08-24. Assessment numbers below are final (Cheng Li’s corrected assessment plan, 2026-08-14, live in Canvas).

Live LD team data (module items, video lengths, production status): see badm554/sync/ — machine-generated snapshots from the canonical Box Excel files, re-synced ~4x/day (check _last-synced.md for freshness). Canonical build detail lives in the course repo (~/teaching/badm554/semesters/fall2026-online/).

Credits: 4 Term: Fall 2026 (Weeks 1-8) Instructor: Vishal Sachdev

Course Vision

Learners arrive able to query data. They leave able to design it. BADM 554 follows an OLAP-first arc: weeks 1–3 build fluency querying a real analytical warehouse (BigQuery public datasets), weeks 4–6 shift to designing and building that warehouse from transactional source data, and weeks 7–8 validate, document, and defend the result. By course end, each team has a working ETL pipeline, a dimensional schema, and a defensible oral presentation — a portfolio artifact, not a tutorial completion.

Learning Outcomes

Two naming schemes are used in parallel: the L-C-E framework (Vishal’s design language, used in week-by-week planning and this file) and CLO numbers (Cheng’s LD framework, used in Course Map.xlsx and the activity roster). The mapping below is the bridge. When Cheng populates the Excel, he uses CLO numbers; this file shows both.

Excel amendment needed (for Cheng): Add an “L-C-E Label” column to the Course Overview CLO table in Course Map.xlsx. Populate with the L-C-E labels below. This lets the sync files carry both schemes and eliminates ambiguity for K-ai.

Source of truth for outcome text (decided 2026-08-13, Vishal): the course repo’s design spec (~/teaching/badm554, 2026-05-10 design spec §”Revised L-C-E learning outcomes”), as carried into the Fall 2026 syllabus. The eight CLOs below are that revised set in the existing slot structure — the slot numbering Cheng’s Course Map already uses is unchanged.

CLO # L-C-E Label Outcome
CLO 1 L1 + L2 Explain why analytical warehouse schemas and transactional source schemas differ and what each is optimized for; read a SQL query (joins, aggregates, window functions) and say what business question it answers
CLO 2 L3 Recognize when AI-generated code, schemas, or analyses are confidently wrong, and name the verification step that would catch it
CLO 3 C1 Frame a stakeholder’s question as a query plan and answer it with SQL against a real warehouse dataset
CLO 4 C2 Design a star schema for a given analytical question — choosing the grain and justifying denormalization trade-offs — and read a normalized transactional schema well enough to explain why it is built that way
CLO 5 C3 + C4 Build an ETL notebook that turns a transactional source into a dimensional, analytics-ready warehouse, moving between SQL and pandas as the step requires
CLO 6 C5 Choose an appropriate way of reaching the database for a given task (notebook or command line) and articulate the trade-off
CLO 7 E1 + E2 + E3 Defend a schema-design or access-pattern decision under questioning, and validate a data product end to end so another analyst could repeat the checks
CLO 8 P1 Document how at least one design decision evolved across the term, including the pre-AI approach, what AI contributed, and what prompted each revision

L-C-E detail (for planning, pedagogy, and AIAS labels)

Literacy (Foundational Awareness)

Competency (Applied Skills)

Expertise (Advanced Application)

Process Learning Outcome (Black Box — see assessment_strategy.md §3.5)

Note on query optimization: Learners gain awareness that slow queries exist and why (indexes, scan patterns), covered in Week 7. Deep optimization is a DBA/data engineering concern beyond the scope of this first course — revisited in later courses as appropriate.

Course Arc (3 + 3 + 2)

Phase Weeks Schema lens What learners do
1 — Question + Warehouse SQL 1–3 Read OLAP / dimensional Query BigQuery public datasets; learn SELECT through window functions + CTEs; pick a stakeholder question and dataset
2 — Build Analytical Data Product 4–6 Build OLAP star; read OLTP 3NF Design a dimensional schema; read Sakila (3NF reference) + personal-data exports (real-world source); build ETL pipeline into DuckDB; provider-group GitHub schema negotiation
3 — Validate + Defend + Awareness 7–8 Validate end-to-end; NoSQL awareness Add validation checks + documentation; oral defense; awareness of when relational isn’t the answer

Every week: 4 framing videos (5–8 min) → curated practice → Jupyter lab → Live Session (90 min) → Project Studio (90 min) → concept assignment → Canvas discussion post → project milestone.

Week-by-Week Breakdown

Week Topic CLO focus Project milestone
1 Warehouse SQL: SELECT / WHERE / GROUP BY on BigQuery public data CLO 1, 3 M1 Project Launch (1%) — team reg + dataset scoping
2 JOINs + aggregates: combining tables to answer real questions CLO 1, 3 M2 Initial Pitch (3%) — 5-min team pitch video: stakeholder, framing, 3–5 questions
3 Window functions, CTEs + parallel access patterns (CLI/MCP introduced) CLO 3, 6 M3 Proposal Draft (2%) — mentor-graded in Canvas
4 Dimensional modeling: star schemas, grain, facts and dimensions CLO 4 M4 Proposal Final (12%) — schema design + ETL plan + feedback-closure memo; M4 teammate eval (1%)
5 Source-data reality: 3NF (Sakila) vs personal-data exports (real-world mess) CLO 2, 4, 5 M5 Analysis Draft (1%) — progress check-in
6 ETL: 3NF → DuckDB warehouse + provider-group GitHub schema negotiation CLO 5, 6 M6 Graded Peer Review of draft data product (4%, Peerceptiv)
7 Validation, documentation, AI Attribution Log wrap-up CLO 2, 7 M7 Revision + Rehearsal (1%) — revised draft + completion-graded rehearsal video
8 NoSQL awareness + Final Deliverable + Individual Oral Defense CLO 7, 8 Final (15%) + Oral (20%) + M8 teammate eval (1%)

Team Project: Stakeholder-Driven Analytical Data Product (Team of max 3, Weeks 1–8)

Learner teams pick a stakeholder + dataset pairing from a curated list or propose their own (subject to instructor approval). The pairing defines a real analytical question the team will answer with a designed, built, and validated data product.

Curated dataset pairings (BigQuery public datasets — no download required):

Weeks 5–6 source data: Learners also work with their own personal-data exports (Spotify, YouTube, Strava) as a real-world messy source — uniquely AI-resistant because the data is individual. Provider-groups (learners with the same export source) collaborate via a shared GitHub repo to negotiate a shared schema.

3NF reference dataset (Weeks 5–6): Sakila via the Pagila Postgres port — a clean textbook 3NF schema used as a contrast to the real-world mess of personal-data exports.

Project Milestones

Weights below are final (Cheng Li’s corrected assessment plan, 2026-08-14; 1 Canvas point = 1% of the course grade). The project group totals 59%.

Milestone Week Weight Deliverables
M1 Project Launch 1 1% Team registration, dataset + stakeholder pick, 1-page scoping notes
M2 Initial Pitch 2 3% 5-min team pitch video to Canvas: stakeholder, 3–5 questions, dataset rationale (Live Session is the optional practice run)
M3 Proposal Draft 3 2% Draft proposal, mentor-graded in Canvas
M4 Proposal Final 4 12% Final proposal: stakeholder, questions, dimensional schema design, ETL plan, AI Attribution checkpoint, feedback-closure memo
M4 Teammate Evaluation 4 1% Formative teammate rating + early warning (Peerceptiv)
M5 Analysis Draft 5 1% Progress check-in
M6 Graded Peer Review 6 4% Structured peer review of another team’s draft data product (Peerceptiv)
M7 Revision + Rehearsal 7 1% Revised draft + feedback-closure note + completion-graded rehearsal video (5-min check-in + hardest expected questions, on camera)
Final Deliverable 8 15% Working ETL + dimensional schema + 2–3 validated analyses + GitHub repo + README
Individual Oral Defense 8 20% Individual, ~5-min defense + follow-up Q&A on your own contribution and AI-use evolution (AIAS 0, no AI). Live only — Project Studio, mentor office hours, or professor office hours; no Canvas video option
M8 Teammate Evaluation 8 1% Summative teammate rating (Peerceptiv)

Final Deliverable (15% of grade) — key components

Rubric (6 dimensions):

Dimension Excellent (A) Proficient (B) Developing (C)
Schema Design Star schema properly grained, dimensions clean, trade-offs justified Mostly correct schema with minor issues Grain problems or missing relationships
ETL Pipeline Modular, idempotent, documented, robust validation checks Good structure, handles happy path Fragile pipeline, minimal error handling
Source-Data Handling Personal-data and/or BigQuery source read correctly; provenance documented Source read works, provenance partial Source integration missing or broken
Documentation Complete architecture diagram, clear setup guide, code comments Adequate documentation Incomplete or unclear
Business Understanding Analyses answer the stakeholder question; decisions explained in business terms Mentions business context No business rationale
Learning Trajectory Clear pivot or error recovery with explicit rationale — shows how thinking evolved M1 → Final Some revision with partial explanation Final product with no visible iteration

Oral Defense (20% of grade)

Week 8, live only (decided 2026-08-15). This is a purely individual assessment, distinct from the team presentation scored under the Final Deliverable — each student defends their own contribution and their AI-use evolution across the term in a single ~5-minute defense plus follow-up Q&A, sat live with a person (Project Studio, mentor office hours, or professor office hours). There is no Canvas-recorded submission path. The session is recorded and graded from that recording. Slots are booked during Week 7 for Week 8.

This is a deliberate exception to the program’s “no graded milestone is attendance-gated” default, made for cohort 1 specifically: at ~30 learners the personal, high-touch version is affordable and the enrollment payoff of a strong first-year experience outweighs the scalability cost. Anyone who cannot make a published slot is guaranteed an accommodation slot — attendance is handled, not left to chance. The in-house viva tool (not yet named to learners) is expected to supply an async accommodation path when it ships; it will not replace the live default.

Rubric (3 dimensions):

Dimension Excellent (A) Proficient (B) Developing (C)
Technical Explanation Clear walkthrough of design decisions, trade-offs articulated Adequate explanation of system Unclear or surface-level description
Live Demo System works end-to-end, handles follow-up queries confidently Demo works with minor issues Demo fails or cannot answer questions
Individual Contribution Each member explains their role with depth Members can describe their work Uneven participation or vague answers

Weekly Assignments (31%), Progress Checks (4%), Discussion (4%)

Each of Modules 1–7 has one graded individual weekly assignment run as a pre-AI / AI / post-AI cycle with an AI Attribution Log, in Jupyter (M1 3% · M2 5% · M3 3% · M4 5% · M5 5% · M6 5% · M7 5% = 31%; Module 8 has none). Progress-check quizzes in M1 and M3 (2% each, AIAS 0). One Live Session discussion post per module (0.5% × 8 = 4%). Learner-facing vocabulary is always practice → assignment → project (never “lab”).

AI Tools Integration

Where AI Accelerates Learning:

  1. Week 1 (SQL Fundamentals): Use Claude/ChatGPT to:
    • Explain SQL errors (“Why does this JOIN return NULL?”)
    • Generate sample SQL queries for practice
    • Validate your schema design
    • Suggest refactoring for clarity
  2. Week 4-5 (Python ETL): Use AI to:
    • Debug pandas errors (“How do I reshape this DataFrame?”)
    • Optimize code performance (“How do I vectorize this loop?”)
    • Generate error handling patterns
    • Suggest pandas functions for complex transformations
  3. Week 7-8 (Validation & Defense Prep): Use AI to:
    • Propose validation checks for the ETL pipeline (then verify them yourself)
    • Stress-test documentation (“could another analyst repeat these checks?”)
    • Generate hard practice questions for the oral-defense rehearsal
    • Review query correctness on edge cases

Three-Options Prompting Protocol (Weeks 4-8, AIAS 2):

For all milestone submissions (M4-M7) and the final deliverable, students must use the three-options format when using AI for design or architectural decisions:

  1. Prompt AI for three distinct options (e.g., “Give me three different approaches to structuring this ETL pipeline”)
  2. Summarize each option briefly
  3. Select one and justify the choice in 2-3 sentences based on your specific project context

Submission format:

AI Prompt: [paste exact prompt]
Option A: [1-sentence summary]
Option B: [1-sentence summary]
Option C: [1-sentence summary]
Chosen: [A/B/C]
Rationale: [why this option fits your context better than the others]

This replaces generic AI attribution (“I used Copilot”) with evidence of evaluative judgment. Options must differ on at least one meaningful dimension (e.g., performance vs. simplicity vs. maintainability). The three-options protocol does not apply to Weeks 1-3 (pre-AI / AIAS Level 0-1 phases).

Per-Milestone AI Attribution Log (program-standard format):

Aligned with design/project_milestone_template.md. Required at four checkpoints:

Milestone What to document
M2 (Wk 2) ERD design decisions — what AI suggested, what you changed, why
M4 (Wk 4) Extraction script — AI-suggested approaches accepted vs. rejected, with rationale
M6 (Wk 6) ETL pipeline — which logic was AI-generated vs. human-specified; flag any AI output accepted without verification
Final (Wk 8) Full Attribution Log covering all AI use across the project + 1-paragraph reflection: how did AI use evolve from M1 to Final? What would you do differently?

Format for each entry:

Milestone: [M2 / M4 / M6 / Final]
AI Tool Used: [Claude / Copilot / ChatGPT / other]
Prompt (summarized): [what you asked]
Output Used: [what you kept]
Output Rejected: [what you discarded and why]
Human Decision: [what judgment you applied that AI could not]

Studio Session Topics:

Assessment Summary

Final weights (Cheng Li’s corrected assessment plan, 2026-08-14; matches live Canvas, 1 point = 1%):

Component Weight Notes
Weekly assignments (M1–M7) 31% Individual, pre-AI/AI/post-AI + Attribution Log, AIAS 2
Progress-check quizzes (M1, M3) 4% Individual, AIAS 0
Live Session discussion posts (M1–M8) 4% Individual; attendance not required
Project milestones (M1–M7: launch, pitch, draft, proposal, analysis draft, peer review, revision+rehearsal) 24% Team (M6 peer review 4% in Peerceptiv)
Final project deliverable (Wk 8) 15% Team
Individual oral defense (Wk 8) 20% Individual, AIAS 0, live only (Studio / mentor / professor office hours)
Teammate evaluation (Peerceptiv) 2% Wk 4 formative 1% + Wk 8 summative 1%

Total: 100%. No traditional exam. All assessment is project-based, skills-based, and participation.

AI Usage Levels (AIAS)

Assessment AIAS Level AI Permitted
Progress-check quizzes (M1, M3) 0 No AI
Weekly assignments (M1–M7) 2 AI-mediated with attribution (pre-AI / AI / post-AI cycle)
Project milestones + final deliverable 2 AI for code assistance, design options, and debugging, with attribution
Initial pitch + oral defense Q&A 0 Prep may use AI (AIAS 2); the pitch/defense itself is your own answers, no AI
Teammate evaluations 0 No AI
Studio sessions (ungraded — output folds into milestones) 1 AI for exploration during exercises

Technology Stack

Prerequisites & Assumptions

Preparatory Coursework Alignment (Staggered Model, 2026-05-18)

Per the program-wide staggered preparatory coursework model (see program/curriculum.md “Preparatory Coursework” and the 2026-05-18 decision, recorded in the internal registry), students arrive in BADM 554 with the following prep status:

Module Required by Tier
GitHub + VS Code (DataCamp) Before Day 1 (BADM 554 Week 1) Gated — completion required
Gen AI literacy End of Week 4 Gated — runs in parallel with this course
Python & Tools prep (Coursera — specific course TBD) Before BDI 513 (Week 5) Self-diagnostic; students may still be working through during weeks 1-4
Stats MOOCs 1 + 2 Before FIN 550 (Week 9) Self-diagnostic

Implications for BADM 554:

Faculty office hours during Weeks 1-3 are recommended for students still finishing Python self-diagnostic prep — see Ron Guymon’s 2026-04-15 note on zero-coding-experience learners.


Pedagogical Notes for Faculty

Design suggestions grounded in program research — not requirements. Adapt to your course and teaching style. Full references in reference/articles/.

The scenic route (cognitive friction) Pre-AI phases aren’t punishment — they’re the scenic route. Schultz’s dopamine research shows the brain registers learning through prediction errors (the gap between expectation and outcome), not through frictionless delivery. When students design their ERD by hand before asking Copilot to generate CREATE TABLE statements, that struggle is the mechanism — not the obstacle. Consider where in each week students should formulate their own approach before consulting AI. The AIAS progression in this course (0→1→2→3 across weeks) already builds this in; the key is framing it for students as investment, not restriction. → Machulla (2026), Schultz et al. (1997)

The IKEA effect (completion matters) Students value what they build — but only when labor leads to a finished artifact. Building and then discarding produces no ownership effect. The M1→M7→Final Deliverable pipeline already does this well: each milestone closes a loop. The oral defense is the ultimate completion signal — the moment where effort converts to demonstrated competence. When possible, make each milestone feel like a working thing (a queryable schema, a running pipeline), not just a checkpoint document. → Norton, Mochon & Ariely (2012)

Prompt sophistication as skill progression Anthropic data shows r > 0.92 correlation between prompt sophistication and output quality, and multi-turn interactions succeed 67% of the time vs. 49% for single-turn. This course is most students’ first encounter with AI-assisted coding (AIAS 1→3). Consider explicitly teaching prompt patterns early — not as a sidebar, but as a core data foundations skill: “asking the right question of an AI is the same skill as writing the right SQL query.” → Means (2026, “How You Ask”)

Three AI iterations before human review For project milestones M4-M7 (ETL pipeline work at AIAS 2-3), consider requiring students to iterate with AI at least 3 times before submitting for human feedback. This builds the habit of treating AI as a drafting partner, not an answer machine, and produces richer AI Attribution Logs. → Means (2026, “Practice Gap”)

Attack your assessments Before the semester starts, have a confident AI user (TA, LD, or yourself) attempt each major assignment using current AI tools from a student’s perspective. Where can AI complete the task without genuine understanding? Those are the spots to add pre-AI phases or shift weight toward the oral defense. Repeat each semester — AI capabilities shift fast. → Furze (2026)

The cognitive offloading U-curve (Zone 2 is the danger zone) Learning outcomes follow a three-zone pattern: Zone 1 (no AI) = full cognitive load, slow but real learning; Zone 2 (scattered, unstructured AI use) = worse outcomes than no AI at all — coordination overhead without genuine cognitive reallocation; Zone 3 (committed, structured delegation) = superior outcomes through deliberate offloading of routine work while investing freed capacity in higher-order reasoning. The AIAS progression in this course (0→1→2→3) is designed to move learners through Zone 1 early, through Zone 2 quickly, and into Zone 3 by Weeks 5-8. The risk is AIAS Level 1 (“AI for ideation only”) becoming a Zone 2 trap if students engage with AI vaguely and inconsistently. Mitigate by giving Level 1 tasks a clear structure: “Before using AI, write your own answer. Then ask AI. Then compare and justify any changes.” That sequence is Zone 3 behavior, even at low AI permission levels. Separately: a 2026 study of 912 learners found that “partnership orientation” — treating AI as a thinking partner rather than an answer machine — independently predicted deeper learning regardless of how much AI was used. The mental model learners bring to AI matters as much as the permission level. Frame Week 1’s AI orientation around this: “Using AI well is the same skill as writing the right SQL query — the quality of your question determines the quality of the answer.” → Hardman (2026), “The Cognitive Offloading Paradox”; Wang & Zhang (2026)

AI as material for thinking, not a shortcut around cognition (design thinking checklist) Before finalizing any assignment or milestone, apply the three diagnostic questions from assessment_strategy.md §3.4: (1) What is the learning outcome this task targets? (2) What cognitive work must remain with the student to achieve that outcome? (3) What kind of AI assistance makes that cognitive work more visible, more rigorous, or more equitable — rather than replacing it? For BADM 554: the cognitive work that must stay with students is schema design judgment (Wks 1-3) and pipeline architecture decisions (Wks 4-7). AI assistance at those stages should make the student’s reasoning more explicit — not substitute for it. The per-milestone attribution log is the mechanism. → Vander (2026), “Design Thinking for AI-Integrated Pedagogy”

Faculty action required: Before T&L course design review, document your answers to the three diagnostic questions above in writing for this course. See assessment_strategy.md §3.4 for the submission format.


Recruiting Talking Points

Drafted by Vishal Sachdev. Adopted as the model for program-wide faculty talking points (Ravi Mehta directive, 2026-05-27). Recorded in the internal decision registry as “Faculty course talking points: three-question framework for recruiting” — full record from K-ai.

What does the course cover? BADM 554 is the first course of MSBAi: 8 weeks of database fluency for working professionals pivoting toward analyst work. Students learn SQL, schema design, and ETL on real enterprise data; no prior SQL required. Each week pairs a stakeholder conversation with a pre-AI / AI-mediated / post-AI workflow that treats AI as a thinking partner, not a shortcut.

What do students build? Teams of three adopt a stakeholder and a real dataset in Week 1, then spend 8 weeks building a portfolio GitHub repository: a dimensional warehouse, an ETL notebook ingesting from a live public data source, an AI Attribution Log, and a cloud-hosted version of the warehouse with a queryable URL. Each student defends the team’s work in an individual oral examination at Week 8.

Why does it matter for careers? SQL and schema judgment are baseline. What separates hireable analysts is talking to a non-technical stakeholder, working with AI without being replaced by it, and shipping a defensible artifact another person can audit. BADM 554 builds those three skills explicitly.


Course Sequence: ← Previous: (first course) Next: BDI 513 — Data Storytelling →