University of Maryland · Fall 2026

Multimodal
foundation
models.

A research-driven course on the models that connect language, vision, audio, and action—from the transformer’s first principles to capable multimodal systems.

WhenTue / Thu · 12:30–1:45 PM
WhereCSI 3117
InstructorJia-Bin Huang

Lecture
schedule.

#DateLectureSlidesSupp video
Module 01 Transformer foundations
L01Tue · Sep 1Introduction
  • Course logistics
  • Multimodal foundation model applications
Slides ↗
L02Thu · Sep 3Transformer
  • Tokenization
  • Sequence-to-sequence architecture
Slides ↗
L03Tue · Sep 8Position embedding
  • Absolute position
  • Relative position
  • Rotary position embeddings
Slides ↗
L04Thu · Sep 10Attention design
  • Key-value cache
  • Multi-head attention
  • Multi-query attention
  • Grouped-query attention
  • Multi-head latent attention
  • DeepSeek Sparse attention (DSA)
Slides ↗
L05Tue · Sep 15Flash Attention
  • GPUs
  • Tiling algorithm
  • Online softmax
  • FlashAttention
L06Thu · Sep 17Linear attention
  • Linear attention
  • Chunkwise parallel training
  • Test-time regression perspective
  • Delta update rule
L07Tue · Sep 22Mixture of Experts
  • Feedforward layers
  • Activations
  • Load balancing
  • Router z-loss
L08Thu · Sep 24Residual connections
  • Residuals
  • Hyper connections
  • Attention residuals
  • Gated residual
L09Tue · Sep 29Embedding scaling
  • N-gram
  • Emgram
M1Thu · Oct 1Midterm 1 — in-class exam
Module 02 Large Language Models
L11Tue · Oct 6Pretraining and scaling laws
  • Next-token prediction
  • Data, model, and compute scaling
  • Chinchilla-optimal training
L12Thu · Oct 8Prompting and PEFT
  • Low-Rank Adaptation
  • LoRA+
  • Quantized Low-Rank Adaptation
  • Vector-based Random Matrix Adaptation
  • Weight-Decomposed Low-Rank Adaptation
L13Tue · Oct 13Fall Break
L14Thu · Oct 15Post-training
  • Reinforcement Learning from Human Feedback
  • Proximal Policy Optimization
  • Direct Preference Optimization
L15Tue · Oct 20Reasoning
  • Chain-of-thought
  • Test-time scaling
  • GRPO and GSPO
L16Thu · Oct 22Efficient inference
  • Quantization
  • Speculative decoding
  • PagedAttention
L17Tue · Oct 27Efficient training
  • Parallelism
  • Mixed precision
  • fp16, bf16, fp8
Module 03 Multimodal models
L18Thu · Oct 29Large multimodal models
  • Example state-of-the-art models
L19Tue · Nov 3Self-supervised representation learning
  • SimCLR
  • DINO
  • Masked autoencoders
  • CLIP
  • JEPA
M2Thu · Nov 5Midterm 2 — in-class exam
L21Tue · Nov 10Diffusion
  • Variational autoencoders
  • Diffusion training
  • Guidance
  • Latent diffusion
L22Tue · Nov 17Flow matching
  • Flow matching foundations
  • MeanFlow
  • Efficient sampling
Module 04 Systems & applications
L23Thu · Nov 19Applications · robot learning
  • Vision-language-action
  • Policy learning
  • World models
L24Tue · Nov 24Agentic AI 1 — reasoning & tool use
  • Test-time scaling
  • Tool calling and function APIs
  • Web search, code execution, and computer use
Nov 25–29Thanksgiving recess
L25Tue · Dec 1Agentic AI 2 — agent harnesses
  • Planning and orchestration
  • Memory and state
  • Tracing, evaluations, and guardrails
M3Thu · Dec 3Midterm 3 — in-class exam
L26Tue · Dec 8Applications · video & audio
  • Multimodal models for video
  • Multimodal models for audio
L27Thu · Dec 10Applications · 3D
  • Multimodal models for 3D understanding
  • 3D creation

Learn by
looking closely.

Coursework emphasizes clear technical thinking, regular engagement with research papers, and an original project in multimodal foundation models.

100%coursework

People.

Questions about the course are best directed through the course communication channel once it is announced. The instructor’s office is IRB 4234.