Learn Before
Relation

Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Learners explore the foundational architecture and empirical evaluation of modern language and multimodal models. The materials cover the core mechanisms of Transformers, including multi-head attention and sequence transduction, alongside in-depth analyses of GPT-4's performance on human exams, multilingual comprehension, and visual reasoning. Participants will gain insight into predictable scaling laws, post-training alignment via RLHF and rule-based reward models, and techniques for auditing dataset contamination.

0

1

Updated 2026-09-07

Tags

Prep Sessions