Learn Before
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Learners explore the foundational architecture and empirical evaluation of modern language and multimodal models. The materials cover the core mechanisms of Transformers, including multi-head attention and sequence transduction, alongside in-depth analyses of GPT-4's performance on human exams, multilingual comprehension, and visual reasoning. Participants will gain insight into predictable scaling laws, post-training alignment via RLHF and rule-based reward models, and techniques for auditing dataset contamination.
0
1
Tags
Prep Sessions
Learn After
Ch.1 Transformer Architecture Fundamentals - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
attention-first3-last3.pdf
Ch.2 Model Scaling and Capability Evaluation - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
gpt4-selected-pages.pdf
Ch.3 Model Alignment and Safety - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor