1Cademy - Core Problem Types in NLP Pre-training

Learn Before

The Pre-training and Fine-tuning Paradigm

Classification

Core Problem Types in NLP Pre-training

Pre-training in Natural Language Processing primarily addresses two categories of problems: sequence modeling, also referred to as sequence encoding, and sequence generation. Despite their distinct forms, these two problem types can be conceptually unified and described using a single, general model formulation for the sake of simplicity.

Updated 2026-04-14

Contributors are:

Who are from:

References

Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course

Learn After

Sequence Encoding Models
Sequence Generation Models
Architectural Differences Between Sequence Encoding and Generation Models
General Formulation of a Sequence Model
A large language model is pre-trained on a vast text corpus. Its training objective is to take a sentence, randomly mask 15% of the words, and then predict only the original masked words by looking at all the surrounding unmasked words (both to the left and right). Which statement best analyzes the primary goal of this specific pre-training approach?
Analyzing Pre-training Objectives
Match each Natural Language Processing (NLP) task with the primary pre-training problem type it is designed to solve.

Learn Before

Related

Learn After