1Cademy - A research team successfully trains a 1-billion-parameter language model. Encouraged by their results, they scale up the exact same architecture and training setup to a 100-billion-parameter version using a much larger dataset. Midway through the training process, the models loss value suddenly becomes `NaN` (Not a Number), and the training crashes. This happens repeatedly despite restarting from previous checkpoints. Which of the following best explains this phenomenon?

Learn Before

Training Instability in Large-Scale LLMs

Multiple Choice

A research team successfully trains a 1-billion-parameter language model. Encouraged by their results, they scale up the exact same architecture and training setup to a 100-billion-parameter version using a much larger dataset. Midway through the training process, the model's loss value suddenly becomes NaN (Not a Number), and the training crashes. This happens repeatedly despite restarting from previous checkpoints. Which of the following best explains this phenomenon?

Updated 2025-10-02

Contributors are:

Who are from:

Learn Before

Related