Learn Before
Concept icon
Concept

General Equations of an N-Gram Model

The general equations of an n-gram model apply the Markov assumption to estimate word probabilities. The conditional probability for the next word is approximated by looking N−1N-1 words into the past:

P(wn∣w1:n−1)≈P(wn∣wn−N+1:n−1)P(w_n|w_{1:n-1}) \approx P(w_n|w_{n-N+1:n-1})

The probability of a complete word sequence is approximated as the product of these conditional probabilities:

P(w1:n)≈∏k=1nP(wk∣wk−N+1:k−1)P(w_{1:n}) \approx \prod^n_{k=1}P(w_k|w_{k-N+1:k-1})

0

1

Concept icon
Updated 2026-06-19

Tags

Deep Learning

Data Science

Machine Learning Yearning @ DeepLearning.AI

Dive into Deep Learning @ D2L

Machine Learning

Supervised Learning