Language Model

« Back to Glossary Index

A language model is a probability distribution over words or word sequences (more accurately, tokens), conducted with Machine Learning. In practice, it gives the probability of a particular word sequence being valid. I.e., to what extent the language model’s output resembles how people write, which is what the language model learns. Since language models are probability distributions, they can’t guarantee grammatical validity.

Natural languages (such as English or Spanish) are ambiguous. For example, “He saw her duck” can mean either that he saw a waterfowl belonging to her, or that he saw her moving to evade something. Thus, again, we can’t speak about one single meaning for a sentence. Instead, we talk about a probability distribution over possible meanings.

Separately, language models learn from text and can be used to produce original text, such as predicting the next word in a text, performing speech recognition, optical character recognition, and handwriting recognition.

Language Model Types

There are two types of language models:

  • Probabilistic language models — Simpler models, with several drawbacks, as:
    • They do not consider context to influence the probability of a word appearing,
    • They scale poorly as the text size increases, and
    • They create a sparsity problem. I.e., word probabilities have few different values, so most of the words have the same probability. As a consequence, the granularity of the probability distribution tends to be relatively low.
  • Neural network-based language models — More complex models that alleviate the issues in the probabilistic models.
    • Neural networks involve numerous matrix computations, and thus, they are expressed as functions. For this reason, they don’t need to store all intermediate calculations. Hence, neural network-based language models scale better, as they produce the probability distribution over the next word without storing counts of previous words.
    • Large Language Models (LLMs) are based on neural networks.
    • Word embedding layers create an arbitrary-sized vector for each word that includes semantic relationships.
      • It adds context to choose the following word.
      • These continuous vectors give rise to the granularity in the probability distribution of the next word, and so simplify the sparsity issue.
« Back to Glossary Index