In the previous post we learned how words can be transformed into feature vectors. However, language models don’t actually work with full words—they work with tokens. A token is the smallest unit of the model that can be a word, subword or even a single character (e.g., punctuation symbols).
For example, the word
unpredictable can be broken into the following tokens:-
un (prefix)-
predict (root)-
able (suffix)Language model is the model that estimates probability of a token or sequence of tokens appear within a longer sequence of tokens.
Example:
When I come home after work, I _________ in the kitchen.
Possible predictions:
- "cook dinner"
- "drink tea"
- "watch TV"
Each option has a different probability, and the model’s task is to pick the tokens with the highest probability to complete the sentence.
One more important element of the LLM is context. Context is helpful information before or after the target token. For example, context can help determine whether the word "orange" refers to a fruit or a color.
When a language model has a large number of parameters (billions of parameters), it’s called a large language model (LLM). LLMs are powerful enough to predict entire paragraphs or even entire essays.
Extending the example above it would be like:
User's question: What can I do in the kitchen after work?
LLM Response: _____
There is no magic - just math, probabilities, and predictions. So if you hear that LLM makes the scientific reasoning about some problem, it's not real reasoning, it's just advanced autocomplete for a given prompt.
#aibasics