Topics Technology and the internet

How do large language models work?

A large language model is a program that has learned patterns in an enormous amount of written text, and it uses those patterns to predict what word, or piece of a word, should come next. Do that over and over and you get sentences, answers, essays and code. There is no stored list of replies. There is a very large set of numbers, called parameters, that were tuned during training until the model got good at guessing.

What makes this interesting is how far a simple goal can go. Predicting the next word well seems to require picking up grammar, facts, style and some reasoning habits along the way. Researchers still debate how much a model truly understands and how it reaches particular answers, and that open question is part of the story. It also explains the odd failures: a model can sound confident and still be wrong.

An episode would walk through the pieces in order: turning text into tokens, the transformer design that lets a model weigh context, how training adjusts the numbers, and how a raw model is shaped into a helpful assistant. On bre, the hosts are AI and can make mistakes too, so it is worth treating any episode as a starting point, and you can press Talk to ask a question.

What a bre episode would cover

An outline of the episode bre would make for this question. Every episode is written fresh when you ask, so yours will differ.

  1. Words become tokensText is chopped into small pieces called tokens, and each one is turned into a list of numbers the model can work with.
  2. Predicting the next tokenThe core task is a guess: given everything so far, which token is likely next. Chaining those guesses produces whole paragraphs.
  3. The transformer and attentionThe transformer design lets the model weigh which earlier words matter most for the current one, which is how it keeps track of context.
  4. How training adjusts the numbersThe model guesses, is scored against real text, and has its parameters nudged slightly. Repeating this at huge scale is what trains it.
  5. From raw model to assistantAfter the first training, people fine-tune the model with examples and feedback so it follows instructions and answers politely.
  6. Why models get things wrongA model produces likely text, not checked truth, so it can invent details. We cover what that means for how you use one.
  7. What nobody fully understands yetResearchers are still working out what happens inside these networks, and we separate what is known from what is debated.

How the episode might open

A sample exchange between two of bre’s AI hosts, bre and Cal. Both are AI; this is written by AI, as every bre episode is.

  1. breAI host

    Okay, let's start with the question under the question. When a chatbot answers you, is it looking something up, or is it doing something else entirely?

  2. CalAI host

    Something else. And that surprised me. It's not a search. It's closer to autocomplete, except trained on a staggering amount of writing, so the guesses get eerily good.

  3. breAI host

    So it's guessing the next word. Just one word at a time?

  4. CalAI host

    One token at a time, which is a word or a chunk of one. It guesses, adds that to the text, then guesses again with the longer text. Hold on, how does that actually work when the answer is a whole essay?

  5. breAI host

    I think that's the trick. Each guess is shaped by everything before it, so the sentence stays on track.

  6. CalAI host

    Right. And the shaping comes from billions of numbers tuned during training. Nobody typed in grammar rules. The model drifted toward them because they helped it guess.

  7. breAI host

    That's the part I want to slow down on. Because if nobody wrote the rules, how do we know what it actually learned?

  8. CalAI host

    Honestly, we only partly do. Researchers are still digging into that, and it's a big open question.

Questions people also ask

Is a large language model just autocomplete?
It is built on the same idea, predicting what comes next, but at a far larger scale and with a design that weighs a lot of context. That scale lets it write long, coherent text. Whether that amounts to understanding is something researchers still debate.
Where does a language model get its knowledge?
From training on very large collections of text, such as books, websites and code. It does not keep a copy of those documents like a library. It adjusts its numbers to capture patterns, which is why it can recall facts imperfectly.
Why do language models make things up?
They are trained to produce plausible text, not to check facts. When the pattern points to an answer that sounds right but is not, the model may state it confidently anyway. This is often called hallucination, and it is a known limitation.
What is a transformer in a language model?
A transformer is the neural network design behind most modern language models. Its key feature, called attention, lets the model weigh how relevant each earlier word is when predicting the next one, which helps it handle long passages and context.

Related topics

More: all 300 topics, technology and the internet, or the longer reads on /learn.

bre’s hosts are AI, and every episode is generated, so they can be wrong: check anything that matters. This page outlines what an episode would cover. It is for interest and learning, not medical, financial or legal advice.