In June 2017, a team of eight researchers working at Google posted a paper with a cheeky title: “Attention Is All You Need.” It proposed a new design for neural networks called the transformer, tested it on translating English into German and French, and noted that the best version had trained in 3.5 days on eight graphics chips. It was a solid result in a niche field. Nobody outside machine learning paid much attention.
Today the “T” in ChatGPT stands for transformer. So does the architecture behind Claude, Gemini, Llama and nearly every other chatbot you’ve used. One of those eight authors, Aidan Gomez, went on to co-found Toronto’s Cohere.
Large language models get described as everything from “fancy autocomplete” to “digital minds.” Neither is quite right. Here’s a plain-language tour of how they actually work, what they’re good and bad at, and what they cost to run, with the hype dialled down.
Step one: chopping text into tokens
An LLM doesn’t read words the way you do. Before anything else, text is split into chunks called tokens. A token might be a whole common word like “the,” a piece of a longer word, a punctuation mark or a space. OpenAI’s rule of thumb for English is that one token is roughly four characters, or about three-quarters of a word. So 100 tokens is around 75 words.
Each token is then converted to a long list of numbers, called an embedding. You can think of it as coordinates in a space with thousands of dimensions, where tokens with related meanings end up near each other. Everything the model does from here on is arithmetic on those numbers.
Tokens explain some of the odd quirks of chatbots. Models have historically struggled with tasks like counting the letters in a word, because they never see individual letters, only chunks. Tokens are also how AI companies bill for their services and measure how much a model can “remember” in one conversation.
Step two: the transformer and attention
Before 2017, the leading language systems processed text one word at a time, in order, like someone reading with a finger under each word. That was slow, and the network tended to lose track of things mentioned far back in a passage.
The transformer’s key idea, called self-attention, lets the model look at every token in a passage at once and decide which ones matter for understanding each of the others. Jay Alammar’s widely used illustrated guide uses this sentence as an example: “The animal didn’t cross the street because it was too tired.” To understand “it,” the model needs to connect it to “animal,” not “street.” Attention is the mechanism that makes that link.
In simplified terms, it works like this:
- For every token, the model computes three sets of numbers, known as a query, a key and a value.
- Each token’s query is compared with every other token’s key to produce a relevance score.
- Those scores decide how much of each token’s value gets blended into the updated representation.
- This happens in many parallel “heads” and is stacked in dozens of layers, each refining the picture.
Because all tokens can be processed at the same time, transformers are very well suited to graphics processors, which are built to do huge numbers of calculations in parallel. That fit between the design and the hardware is a big reason the technology scaled so quickly.

Step three: pretraining, or learning to predict the next token
Here’s the part that surprises many people. The core training task for an LLM is almost absurdly simple: given some text, predict the next token. The model makes a guess, compares it with the real next token, and nudges its internal settings, called parameters or weights, to do slightly better next time. Repeat that trillions of times.
The scale is hard to picture. Meta said its Llama 3 models were pretrained on over 15 trillion tokens from publicly available sources, using two clusters of about 24,000 GPUs each. The largest Llama 3.1 model has 405 billion parameters, according to its model card.
Why does predicting the next word produce something that can summarize a contract or explain a tax rule? Because to predict text well across that much material, a model has to pick up grammar, facts, writing styles and patterns of reasoning. Nobody programs those in directly. They emerge as useful shortcuts for the prediction task.
OpenAI researchers documented in a 2020 paper on scaling laws that a model’s error falls in a predictable pattern as you increase its size, its training data and the computing power used. That finding helped justify the race to build ever-larger models.
Is it “just autocomplete”?
Mechanically, yes: the model outputs one token at a time. But researchers studying what happens inside these models have found more structure than the phrase suggests. In a March 2025 interpretability study, Anthropic reported that its Claude model planned rhyming words before writing a line of poetry, and combined separate facts (“Dallas is in Texas,” “the capital of Texas is Austin”) to answer a two-step question. In our view, “autocomplete” undersells what’s going on, and “thinking” oversells it. The honest answer is that we’re still mapping it.
Step four: fine-tuning into an assistant
A freshly pretrained model is a strange thing. Ask it a question and it might continue with three more questions, because that’s a plausible continuation of text it has seen. Turning it into a helpful assistant takes further training.
The approach that made ChatGPT possible was laid out in OpenAI’s InstructGPT paper in March 2022. It has two main stages:
- Supervised fine-tuning: people write examples of good answers to prompts, and the model is trained to imitate them.
- Reinforcement learning from human feedback (RLHF): people rank several model answers from best to worst, and the model is trained to produce the kinds of answers people prefer.
The results were striking. Human evaluators preferred answers from a 1.3-billion-parameter InstructGPT model over those from the 175-billion-parameter GPT-3, despite it being over 100 times smaller. Size isn’t everything. How a model is shaped after pretraining matters enormously.
Businesses can also fine-tune existing models on their own data, though many find it simpler to give a model relevant documents at the moment of asking, a technique often called retrieval-augmented generation.
Why LLMs make things up
The most important limitation to understand is hallucination: a model stating something false with complete confidence. It’s not a bug in the usual sense. The system is built to produce plausible text, and plausible and true are not the same thing.
OpenAI argued in a September 2025 explainer that the way models are trained and graded makes this worse. Most benchmarks score only accuracy, so a model that guesses is rewarded over one that admits it doesn’t know, much like a student guessing on a multiple-choice test. In one example, a chatbot asked for a researcher’s PhD dissertation title gave three different answers, none correct.
The real-world consequences show up most clearly in courtrooms. In the U.S. case Mata v. Avianca, two lawyers were fined US$5,000 in June 2023 after filing a brief that cited fake cases invented by ChatGPT. Canada has seen the same problem: by mid-June 2026, at least 167 filings with AI-hallucinated citations had turned up across 51 Canadian courts and tribunals, according to a running tally of documented cases.
Practical rules follow from this:
- Treat an LLM’s factual claims as leads to verify, not as answers.
- Ask for sources, then check the sources exist and say what the model claims.
- Be most careful with names, numbers, dates, quotes and legal or medical specifics, which are where invented details hide.
What all this costs: compute and energy
Building and running LLMs takes a lot of computing power, and that has a growing energy footprint. Some sourced figures help put it in perspective.
| Measure | Figure | Source |
|---|---|---|
| Growth in compute used to train frontier AI models | About 4 to 5 times per year (2010 to 2024) | Epoch AI |
| GPU time to train Llama 3.1 405B | 30.84 million H100 GPU-hours | Meta model card |
| Location-based emissions, training Llama 3.1 family | 11,390 tonnes CO2-equivalent | Meta model card |
| Energy for a median Gemini text prompt (May 2025) | 0.24 watt-hours | |
| Global data centre electricity use, 2024 | About 415 TWh, 1.5% of world total | IEA |
| Projected data centre use, 2030 | About 945 TWh | IEA |
A few notes on those numbers. Epoch AI, a research group that tracks the field, found training compute for notable models grew about 4.1 times a year between 2010 and May 2024. Meta’s model card says its market-based emissions were zero because it matched its use with renewable energy, but the location-based figure reflects the actual grids involved.
On the usage side, Google reported in August 2025 that a median Gemini Apps text prompt used 0.24 watt-hours, which it compared to watching TV for under nine seconds, and that energy per prompt had fallen 33-fold in a year. A single query is small. Billions of them a day are not.
The bigger picture comes from the International Energy Agency, which estimates data centres used about 415 terawatt-hours in 2024 and projects that will more than double to roughly 945 TWh by 2030, with AI as the most important driver. For Canada, with its large supply of hydroelectric power, that demand is both an opportunity to attract data centres and a strain on provincial grids already juggling electrification.
What to keep in mind
Strip away the mystique and an LLM is a very large statistical model that turns text into tokens, uses attention to weigh how those tokens relate, and predicts what comes next, shaped by human feedback into something that behaves like an assistant. That recipe turns out to be remarkably capable, and it has real limits: it doesn’t know when it’s wrong, and it consumes a lot of electricity at scale.
The useful stance is neither awe nor dismissal. Use these tools for drafting, summarizing, brainstorming and explaining, where a plausible first pass is valuable and you can check the result. Be sceptical wherever being confidently wrong would be costly. And when someone tells you an AI “understands” or “is just autocomplete,” you’ll know both claims leave out most of the story.
Sources and further reading
- Vaswani et al.: Attention Is All You Need (2017)
- OpenAI Help Center: What are tokens and how to count them
- Jay Alammar: The Illustrated Transformer
- Meta: Introducing Llama 3
- Meta Llama 3.1 405B model card
- Kaplan et al.: Scaling Laws for Neural Language Models (2020)
- Anthropic: Tracing the thoughts of a large language model
- Ouyang et al.: Training language models to follow instructions with human feedback (2022)
- OpenAI: Why language models hallucinate
- Hallucination (artificial intelligence): documented cases
- Epoch AI: Training compute of frontier AI models grows by 4-5x per year
- Google Cloud: Measuring the environmental impact of AI inference
- IEA: Energy and AI, executive summary
Leave a comment