How a large language model works
Prelims and MainsCurrent affairs on this: AI and robotics: technology and India's capability
A large language model is a program that predicts the next piece of text. Text is cut into tokens, each turned into numbers, and a neural network of the transformer type is trained on a vast body of text by one task repeated billions of times: guess the next token, and adjust the model's internal numbers, its parameters, when wrong. In use it answers a prompt one token at a time. Because it produces what is likely rather than what is true, it can state false things fluently: hallucination. "Large" refers to parameters, now in the hundreds of billions, which is why data centres and graphics processors make news.
From text to answer
1 Tokenise
Text is cut into tokens and each becomes numbers.
2 Pretrain
The transformer learns to predict the next token over a huge corpus; the learnt numbers are the parameters.
3 Align
Human feedback teaches it to follow instructions and refuse.
4 Prompt
The user's text plus any retrieved documents go in.
5 Generate
One token at a time, each chosen from the likely next tokens.
Where it breaks: Hallucination: likely is not true.
- Tokens in, one token out, repeated: everything an LLM does is next token prediction over the prompt plus its own output so far.
- Pretraining learns language and facts from text; a second round of training on human feedback teaches it to follow instructions and refuse some requests; retrieval augmented generation, which attaches a search or document store and puts the results in the prompt, adds current facts at use and reduces hallucination.
- Parameters are the learnt numbers; the context window is how much text the model can consider at once; both drive cost.
- Hallucination is structural: a likely sentence is not a checked one, so an LLM is unreliable for facts unless grounded.
- India's models, BharatGen and Sarvam under the IndiaAI Mission, are "sovereign" in the sense of being trained in India on Indian languages and data.
- An open weight model is one whose trained parameters are released for download, so anyone can host it locally or modify it; most closed frontier models are not, and a model whose weights have been published cannot be withdrawn, which limits what export controls on models can do.
Mains: An LLM converts computing power and text into fluent language; its value to India lies in Indian language access to services, and its risk in fluent error, so the policy questions are data, compute and accountability, not the model itself.
UPSC has asked
- Prelims 2026: how Large Language Models work
See also: IndiaAI Mission · BHASHINI · India's data centre capacity · The European Union's AI Act · Global governance of artificial intelligence