What is Large Language Model (LLM)?
An artificial intelligence model trained on massive text corpora using deep transformer architectures to understand, generate, summarize, and reason over human language and code.
Key facts
- Built upon the self-attention transformer mechanism introduced in 2017.
- Operates autoregressively by calculating token probability distributions.
- Context windows in 2026 range from 128k to upwards of 2M+ tokens.
- Powers modern conversational assistants, code generators, and agentic workflows.
Explanation
Large Language Models represent the computational foundation of modern generative AI. They are trained in two distinct phases: self-supervised pre-training on trillions of tokens of text and code, followed by post-training (RLHF, DPO, and supervised fine-tuning) to align responses with human instruction and safety boundaries.
In modern production environments, LLMs act as cognitive reasoning cores. Instead of just answering static questions, models interface with external APIs, execute Python code in sandboxes, and inspect databases to solve multi-step operational tasks.