The maximum amount of text a model can consider at once, an engineering constraint that directly drives the cost and confidentiality of a use case.
The context window is the maximum amount of text a model processes at once, taken in and considered before any reasoning. Its growth, from a few thousand units to several hundred thousand within a few years, changed how applications are built: practice moved from splitting documents and presenting only extracts, a method requiring a prior search and inheriting its errors, to presenting the whole document. Three practical consequences follow, and they bear directly on underwriting a risk. The cost of a call grows with the amount of context supplied, so a large window invites expensive uses no budget anticipated. Confidentiality changes scale, an entire contract transmitted replacing an extract. And reliability does not follow linearly, evaluations showing that attention to information placed in the middle of a very long context declines, so a large window is no guarantee that everything was read.
Google's announcement in February 2024 of a one million token context window for its Gemini 1.5 model marked a jump of roughly two orders of magnitude over the models in wide use two years earlier.
context window, longueur de contexte, contexte long