GIS

Large Language Model

A generative AI model trained on large amounts of text that produces language by predicting the next word. NIST warns it can state false content with confidence.

Detailed Definition

A large language model (LLM) is a generative AI model trained on very large amounts of text to read and write language. GAO lists LLMs among generative AI tools that "can receive a prompt in plain language and generate an output (e.g. text, an image, or a video) that is statistically representative of the training data" (GAO-24-106946, 2024). NIST's glossary says generative pre-trained transformer (GPT) models, "pre-trained through self-supervised learning on large data sets of unlabelled text," are "the current predominant architecture for large language models" (NIST AI 100-2e2025).

Related terms

  • Generative AI: "The class of AI models that emulate the structure and characteristics of input data in order to generate derived synthetic content" (NIST SP 800-218A).
  • Foundation model: "models trained on broad data using self-supervised learning that can be adapted such as through fine-tuning for a variety of downstream tasks" (NIST AI 100-2e2025).
  • Retrieval-augmented generation (RAG): a model "paired with a separate information retrieval system." GAO says RAG "enhances the accuracy and reliability of a generative AI model by retrieving contextual information from sources not included in the initial training data."

Confabulation

NIST's Generative AI Profile (NIST AI 600-1, 2024) defines confabulation as "The production of confidently stated but erroneous or false content (known colloquially as 'hallucinations' or 'fabrications') by which users may be misled or deceived." NIST explains that confabulations "are a natural result of the way generative models are designed," because "LLMs predict the next token or word in a sentence or phrase." Outputs "may also include confabulated logic or citations," and NIST notes that "legal confabulations have been shown to be pervasive in current state-of-the-art LLMs."

GAO adds that models "inherit the characteristics of the data they are trained on, which can include bias, inaccuracy, and impropriety."

Over-reliance

NIST warns that people "may over-rely on GAI systems or may unjustifiably perceive GAI content to be of higher quality than that produced by other sources." It calls this automation bias; see Automation.

Language processing more broadly is covered under Natural Language Processing.

Why it matters for land and mining claim records

Title and mining claim work turns on exact citations, dates, names, and legal descriptions, the kind of detail NIST says an LLM can state wrongly with full confidence. An LLM summary of a deed, a case file, or a regulation is a draft. Every citation, date, and figure in it should be checked against the source document before anyone relies on it.