Theoretical framework analyzing conflicts of interest between a principal and their agent, and the resulting economic costs when their objectives diverge.
Agency theory is an institutional microeconomics framework formalized notably by Jensen and Meckling in their landmark 1976 article in the Journal of Financial Economics. It models the relationship between a principal who delegates a task to an agent, in the presence of information asymmetry and diverging interests. The principal cannot perfectly observe the agent's actions (moral hazard) nor their pre-relationship characteristics (adverse selection). To align behaviors, the principal must bear agency costs: monitoring costs, bonding costs, and residual loss representing value destroyed by the irreducible divergence of interests. In the context of agentic AI, this theory is reinterpreted: algorithmic agents neutralize the traditional interest divergence (they have no union, no personal agenda), but create an information asymmetry of a radically new nature. The opacity of deep neural networks (black box) means the human principal cannot retrace the agent's reasoning in high-dimensional vector spaces. Algorithmic hallucination thus becomes the modern equivalent of residual loss: no longer an interest divergence, but a representational divergence between reality and the agent's output. This reversal of the information structure fundamentally alters the conditions for reliable delegation.
In an insurance company that has automated its claims processing, the claims director (principal) delegates file analysis to an LLM agent. Classical interest divergence disappears: the agent has no personal reason to bias its analysis. But information asymmetry persists in another form: the director cannot retrace why the agent rejected a particular file, which compromises compliance with legal obligations to justify rejection decisions.
agency theory, principal-agent problem, problème agent-principal, Jensen-Meckling