Devices limiting an AI system's undesirable behaviors, whose effectiveness remains imperfect against attempts to circumvent them.
Guardrails are the set of devices intended to limit the undesirable behaviors of an artificial intelligence system, by preventing it from producing dangerous, illegal, biased content or content contrary to the intended use. They take various forms, input and output filters, explicit rules, automatic moderation, constraints built in during training, or verification layers monitoring the model's responses. Their function is to make the system usable by reducing its risks, and they are an essential component of the responsible deployment of AI. Their fundamental limit is that they are neither perfect nor infallible. Circumvention techniques, jailbreak or prompt injection, evolve continually to defeat them, and a guardrail can also be weakened unintentionally, for example during a poorly controlled fine-tuning. This imperfection feeds a particular risk, that of a false sense of security, since an organization may believe itself protected by guardrails it overestimates, and relax its vigilance accordingly. For insurance and liability, the question becomes that of the expected standard, namely what level of guardrails an organization must reasonably put in place, and how to assess its diligence when a circumvented system causes harm.
A company deploys an assistant fitted with guardrails meant to block sensitive content. A user circumvents them through a clever jailbreak, and the system produces a harmful response, raising the question of whether the guardrails were equal to the risk.
garde-fous, guardrails, barrières de sécurité, filtres de sécurité