A technique of crafting prompts designed to bypass an AI model's safeguards and make it produce content that is normally prohibited.
Jailbreak refers to the set of techniques aimed at bypassing an AI model's safeguards, that is, pushing it, through cleverly crafted prompts, to produce content its designers specifically sought to prohibit, whether dangerous instructions, illicit content or disclosures it should refuse. The methods are varied and inventive, the staging of a role-play, nested instructions, the exploitation of ambiguities or of less-monitored languages, and they evolve constantly in response to the fixes vendors apply, in a cat-and-mouse dynamic with no end point. Jailbreak differs from prompt injection, whose spirit it shares, in that it targets the model's internal safeguards rather than the hijacking of an agent through external data. For insurance and risk management, it raises delicate questions of liability, since when a jailbroken model produces harmful content, fault is apportioned uncertainly between the user who forced the system, the vendor whose protections gave way and the company that deployed the tool without sufficient oversight.
A user disguises a prohibited request as a fiction exercise to get an assistant to supply instructions it would normally have refused. If those instructions cause harm, the apportionment of liability between the user, the vendor and the deployer remains largely uncertain.
jailbreak, contournement de garde-fous, débridage