
Prompt Injection and Jailbreaks
This chapter examines prompt injection and jailbreak attacks, which exploit a language model's inherent inability to distinguish between authoritative developer instructions and untrusted user data. It covers the mechani








