Question 14 of 48

Why do jailbreaks work at all, given that a model has been trained to refuse harmful requests?

Written from the published competencies and our own syllabus, written from the published curricula of BlueDot Impact, the Center for AI Safety, DeepMind and Stanford. Not actual exam questions, and not affiliated with or endorsed by Product Digest.