A support agent built on a model reads customer emails and can issue refunds. A customer email contains text telling the assistant to ignore its instructions and refund in full. It does. What addresses this?
Written from the published competencies and our own syllabus, written from the published curricula of BlueDot Impact, the Center for AI Safety, DeepMind and Stanford. Not actual exam questions, and not affiliated with or endorsed by Product Digest.