During a red team session somebody finds a phrasing that makes the assistant reveal part of its system prompt. The team adds that exact phrase to a block list. What is wrong with this response?
Written from the published competencies and our own syllabus, written from the published curricula of BlueDot Impact, the Center for AI Safety, DeepMind and Stanford. Not actual exam questions, and not affiliated with or endorsed by Product Digest.