Interpretability researchers identify a feature inside a model and then amplify it, changing the model's output in the predicted direction. Why does the steering result matter?
Written from the published competencies and our own syllabus, written from the published curricula of BlueDot Impact, the Center for AI Safety, DeepMind and Stanford. Not actual exam questions, and not affiliated with or endorsed by Product Digest.