Domain 3 of 6

Failures of Specification and Generalisation

Specification gaming, reward hacking and goal misgeneralisation. Getting exactly what you asked for, and why that is the harder problem rather than the model disobeying.

2
Concepts
~17%
Of the exam
8
Practice questions
Concepts in this domain
01Specification gaming and goal misgeneralisationTwo ways a system does exactly what it was trained to do and still fails, and why that is harder to fix than disobedience would be.02Responsible AI in practiceBias, fairness, robustness, safety and veracity treated as things you can measure and act on, plus the legal exposure that arrives with generative output.
Try a question from this domain

An agent trained to maximise score in a boat race learns to circle a lagoon collecting bonus targets rather than finishing the course. What has happened?

  • ASpecification gaming, because the objective was a proxy and the optimiser found where the proxy and the intention come apart.
  • BThe model malfunctioned and needs retraining.
  • CGoal misgeneralisation, because the agent learned a different goal from the one intended.
  • DOverfitting, because the agent memorised the training track.
8 questions on this domain.

One per page, with a worked explanation.

Start the set