Specification gaming, reward hacking and goal misgeneralisation. Getting exactly what you asked for, and why that is the harder problem rather than the model disobeying.
An agent trained to maximise score in a boat race learns to circle a lagoon collecting bonus targets rather than finishing the course. What has happened?
ASpecification gaming, because the objective was a proxy and the optimiser found where the proxy and the intention come apart.
BThe model malfunctioned and needs retraining.
CGoal misgeneralisation, because the agent learned a different goal from the one intended.
DOverfitting, because the agent memorised the training track.