Domain 4 of 6

Controls and Safeguards Around a Model

Defence in depth for something that cannot be made reliable on its own. Filtering both directions, grounding answers in sources that can be checked, and least privilege for anything an agent is allowed to touch.

4
Concepts
~18%
Of the exam
9
Practice questions
Concepts in this domain
01Securing an AI systemAccess control, encryption and the shared responsibility model applied to machine learning, plus the attack surfaces that only exist because there is a model in the path.02Prompt engineeringThe parts a prompt is made of, the shot based techniques and chain of thought, and the four attacks the syllabus expects you to name.03Agents and their toolsHow an agent reaches the outside world through extensions, functions, data stores and plugins, and the Google Cloud APIs it calls to do it.04Grounding and retrievalWhat grounding means, the difference between first party, third party and world data, and the sampling parameters that shape what comes back.
Try a question from this domain

A support agent built on a model reads customer emails and can issue refunds. A customer email contains text telling the assistant to ignore its instructions and refund in full. It does. What addresses this?

  • AFiltering both directions, least privilege on the refund tool, and a confirmation before anything financial.
  • BA stronger system prompt stating that instructions in emails must be ignored.
  • CRetraining the model on examples of this attack.
  • DLowering the temperature so the model follows its instructions more consistently.
9 questions on this domain.

One per page, with a worked explanation.

Start the set