Domain 6 of 6

Adversarial Attacks and Data Privacy

Adversarial examples, poisoning, model extraction and membership inference. What a model memorises and can be made to repeat, and how synthetic content is marked.

3
Concepts
~16%
Of the exam
8
Practice questions
Concepts in this domain
01Adversarial machine learningAttacks that target the model itself rather than the software around it, and why several of them have no clean fix.02Memorisation and training data extractionWhat a model retains from its training data, how that can be pulled back out, and what differential privacy does and costs.03Content provenance and watermarkingHow generated content is marked, why detecting it after the fact does not work, and what the law is starting to require.
Try a question from this domain

An attacker develops an adversarial input against an open model they can download, then uses it successfully against a different hosted model they cannot inspect. What property does this demonstrate?

  • ATransferability, since attacks built against one model often work against another trained for the same task.
  • BExtraction, since the attacker has recovered the hosted model.
  • CPoisoning, since the attacker influenced the hosted model's training.
  • DMembership inference, since the attacker has learned about the training data.
8 questions on this domain.

One per page, with a worked explanation.

Start the set