Product Digest AI Safety Foundation Certification

Our own certification, covering what the AI safety field is actually about. Six modules and an assessment of 48 questions. Sits beneath the Practitioner certification.

At a glance
Cost
Free No exam fee, no course fee, and nothing to upgrade to.
Format
Six modules and one assessment of 48 questions covering all of them. The reading is open to anyone. The assessment needs an account.
Pass mark
70 per cent, or 34 of 48, across the whole assessment rather than per module.
Prerequisites
None. It assumes no machine learning background and does not ask you to read any maths.
Renewal
None. The material will date as the field moves, so a result records what you knew when you took it. The Practitioner certification is the next one, covering what to do about all of this when you ship something.
Based on
The our own syllabus, written from the published curricula of BlueDot Impact, the Center for AI Safety, DeepMind and Stanford
What it covers

What it covers. This is ours. Product Digest issues it, no external body accredits it, there is no proctor and nobody checks who is at the keyboard. It is worth what the reading is worth.

Foundational and Practitioner divide the subject rather than repeat it. This one explains how a model is made to behave, where that goes wrong, what evaluation and interpretability can establish, and what the adversarial and privacy surface actually is. Practitioner covers what to do about all of it when you ship something.

It is written for people with no machine learning background. There is no maths, and nothing here requires you to have trained a model or to intend to.

Every question is a scenario or asks you to tell two ideas apart. None asks for a definition.

The syllabus6 domains · free

Product Digest’s own domains. Each opens its own page listing the concepts beneath it, and each concept has a page of its own.

01Fundamentals of AI SafetyThe three kinds of risk the field separates, misuse, accidents and structural harm, why they need different remedies, and where the disagreements inside the field actually lie.What AI safety means · How AI features fail · What generative AI is good and bad at3 concepts
~19%
02Training and Alignment of ModelsPre training, supervised fine tuning, reinforcement learning from human feedback and constitutional methods. What each stage adds, and why a helpful model is a trained artefact rather than an emergent one.How models are trained to behave · Foundation models and how they work · Supervised, unsupervised and reinforcement learning3 concepts
~19%
03Failures of Specification and GeneralisationSpecification gaming, reward hacking and goal misgeneralisation. Getting exactly what you asked for, and why that is the harder problem rather than the model disobeying.Specification gaming and goal misgeneralisation · Responsible AI in practice2 concepts
~17%
04Capability, Scaling and EvaluationScaling laws and why capability arrives unevenly, what an evaluation can and cannot establish, elicitation, and the dangerous capability evaluations frontier labs now run.Scaling and emergence · Capability evaluations2 concepts
~17%
05Interpretability and TransparencyInterpretability, what it has established and what it has not, and the documentation that stands in for understanding while the research catches up.Interpretability · Transparency and explainability2 concepts
~12%
06Adversarial Attacks and Data PrivacyAdversarial examples, poisoning, model extraction and membership inference. What a model memorises and can be made to repeat, and how synthetic content is marked.Adversarial machine learning · Memorisation and training data extraction · Content provenance and watermarking3 concepts
~16%
Try a question

A company finds that its recruitment model has been used by a manager to screen candidates in a way the company never authorised, using a workflow the manager built themselves. The model works exactly as designed. Which category of risk is this?

  • AMisuse, because the system worked as intended and somebody used it to cause harm.
  • BAn accident, because the model was applied to a task it was not evaluated for.
  • CStructural harm, because it reflects competitive pressure inside the company.
  • DMisalignment, because the model pursued an objective its designers did not intend.
48 questions across the 6 domains.

One per page, with a worked explanation.

Start the set