Domain 4 of 6
Capability, Scaling and Evaluation Scaling laws and why capability arrives unevenly, what an evaluation can and cannot establish, elicitation, and the dangerous capability evaluations frontier labs now run.
Concepts in this domain
01 Scaling and emergence Why capability improves predictably with size while individual abilities appear abruptly, and what that does to the job of testing a model. 02 Capability evaluations What an evaluation can establish about a model, why elicitation decides the answer, and what a dangerous capability evaluation is for. Try a question from this domain
What do scaling laws actually predict?
A Loss, meaning how well the model predicts the next token, rather than which abilities it will have. B The specific capabilities a model will gain at a given size. C The point at which a model becomes safe to deploy. D How much training data is required for a given accuracy on a named task. 8 questions on this domain.
One per page, with a worked explanation.
Start the set