A team is evaluating a summarisation feature and needs an automated metric to detect whether a new model version has made summaries worse. Which metric is the conventional choice for summarisation, and what is its limitation?
Written from the published competencies and the AIF-C01 exam guide, version 1.4, which lists five content domains with published weightings and a task statement breakdown under each. Not actual exam questions, and not affiliated with or endorsed by Amazon Web Services.