Question 1 · choose 1
A team evaluates a model that writes summaries of news articles. For each article it has a reference summary written by an editor, and it wants a metric that counts the overlapping word sequences (n-grams) between the model's summary and the reference. Which metric fits?
- ABERTScore
- BROUGE
- CToxicity score
- DClassification accuracy
Show the answer and why
ABERTScore
Incorrect
BERTScore compares embeddings of the generated and reference text, not overlapping word sequences.
BROUGE
Correct
ROUGE-N computes the n-gram overlap between the reference summary and the model's summary, and it is used to score summarization accuracy.
CToxicity score
Incorrect
A toxicity score comes from a toxicity detector model and measures harmful content, not agreement with a reference.
DClassification accuracy
Incorrect
Accuracy is the ratio of correctly classified items to all items. A summary is generated text, not a class label.
ROUGE counts shared n-grams with a reference; BERTScore compares meaning through embeddings. Both are used to judge summaries against a gold standard.
AWS documentation