Question 1 · choose 1
A company has two million support articles. It wants a search feature that finds articles with the same meaning as a customer's question even when the article uses different words. What should the company create from the articles to make this possible?
- AEmbeddings, numeric vectors that can be compared for similarity
- BToken lists that are matched against the question word for word
- CA higher temperature setting so that the model considers related words
- DA model with a context window large enough to hold every article
Show the answer and why
AEmbeddings, numeric vectors that can be compared for similarity
Correct
An embedding turns input into a vector of numbers so that objects can be compared; sentences can be compared to determine how similar their meaning is, even when the wording differs.
BToken lists that are matched against the question word for word
Incorrect
A token is a unit of text a model reads, such as a word or part of a word. Matching tokens exactly still needs the same words, which is the limitation the company wants to remove.
CA higher temperature setting so that the model considers related words
Incorrect
Temperature changes how random a model's generated output is. It has no role in finding which stored articles are similar to a question.
DA model with a context window large enough to hold every article
Incorrect
The context window is the number of tokens a model can take into account at one time. Two million articles will not fit, and a bigger window does not provide a similarity search.
Semantic search compares meaning, and meaning is compared through embedding vectors: similar texts get vectors that are close to each other.
AWS documentation