Company: Deloitte USI_5sep
Difficulty: medium
A dataset is split into training, validation, and test sets. After several iterations, a model performs very well on training and validation data but noticeably worse on the untouched test set. Which interpretation is most appropriate?