A benchmark that tests if vision-language models can spot things they never learned.
VLOSE is a pop quiz with a weird fruit. Saying “apple” is easy. Saying “I have no clue” earns the gold star.
It shows a VLM new objects. It checks if the model admits when it does not know.
VLM
VLOSE tests if a VLM can spot an unknown object.
Hallucination
VLOSE catches models that wrongly name strange objects.
Model Calibration
VLOSE checks if a model admits when it is unsure.