A Survey on Predicting the Locations of Unseen Objects in Household Environments Using Vision-Language Models
Keywords:
Abstract
Locating a household object that is not currently in view is a different and harder problem than detecting one that is. A useful domestic robot or assistive system must predict where a misplaced remote, medication box, or set of keys is likely to be, drawing on semantic priors, memory of past observations, and knowledge of how objects move through a home over time. This review surveys the emerging body of work on AI-driven predictive household object localization, with particular attention to systems that combine vision-language models (VLMs) with heuristic and commonsense reasoning. We first formalize the predictive localization problem and distinguish it from object detection, object goal navigation, and semantic search, organizing the literature along three axes: known-object retrieval, unseen-object grounding, and temporal prediction. We then examine the two technical pillars of current frameworks, open-vocabulary perception built on models such as CLIP, Grounding DINO, and multimodal large language models, and reasoning modules that convert commonsense knowledge into ranked location predictions. A consolidated account of datasets, simulators, benchmarks, and evaluation metrics follows, including where present evaluation practice fails to test prediction of out-of-view, relocated objects. We review applications in domestic service robotics, assistive technology for elderly and memory-impaired users, and smart-home systems, and close with a research agenda covering personalization, temporal reasoning, uncertainty calibration, privacy, and deployment on resource-limited hardware.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License.