Embodied Intelligence Security with Vision-language Models: A Survey
-
Abstract
Embodied intelligence (EI), integrating vision-language models (VLMs) with action-oriented capabilities, presents transformative potential for autonomous systems. However, deploying VLMs in safety-critical applications like self-driving cars and collaborative robots introduces significant challenges. Key concerns include adversarial attacks and content manipulation, through which maliciously altered visual or textual inputs could lead to incorrect interpretations and hazardous actions in EI systems. Furthermore, integrating VLMs into embodied agents amplifies these risks, as real-world physical interactions introduce complex safety-critical scenarios where errors in perception or decision-making can have immediate and severe consequences. While VLM integration amplifies safety concerns, these models simultaneously offer a foundation for defensive strategies aimed at enhancing the reliability of EI systems. Additionally, VLM-based EI systems raise critical ethical concerns regarding accountability, embedded biases and societal displacement. This review systematically analyzes current applications, safety risks, protection methods and ethical implications, while proposing research directions to advance trustworthy EI systems.
Full Text Translation (powered by iFLYTEK)
-
-