
רונן ברפמן
אקדמי בכיר
Visual Language State Machine Robot∗
Autonomous mobile robots in human environments must balance efficiency with social compliance and safety. While Vision-Language Models (VLMs) enable semantic scene understanding, existing frameworks often lack dynamic reconfiguration in response to changing scenarios, gestures, and hazards. We present zero-shot VLM integration without fine-tuning using a modular state-machine architecture that dynamically adapts navigation, gesture-triggered transitions, and safety-critical overrides. Experiments show robust gesture recognition, object detection, and scenario adaptability, with seamless behavior switching at expected transition points.
| שפת פרסום | אנגלית |
| דפים | 354-361 |
| סטטוס פרסום | פורסם - 01.01.2026 |
Keywords
Cognitive Robotics
Robot and Multi-Robot Systems
Task Planning and Execution
Vision and Perception
ASJC Scopus subject areas
Artificial Intelligence