רונן ברפמן

אקדמי בכיר

Visual Language State Machine Robot

Elias Goldsztejn, Dan Rouven Suissa, Ronen Brafman

Autonomous mobile robots in human environments must balance efficiency with social compliance and safety. While Vision-Language Models (VLMs) enable semantic scene understanding, existing frameworks often lack dynamic reconfiguration in response to changing scenarios, gestures, and hazards. We present zero-shot VLM integration without fine-tuning using a modular state-machine architecture that dynamically adapts navigation, gesture-triggered transitions, and safety-critical overrides. Experiments show robust gesture recognition, object detection, and scenario adaptability, with seamless behavior switching at expected transition points.

שפת פרסום אנגלית
דפים 354-361
סטטוס פרסום פורסם - 01.01.2026

Keywords

Cognitive Robotics
Robot and Multi-Robot Systems
Task Planning and Execution
Vision and Perception

ASJC Scopus subject areas

Artificial Intelligence
גישה למסמך
10.5220/0014310800004052
קבצים וקישורים אחרים
Link to publication in Scopus