רונן ברפמן

אקדמי בכיר

Visual Language State Machine Robot∗

Elias Goldsztejn, Dan Rouven Suissa, Ronen Brafman

Autonomous mobile robots in human environments must balance efficiency with social compliance and safety. While Vision-Language Models (VLMs) enable semantic scene understanding, existing frameworks often lack dynamic reconfiguration in response to changing scenarios, gestures, and hazards. We present zero-shot VLM integration without fine-tuning using a modular state-machine architecture that dynamically adapts navigation, gesture-triggered transitions, and safety-critical overrides. Experiments show robust gesture recognition, object detection, and scenario adaptability, with seamless behavior switching at expected transition points.

שפת פרסום אנגלית
דפים 354-361
סטטוס פרסום פורסם - 01.01.2026

Keywords

Cognitive Robotics
Robot and Multi-Robot Systems
Task Planning and Execution
Vision and Perception

ASJC Scopus subject areas

Artificial Intelligence
גישה למסמך
10.5220/0014310800004052
קבצים וקישורים אחרים
Link to publication in Scopus