Architecture
The system is layered so that foundation models can only influence the car through well-defined, checked interfaces:
- Perception and localization: ZED camera and a particle filter.
- Control: a Stanley controller for stable lane tracking and path following.
- Decision making: a rule-based node augmented with VLM and LLM reasoning.
- Safety guard: a finite state machine between the foundation models and the low-level controller, so model outputs are always checked before they reach the car.
- Interface: a ROS bridge connecting the car to an external server that runs the models.
What it can do
- VLM overtaking: a vision-language model looks at the camera image and decides when to overtake.
- Voice commands: spoken commands go through speech recognition to an LLM that picks the driving behavior. Asked "I am driving in the UK, which lane should I be in?", the car moves from the right lane to the left. Told "I am in a hurry", it switches to an aggressive driving profile.
Validation
The full stack was integrated and tested end to end on the real F1TENTH car, checking navigation and control performance under dynamic conditions.