EXIMO: VLM Guided Exploration of VLA Policies

· AstraNL · robotics

A new method called EXIMO uses vision-language models to direct the exploration of vision-language-action policies. It targets the challenge of fine-tuning large VLA models, which are typically trained through behavior cloning on extensive teleoperation datasets, so that robot systems can adapt to new tasks without full retraining.

This matters for robotics and automation coordination because current VLA-based manipulation policies have scaled to billions of parameters yet remain difficult to update efficiently on deployed hardware. Integrators working with logistics or multi-robot setups often need policies that adjust to task variations using limited new data rather than repeated large-scale data collection.

The method focuses specifically on the exploration phase during fine-tuning of existing VLA policies in manipulation domains.