Conference item
VIKI‑R: coordinating embodied multi-agent cooperation via reinforcement learning
- Abstract:
- Coordinating multiple embodied agents in dynamic environments remains a core challenge in artificial intelligence, requiring both perception-driven reasoning and scalable cooperation strategies. While recent works have leveraged large language models (LLMs) for multi-agent planning, a few have begun to explore vision-language models (VLMs) for visual reasoning. However, these VLM-based approaches remain limited in their support for diverse embodiment types. In this work, we introduce VIKI-Bench, the first hierarchical benchmark tailored for embodied multi-agent cooperation, featuring three structured levels: agent activation, task planning, and trajectory perception. VIKI-Bench includes diverse robot embodiments, multi-view visual observations, and structured supervision signals to evaluate reasoning grounded in visual inputs. To demonstrate the utility of VIKI-Bench, we propose VIKI-R, a two-stage framework that fine-tunes a pretrained vision-language model (VLM) using Chain-of-Thought annotated demonstrations, followed by reinforcement learning under multi-level reward signals. Our extensive experiments show that VIKI-R significantly outperforms baselines method across all task levels. Furthermore, we show that reinforcement learning enables the emergence of compositional cooperation patterns among heterogeneous agents. Together, VIKI-Bench and VIKI-R offer a unified testbed and method for advancing multi-agent, visual-driven cooperation in embodied AI systems.
- Publication status:
- Published
- Peer review status:
- Peer reviewed
Actions
Access Document
- Files:
-
-
(Preview, Version of record, pdf, 8.0MB, Terms of use)
-
- Publisher copy:
- 10.52202/085713-3935
Authors
+ Science and Technology Commission of Shanghai Municipality
More from this funder
- Funder identifier:
- https://ror.org/03kt66j61
- Publisher:
- Curran Associates
- Host title:
- Advances in Neural Information Processing Systems 38
- Volume:
- 38
- Pages:
- 130799-130830
- Publication date:
- 2026-08-06
- Event title:
- 39th Annual Conference on Neural Information Processing Systems (NeurIPS 2025)
- Event location:
- San Diego, CA, USA
- Event website:
- https://neurips.cc/Conferences/2025
- Event start date:
- 2025-11-30
- Event end date:
- 2025-12-05
- DOI:
- ISSN:
-
1049-5258
- EISBN:
- 9798331338275
- Language:
-
English
- Pubs id:
-
2451981
- Local pid:
-
pubs:2451981
- Deposit date:
-
2026-10-05
- ARK identifier:
Terms of use
- Copyright holder:
- Kang et al and NeurIPS
- Copyright date:
- 2025
- Rights statement:
- © (2026) by individual authors and Neural Information Processing Systems Foundation Inc. All rights reserved.
- Notes:
- This paper was presented at the 39th Annual Conference on Neural Information Processing Systems (NeurIPS 2025), 30th November - 5th December 2025, San Diego, CA, USA.
If you are the owner of this record, you can report an update to it here: Report update to this record