RMMBench: A Comprehensive Benchmark for
Robotic Mobile Manipulation

1Shandong University    2Meituan Group

Abstract

Although the advancement of vision-language models (VLMs) has endowed robots with enhanced environmental understanding and task reasoning, a comprehensive evaluation methodology is important to advance the integration of VLMs in robotic navigation and manipulation. However, current benchmarks lack a comprehensive method to evaluate diverse robotic tasks, and evaluation metrics remain relatively constrained, making it difficult to assess the embodied capabilities of VLMs in a thorough and fine-grained manner.

To address this issue, we propose RMMBench, an evaluation benchmark that requires robots to understand language instructions and perform long-horizon tasks in continuous spaces. RMMBench seamlessly integrates high- and low-level embodied tasks into a unified framework, constructing a “navigation–manipulation” task suite comprising 70 canonical task scenarios that range from localized manipulation to long-horizon composite navigation. The results reveal that leading VLMs still face major challenges in spatial localization when coping with mobile manipulation tasks, and also highlight the necessity of enhancing the spatial perception capability of robots during long-period interactions.

Overview of the RMMBench (animated)
Overview of the RMMBench. RMMBench is a comprehensive mobile manipulation benchmark with a standardized skill library to evaluate the embodied performance of vision-language models.

The RMMBench Benchmark

RMMBench requires VLM-based models to execute continuous-space, long-horizon mobile manipulation tasks based on language instructions and raw visual feedback. Evaluating this capability requires tight integration of sequential planning, spatial grounding, and closed-loop visual verification.

Overview of the RMMBench evaluation pipeline
Overview of the RMMBench evaluation pipeline. It ingests head and wrist camera views, complemented by an optional bird’s eye view (BEV) modality. Navigation is executed over a Hanan grid dynamically bounded by the workspace geometry. For manipulation tasks, the pipeline overlays algorithmic candidates from GraspNet and a human-verified baseline fallback onto the wrist view for the VLM to select. Decision-making is conditioned on the historical context.

Video Demonstrations

The demonstrations are organized into two parts: manipulation tasks and navigation tasks.

Manipulation Tasks

cook_chicken_breast_0
cook_sausage_0
fix_burnt_bread_0
pick_bottle_opener_0
pick_chip
pick_chocolate
pick_no_sugar_drink
pick_non_alcoholic_drink
pick_whisk_0
select_bagged_snacks
select_bread_toaster_0
select_bread_toaster_1
select_bread_toaster_2
select_same_cake
store_yogurt
uncover_fruit
wash_place_carrot
wash_place_cucumber

Navigation Tasks

Interactive Rollout Logs

These interactive execution logs record the complete multi-step decision process of a VLM agent (Gemini-3-Flash) on RMMBench tasks, including the natural-language instruction, per-step visual observations, skill selections, bounding-box predictions, and environmental feedback. Each panel below is a scrollable mini window — scroll up and down inside a panel to walk through the episode step by step, or click “Open full log” to view the complete page.

Manipulation Tasks

select_snacks · Gemini-3-Flash Open full log ↗
Scroll inside the window to inspect each step.
uncover_fruit · Gemini-3-Flash Open full log ↗
Scroll inside the window to inspect each step.

Composite Navigation Tasks

organize_snacks_to_tray · no BEV, no LiDAR rays Open full log ↗
Composite navigation without BEV and without LiDAR auxiliary rays — purely egocentric decision-making.
organize_snacks_to_tray · no BEV, with LiDAR rays Open full log ↗
Composite navigation without BEV but with LiDAR auxiliary rays overlaid on the view.

Citation

@article{li2026rmmbench,
  title   = {RMMBench: A Comprehensive Benchmark for Robotic Mobile Manipulation},
  author  = {Li, Huapeng and Feng, Fuxiang and Fan, Jinqiu and Yang, Shuo and
             Chen, Fengjiao and Cao, Xuezhi and Song, Ran and Zhang, Wei},
  journal = {arXiv preprint arXiv:XXXX.XXXXX},
  year    = {2026}
}

Acknowledgments

This work was supported in part by the National Natural Science Foundation of China under Grant U22A2057, in part by the Key R&D Program of Shandong Province under Grant 2025CXGC010210, and in part by the Municipal-University Collaborative Development Project of Jinan under Grant JNSX2025002.