Prompt
You are a robotic arm composed of seven links with a white movable chassis in given room. You need to complete the tasks according to human instructions.
We provide an Available_Actions set and the corresponding explanations for each action. Each step, you should select one action from Available_Actions.
I will provide you with images from two different perspectives:
1.HEAD CAMERA: Provides a Head View (human-like perspective) showing the entire scene from the robot's eye level. This perspective is used to locate the position of the container.
2.WRIST CAMERA: Mounted on the robotic arm's end effector, this camera moves in real-time with the arm and provides a detailed close-up perspective for precise manipulation.
Use the exact object names from the list: ['bread', 'candy', 'plate', 'bar', 'table_shelf']
This is your task: My snacks fell and got mixed up, please pick up the edible ones and put them onto the plate.
After executing the action, you will receive a new observation image showing the updated state. To complete your task, you should output your action from the Available Action Library.
Available Action Library:
- pick <object>: Move from the current position to a suitable location for grasping the target object and close the gripper to perform the grasp.
- place to <container>: Move from the current position to a suitable location for placing the target object to target container and open the gripper to perform the place.
- lift: Lift your end effector vertically.
- pull: Pull your hand out parallel while maintaining the gripper state.
- push: Push your end effector in while maintaining the gripper state.
- observe: Reset the end-effector to the default position to gain a clear, wide-angle view or to stabilize the camera orientation before executing navigation skills.
- open_door: Given that the gripper has already grasped the handle, open the door, then release the handle.
- close_door: Push or pull the door along its path until fully closed.
- recall <step_x>: Review the key frames of a previous step (e.g. step_0) if you are unsure whether it succeeded.
- end: If you think you have completed the task, please output 'end'.
Please follow these output rules:
- If you find that the action is not executed successfully in the subsequent stage (for example, the object is not picked up successfully after executing pick), you can try again.
- If your action is 'pick', output the bounding box coordinates of the target object in the wrist camera image.
- If your action is 'place', output the bounding box coordinates of the target container in the head camera image.
- If your action is 'recall', output the step index you want to review instead of a bounding box (see Output Format Rules below).
- For all other actions, do not output any bounding box coordinates.
Output Format Rules:
- For pick or place actions: You must output both the action and the normalized bounding box. Format: action: pick <object>, box_2d: [y_min, x_min, y_max, x_max]
- For recall actions: You must output the action and the target step index on separate lines. Format:
action: recall
target: step_x
- For all other actions (e.g., lift, pull, open_door, etc.): Output only the action. Do not include box_2d. Format: action: <action_name>
- Reasoning : For every action, you must provide a brief reason for your decision. This should explain why the target object was chosen or why the specific action is necessary.
Constraints:
-The box_2d must be [y_min, x_min, y_max, x_max] normalized to 0-1000.
-Do not include any image data, base64 strings in the output.
Examples:
- Case 1 (Action with BBox):
action: pick bell_pepper, box_2d: [y_min, x_min, y_max, x_max]
reasoning: your reasoning
- Case 2 (Action without BBox):
action: lift
reasoning: your reasoning
- Case 3 (Action with BBox):
action: place to plate, box_2d: [y_min, x_min, y_max, x_max]
reasoning: your reasoning
- Case 4 (Recall action):
action: recall
target: step_0
reasoning: your reasoning
These are your initial observation. Please analyze the scene and select a manipulation action.
1. HEAD CAMERA:<image_1>
2. WRIST CAMERA:<image_2>
Response
action: pick bread, box_2d: [529, 658, 621, 774]
reasoning: The bread is an edible snack that needs to be moved to the plate. I am starting by picking up the bread, which is clearly visible in the wrist camera view.
Wrist Camera (Pick)

img_0_wrist_pick_pts.png