Recent Projects

AquaCap: A Training-Free Underwater Embodied Agent with Code-as-Policy

Xiaoshi Li1,*, Yule Xu1,*, Chunghiu Kong1, Yizhou Zhou2, Yang Liu2, Hao Yang1, Zihao Huang2,3,†, Yunxiao Shan1,4,†

1 Sun Yat-sen University   2 Dalian University of Technology   3 BIXOCEAN.INC   4 Southern Marine Science and Engineering Guangdong Laboratory

* Equal contribution. † Corresponding authors.

AquaCap uses Code-as-Policy to turn language instructions into underwater navigation and manipulation actions, without task-specific training or parameter updates. It addresses settings where collecting robot interaction data is expensive and visibility is often degraded.

A dual-layer agent produces plans and executable control code. Structured perception supplies semantic and geometric observations together with reliability information, while failure-aware memory helps the agent diagnose unsuccessful actions, revise its code, and replan in a closed loop.

AquaCap achieves a 66.43% success rate in simulation. Physical ROV experiments demonstrate autonomous grasping and object transport, including handling targets moved by hydrodynamic disturbances.

arXiv Paper PDF YouTube
ICRA 2027 · Under Review · arXiv preprint
AquaCap: Underwater Navigation and Manipulation
Labeled demo · 2× speed
IMATP: Iterative Multi-Agent Task Planning for Household Reorganization
A multi-LLM-agent framework for long-horizon task planning under partial observability

IMATP (Iterative Multi-Agent Task Planning) addresses the challenge of long-horizon household reorganization tasks in partially observable environments. Single-pipeline LLM planners suffer from goal confusion, context entanglement, and ungrounded actions when objects are occluded or the scene is partially observed.

IMATP decomposes planning into three specialized LLM agents with structured information flow: a Planner that performs goal decomposition and task scheduling, an Explore agent that actively gathers information about unseen or occluded objects, and an Execute agent that grounds high-level plans to the robot's atomic action space. The iterative perceive-plan-act loop prevents error accumulation and enables robust replanning when unexpected states are encountered.

Validated on real-world household robots (Unitree G1 humanoid with dexterous hands), IMATP demonstrates significant improvements over single-agent baselines on household reorganization benchmarks, achieving higher task success rates with more efficient exploration trajectories and fewer action errors.

Dec. 2025 - Mar. 2026
IMATP Demo: Multi-Agent Task Planning on Unitree G1 Humanoid
FOSGlove: Fiber Optic Sensor Glove for Dexterous Hand Teleoperation
A wearable data glove based on dye-doped polymer optical fiber sensors

FOSGlove is an innovative wearable data glove project designed for the teleoperation of robotic dexterous hands (specifically the Amazing Hand). It utilizes custom-developed R6G-doped NOA85 polymer optical fibers attached to the fingers. As the user bends their fingers, the optical strain (bending angle) causes light intensity loss within the fibers.

The system captures the continuous light intensity signals from the thumb, index, middle, and ring fingers using VEML6035 sensors multiplexed via TCA9548A. These analog signals are processed by an ESP32C3 MCU and transmitted to a host PC. A custom Feedforward Neural Network (FNN) then infers the precise bending angles (0°–90°) of the Proximal Interphalangeal (PIP) and Interphalangeal (IP) joints in real-time. Finally, these predicted joint angles drive a simulated or physical robotic hand with high precision and low latency via a Dora dataflow integration.

Jan. 2025 - Aug. 2026
FOSGlove Prototype
FOSGlove Prototype & Hardware Architecture
Amazing Hand
Amazing Hand