Autonomous reasoning for a humanoid with language models and semantic graphs

A language model can turn "clean the table" into a reasonable list of steps, but it does not know what is on the table or what the robot can do. This project connects the two. A language model proposes goals, and two graphs, one holding what the robot knows to be possible and one holding what it sees at that moment, determine which actions are feasible. In 140 executions with achievable actions the system produced a valid plan in 81% of cases, and in no case did it generate a plan containing a prohibited action.
Context
There are two common approaches to making a robot carry out natural-language instructions. Symbolic planners are verifiable but rigid, and fail when the instruction does not fit their rules. End-to-end trained models are flexible, but their behavior is hard to bound. In a humanoid operating near people, that lack of guarantees is a practical problem.
The work opts for an intermediate architecture in which the language model proposes and the symbolic structure decides.
Architecture
System flow, from the camera and the user's instruction to the action plan
Perception. An RGB-D camera, which is the sensor available on the Unitree G1 humanoid the system was designed for. Objects are detected with an open-vocabulary model.
Semantic graph. Holds long-term knowledge: which objects exist, which actions each one admits, and which are prohibited. The robot's actions are defined in a JSON file with their preconditions and effects, which allows the system to be moved to another platform.
Semantic graph: object-action-object relations and constraints
Perceptual graph. Holds the context of the moment: the objects detected in the current frame and their spatial relations.
Perceptual graph of one frame: detected instances and spatial relations
Reasoning. The language model (Llama 3.1 with 8 billion parameters) interprets the instruction and proposes goals. Goals are prioritized, checked against both graphs and turned into a sequence of actions. If there is no plan, fallback mechanisms reformulate the goal, replan, or terminate without acting.
Ethical constraints are not a filter applied at the end. They are built into the semantic graph, so that an unsafe action is not part of the feasible action space.
Evaluation
Fifteen instructions were used, each repeated 10 times, on an office desk simulated in Gazebo and on a real one. The instructions cover direct manipulation, multi-step tasks, conditionals, abstract commands ("write me an essay"), social tasks ("greet my guests"), a deliberately prohibited command, and one in Spanish.
Desk environment simulated in Gazebo
Performance was measured with three indicators: task success, the number of reasoning steps as a hardware-independent measure of complexity, and the type of failure.
Results
- Success. 113 of 140 executions with achievable actions, that is, 81%.
- Safety. The language model occasionally proposed ethically questionable goals. In every case they were discarded before execution.
- Failures. They concentrate in abstract or ambiguous commands and come mostly from the language model: poorly formulated or invented goals. Perception errors, from unstable or missing detections, are the second cause.
Distribution of failures by task category: perception, planning and language model
Success rate against the average number of reasoning steps
When the system fails, it does so without acting. Several failures also occurred after part of the task had been completed.
What is missing
"Success" measures that a complete, valid plan is generated, not that a physical robot executes it. The scope of the validation is the reasoning, with real and simulated perception, and it does not include manipulation with the humanoid. The robot is assumed stationary in front of its workspace. Performance depends on how well the actions are described and on the language model used, which is small; commands in English worked better than in Spanish. Continuous autonomy, in which the robot generates its own goals indefinitely, is raised but not solved.
How it fits in Robiolab
This project opens the decision level in the lab, which had not been addressed so far: the group's other work deals with the robot's body, its control, or decoding human intention. It shares with them the interest in interaction with people, and its most reusable contribution is the way it treats safety as a property of the structure rather than as an after-the-fact check.
