CAMBRIDGE — Researchers at MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) have developed a new method called Masked Inverse Reinforcement Learning (Masked IRL). This approach was designed to improve how artificial intelligence systems interpret ambiguous human instructions using minimal demonstration data.
The Masked IRL system utilizes two large language models. One model works to clarify ambiguous user instructions based on demonstration data. The second model identifies and disregards environmental details that are not relevant to a task. The system requires nearly five times less demonstration data compared to other similar baseline methods.
Minyoung Hwang, an MIT PhD student and CSAIL researcher, is a co-author on the paper. "Our approach could come in handy when a human interacts with a robot but doesn't want to spell out all the details of a task," Hwang said. "We're minimizing human effort by enabling machines to get to the bottom of what users really want."
In various tests, Masked IRL identified users' unstated preferences up to 15 percent more often than comparable baselines. These results were observed in both 3D and real-world demonstrations. For instance, in real-world tests, a robotic arm was trained using 50 kinesthetic demonstrations. After receiving the general instruction to "stay away," the arm successfully moved a cup toward a human while avoiding a computer.
The project is scheduled for presentation at the 2026 IEEE International Conference on Robotics and Automation in June. Funding for this research was provided in part by the Tata Group through the MIT Generative AI Impact Consortium Award and the Department of Defense.
Why It Matters
The development of Masked IRL addresses challenges in human-AI interaction by enabling AI systems to interpret ambiguous commands with less data. By requiring less demonstration data, this method could reduce the effort needed to train AI systems for practical applications. This approach aims to improve the efficiency and adaptability of AI in scenarios where precise, detailed instructions from human operators may not always be feasible, fitting into ongoing efforts to make AI systems more intuitive and responsive to user needs.
forum Comments (0)
No comments yet. Be the first to comment.