CAMBRIDGE, MASSACHUSETTS — Researchers at the Massachusetts Institute of Technology developed an automated framework called SEED-SET to systematically evaluate the ethical alignment of autonomous systems, balancing measurable outcomes with qualitative human values. The research will be presented at the International Conference on Learning Representations.
The SEED-SET framework tackles the evaluation challenge by splitting the problem into a hierarchical structure with an objective model for tangible metrics and a subjective model for stakeholder judgments. It does not require pre-existing evaluation data and adapts to multiple objectives, using an adaptive process to select the best scenarios for further evaluation and streamlining a process that typically requires manual effort.
"We can insert a lot of rules and guardrails into AI systems, but those safeguards can only prevent the things we can imagine happening. It is not enough to say, 'Let's just use AI because it has been trained on this information.' We wanted to develop a more systematic way to discover the unknown unknowns and have a way to predict them before anything bad happens," said Chuchu Fan, an associate professor in the MIT Department of Aeronautics and Astronautics and a principal investigator in the MIT Laboratory for Information and Decision Systems.
For its subjective assessment, the framework uses a large language model as a proxy for human evaluators to capture and incorporate stakeholder preferences. Researchers encode the preferences of each user group into a natural language prompt for the large language model, which then uses the prompts to compare two scenarios and select the preferred design based on ethical criteria. The framework uses the selected scenarios to simulate the overall system and guide its search for the next candidate scenarios.
"The objective part of our approach is tied to the AI system, while the subjective part is tied to the users who are evaluating it. By decomposing the preferences in a hierarchical fashion, we can generate the desired scenarios with fewer evaluations," said Anjali Parashar, a mechanical engineering graduate student at MIT.
Researchers tested the framework on realistic autonomous systems, including an AI-driven power grid and an urban traffic routing system. In the power grid domain, the framework pinpointed power distribution cases that prioritize higher-income areas during peak demand, leaving underprivileged neighborhoods more prone to outages. In a power grid system, evaluating the ethical alignment of an AI model's recommendations while considering all objectives is especially difficult.
The framework generated more than twice as many optimal test cases as baseline strategies in the same amount of time while uncovering many scenarios other approaches overlooked. The test cases produced by the framework can show scenarios where autonomous systems align with human values, as well as scenarios that fall short of ethical criteria. When researchers shifted user preferences, the set of scenarios generated by the framework changed accordingly.
"We don't want to spend all our resources on random evaluations. It is very important to guide the framework toward the test cases we care the most about," said Yingke Li, a postdoctoral researcher in the Department of Aeronautics and Astronautics at MIT.
forum Comments (0)
No comments yet. Be the first to comment.