Master's Thesis: Explainability in Multi-Objective Reinforcement Learning Agents
Join our Team
About this opportunity:
Reinforcement Learning (RL) has shown strong performance in fields such as robotics and the alignment of large language models. An RL agent optimises its policy by maximising a cumulative reward, typically represented as a scalar value quantifying the overall goal.
Real-world problems are often more complex and involve multiple competing objectives. Multi-Objective Reinforcement Learning (MORL) addresses this by learning policies that map to the Pareto-optimal frontier while preserving trade-offs between objectives. For example, a legged robot may need to maximise speed while minimising energy consumption, while an antenna-tilt optimisation agent may need to minimise signal interference while maximising coverage area.
State-of-the-art methods such as Envelope Q-Learning and concave-augmented Pareto Q-learning enable humans to adjust priorities among different objectives at runtime. However, the decision-making process of these agents remains largely a black box.
The goal of this thesis is not only to implement and train high-performing MORL agents, but also to systematically extract the reasoning behind their decisions using causality or explainable AI (XAI) methods. The work corresponds to 30 hp for one student, with a preferred starting date in December 2026 or January 2027. The location is Kista.
What you will do:
- Conduct a literature review and reproduce state-of-the-art MORL algorithms.
- Design and train MORL agents in an open-source environment such as MO-Gymnasium and, optionally, in an internal telecommunications simulator.
- Generate explanations for the trained MORL agents using one of the following approaches:
- Post-hoc explainability: implement a technique for generating explanations from a trained MORL agent.
- Causality: train a causality-based agent or distil the trained policy into a causal model.
- Analyse the results using available benchmarks, such as the Open RL Benchmark.
- Document the findings in an academic thesis report, present the results, and potentially contribute to a research paper.
The skills you bring:
- You are a final-year Master's student in Mathematical Statistics, Applied Mathematics, Control Engineering, Electrical Engineering, Robotics, Computer Engineering, Computer Science, Machine Learning, or a related field.
- You have good theoretical knowledge of reinforcement learning.
- You have excellent programming skills in Python and PyTorch.
- You can work independently, structure an open research question, and communicate your findings clearly.
Experience with reinforcement-learning libraries such as Gymnasium, Ray RLlib, Stable-Baselines3, or PyTorch is considered a plus. Relevant coursework or project experience in reinforcement learning, deep learning, or similar areas is also beneficial.
Keywords: multi-objective reinforcement learning, explainable AI, causality, and model distillation.
Selected references: R. Yang, X. Sun, and K. Narasimhan, “A generalized algorithm for multi-objective reinforcement learning and policy adaptation,” NeurIPS 2019; H. Lu, D. Herman, and Y. Yu, “Multi-Objective Reinforcement Learning: Convexity, Stationarity and Pareto Optimality,” ICLR 2022.
Why join Ericsson?At Ericsson, you´ll have an outstanding opportunity. The chance to use your skills and imagination to push the boundaries of what´s possible. To build solutions never seen before to some of the world’s toughest problems. You´ll be challenged, but you won’t be alone. You´ll be joining a team of diverse innovators, all driven to go beyond the status quo to craft what comes next.
What happens once you apply?Click Here to find all you need to know about what our typical hiring process looks like.Encouraging a diverse and inclusive organization is core to our values at Ericsson, that's why we champion it in everything we do. We truly believe that by collaborating with people with different experiences we drive innovation, which is essential for our future growth. We encourage people from all backgrounds to apply and realize their full potential as part of our Ericsson team. Ericsson is proud to be an Equal Opportunity Employer. learn more.
Primary country and city: Sweden (SE) || Stockholm
Req ID: 791199

