Verified from career page · Posted yesterday
Master's Thesis: Human-Centred Evaluation of Explainable Reinforcement Learning New
Ericsson · 75 Technology & Research
Stockholm
- Posted
- yesterday
- Workplace
- On-site
- Salary
- Not disclosed
- Visa sponsorship
- Not specified
Posted on 28 September 2026
Work model: On-site
Salary range not shared by the company
Visa sponsorship details unknown
Join our Team
About this opportunity:
Reinforcement Learning (RL) is increasingly used for sequential decision-making in areas such as robotics, recommender systems, autonomous control, and telecommunications. In telecom, RL is being explored for radio resource allocation, traffic steering, energy saving, and network self-optimisation. However, RL policies are typically opaque, making it difficult for operators, engineers, and researchers to understand why an agent selected a particular action.
This is a serious obstacle to deployment in telecom, where trust, accountability, and the ability to diagnose misbehaviour are essential. Explainable Reinforcement Learning (XRL), a sub-field of Explainable AI (XAI), aims to make agent behaviour interpretable to humans.
This thesis will design and run a user study comparing Feature Importance (FI) explanations with Temporal Policy Decomposition (TPD), which explains actions through predicted future outcomes. The study will investigate whether outcome-based explanations are more useful to humans than feature-attribution explanations in an RL context.
The work corresponds to two students, 30 hp each, and can be organised into two subtracks. The students will collaborate on the user-study infrastructure and codebase. The location is Stockholm, Kista, and the preferred starting period is October 2026 to January 2027.
What you will do:
Review XAI and XRL literature and identify appropriate metrics and evaluation protocols for explanation quality.
Design a user-study protocol based on four conditions:
No explanation: participants see only the agent's actions.
FI only: participants see feature-importance explanations.
TPD only: participants see temporal-outcome explanations.
TPD + FI: participants see both explanation types.
Extend an existing web application to support the required XRL methods, the combined condition, and new measurements.
Run a pilot study, refine the protocol, recruit participants, and conduct the main user study.
Analyse the results statistically and evaluate the effectiveness of each explanation method individually and in combination.
Write the thesis report and present the results to the research team.
The skills you bring:
You are a Master's student in Computer Science, Human–Machine Interaction, Machine Learning, Data Science, or a related field.
You have a foundation in machine learning, basic statistics, and data analysis.
You have good programming skills in JavaScript and Python.
You have good English proficiency and can communicate your findings clearly.
About Ericsson
Ericsson's European engineering spreads across Budapest, Stockholm, Kraków and Reading. Telecom infrastructure at global scale — Stockholm carries the founding engineering culture, Budapest and Kraków the larger current headcount.