Modoante
SwedenFull TimeOther

Master's Thesis: Human-Centred Evaluation of Explainable Reinforcement Learning

Ericsson | Stockholm, Stockholm County, Sweden | Salary not specified

Source: JobsPipe

Required Skills

javascriptpythondata-sciencedata analytics
View company

Role snapshot

Work model
On-site
Language
Not specified
Experience
Not specified
Posted
Posted 1 day ago
Application deadline
2026-10-13

What you'll do

Join our Team About this opportunity: Reinforcement Learning (RL) is increasingly used for sequential decision-making in areas such as robotics, recommender systems, autonomous control, and telecommunications. In telecom, RL is being explored for radio resource allocation, traffic steering, energy saving, and network self-optimisation. However, RL policies are typically opaque, making it difficult for operators, engineers, and researchers to understand why an agent selected a particular action. This is a serious obstacle to deployment in telecom, where trust, accountability, and the ability to diagnose misbehaviour are essential. Explainable Reinforcement Learning (XRL), a sub-field of Explainable AI (XAI), aims to make agent behaviour interpretable to humans. This thesis will design and run a user study comparing Feature Importance (FI) explanations with Temporal Policy Decomposition (TPD), which explains actions through predicted future outcomes. The study will investigate whether outcome-based explanations are more useful to humans than feature-attribution explanations in an RL context. The work corresponds to two students, 30 hp each, and can be organised into two subtracks. The students will collaborate on the user-study infrastructure and codebase. The location is Stockholm, Kista, and the preferred starting period is October 2026 to January 2027. What you will do:

  • Review XAI and XRL literature and identify appropriate metrics and evaluation protocols for explanation quality.
  • Design a user-study protocol based on four conditions:
  • No explanation: participants see only the agent's actions.
  • FI only: participants see feature-importance explanations.
  • TPD only: participants see temporal-outcome explanations.
  • TPD + FI: participants see both explanation types.
  • Extend an existing web application to support the required XRL methods, the combined condition, and new measurements.
  • Run a pilot study, refine the protocol, recruit participants, and conduct the main user study.
  • Analyse the results statistically and evaluate the effectiveness of each explanation method individually and in combination.
  • Write the thesis report and present the results to the research team.

The skills you bring:

  • You are a Master's student in Computer Science, Human–Machine Interaction, Machine Learning, Data Science, or a related field.
  • You have a foundation in machine learning, basic statistics, and data analysis.
  • You have good programming skills in JavaScript and Python.
  • You have good English proficiency and can communicate your findings clearly.

Work resources

Helpful work resources for this job

Sign in to read

Application checklist

Review your application before submitting

A quick checklist to help candidates submit a clearer, more complete application.

Read resource