Modoante
SwedenFull TimeSoftware Developers

Master Thesis: Cross-Lingual Safety Alignment & System Identity Injection via DPO

AI Sweden | Stockholm, Stockholm County, Sweden | Salary not specified

Source: JobsPipe

Required Skills

pythonlinuxdata-science
View company

Role snapshot

Work model
Hybrid
Language
Not specified
Experience
Not specified
Posted
Posted 1 day ago
Application deadline
2026-10-17

What you'll do

As Sweden's national center for applied AI, we're on a mission to accelerate the use of AI to benefit our society, our competitiveness, and everyone living in Sweden. We drive impactful initiatives in areas such as healthcare, energy, and public services while pushing the boundaries of AI research in fields such as natural language processing, machine learning and AI security. Join us in harnessing the untapped value of AI to drive innovation and create sustainable value for Sweden. We are now looking for a master thesis student to join our team. Introduction Ensuring that an open-weight European model maintains a consistent system identity ("OpenEuroLLM") Models often over-refuse or exhibit language-dependent safety leaks when prompted in low-resource European languages. Project Background and Problem Statement T4.6 has established the openeurollm-model-identity dataset and a multilingual safety refusal suite. This thesis asks:

  • Can joint DPO co-training on openeurollm-model-identity and Safe-RLHF pairs eliminate cross-lingual safety leaks and over-refusal without degrading instruction-following performance in Swedish and Icelandic?

Outline

  • Literature Study: Review cross-lingual safety gaps, EU AI Act risk taxonomies, and system prompt adherence in DPO.
  • Implementation: Fine-tune Prelude 9B using openeurollm-model-identity and target-language Safe-RLHF preference pairs using Hugging Face TRL or Megatron-LM.
  • Evaluation: Conduct cross-lingual jailbreak probes (e.g., Do-Not-Answer, X-Alpaca) and evaluate over-refusal rates on benign administrative queries.

Who We’re Looking For We are seeking curious, self-driven MSc students eager to work at the frontier of open-weight European AI research (LLMs). You thrive on empirical discovery, design rigorous experiments, and let data challenge your assumptions.

  • Ongoing Master’s studies in Computer Science, Data Science, Machine Learning, Engineering Physics, or a related quantitative field.
  • Proficiency in Python and hands-on experience with modern deep learning frameworks (PyTorch, Hugging Face ecosystem).
  • Familiarity with LLM post-training alignment (e.g., SFT, DPO, RLHF/RLVR) or context-extension, alongside comfort running distributed GPU training in Linux/HPC environments.

At AI Sweden, we are committed to building diverse and inclusive teams. Some positions may be subject to export control regulations, which means that specific requirements may apply. Why should you do your thesis with AI Sweden? Doing your thesis at AI Sweden means working alongside leading AI scientists and change leaders. AI Sweden is Sweden’s National Center for AI, we drive research questions that have both a long shelf-life and are widely applicable to Swedish industry and the public sector. We aim for publications at the most competitive venues and celebrate a culture of research excellence. As an organization, we’re uniquely positioned at the sweet spot of governmental influence and startup agility. Small enough to stay adaptive and have fun but backed by and in close contact with both the government, academia and private and public sector. Practical details Location: Hybrid (Gothenburg / Stockholm) or Remote. Application Deadline: 2026-11-20 (rolling selection – position may be filled earlier). Start Date: January 2027 Contact If you have any questions or thoughts, don’t hesitate to contact: Amaru Cuba Gyllensten, Senior Research Scientist Danila Petrelli, Senior Data Lead/Research Scientist AI Sweden does not accept unsolicited support and kindly ask not to be contacted by any advertisement agents, recruitment agencies or manning companies. References [1] OpenEuroLLM Consortium, "openeurollm-model-identity & Safety Evaluation Suite," 2026. [2] PKU Alignment Team, "Safe-RLHF: Safe Reinforcement Learning from Human Feedback," 2024

Work resources

Helpful work resources for this job

Sign in to read

Application checklist

Review your application before submitting

A quick checklist to help candidates submit a clearer, more complete application.

Read resource