Er du sikker? How Norwegian pre- and post training of LLMs affect safety guardrails
Interested in AI safety and security, large language models (LLMs) and how well their internal guardrails actually hold up outside English?
Most LLM guardrails are built and evaluated in English, yet models are increasingly trained and fine-tuned for other languages, such as Norwegian. This project investigates Norwegian-focused LLMs, and asks the question: when a model is adapted for Norwegian, what happens to its internal guardrails?
Description
Safety and security guardrails in LLMs are the mechanisms that make a model refuse harmful or manipulative requests. These are typically trained through alignment techniques, such as instruction tuning, and are often developed and evaluated in English. Recent research shows that such safety mechanisms do not necessarily transfer reliably across languages: a model that refuses a harmful request in English may comply with the same request in a lower-resourced language. At the same time, multilingual models are increasingly being adapted to Norwegian through continual pre-training and various stages of post-training, and are released and deployed.
Norway has an active open-model ecosystem, but Norwegian remains a comparatively low-resourced language, potentially placing these models in the risk zone described above. Few studies have examined how Norwegian adaptation affects model safety. This project investigates this gap and aims to understand the safety implications of adapting LLMs to Norwegian, thus contributing to a safer deployment of AI systems in Norwegian.
Goals
The project can be scoped and adjusted to a subset of the following, depending on the student's interests:
- Compare Norwegian models at different stages of training against their original base models to isolate the effect of Norwegian adaptation on safety, i.e. measure guardrail drift.
- Measure internal guardrail behaviour (refusal rate, compliance and degree of harmful disclosure) for a selection of open models, comparing Norwegian and English prompts.
- Correlate the size of a model's Norwegian pre-training data with its measured guardrail robustness, testing whether more Norwegian data equals safer in Norwegian.
- Investigate the differences between bokmål and nynorsk: does the safety gap differ between the two written standards?
Relevant models to investigate can be (but are not limited to): NorMistral (https://www.mn.uio.no/ifi/english/research/groups/ltg/llms-for-norwegian/) and Borealis (https://ai.nb.no/borealis/) models.
Learning outcome
- A solid understanding of LLM safety and security, how guardrails are trained into models and how they can fail.
- Strong familiarity with the Norwegian language model landscape.
- Practical experience on the challenges of deploying LLMs safely in a lower-resourced language.
- Reflecting on important ethical considerations in security and language technology research.
Qualifications
- Motivation and curiosity about AI safety, security and language models.
- Fluency in Norwegian, as the project involves authoring and evaluating Norwegian-language test material.
- Proficiency in Python.
- Enrolled as a master student in Language Technology (UiO).
Supervisors
- Lilja Øvrelid, UiO (main supervisor)
- Annika Willoch Olstad
- Michael Riegler
Collaboration partners
- Language Technology Group at University of Oslo
References
- Relevant reports from the Centre of AI Security and Safety: https://www.simula.no/research/research-centres/caiss-centre-ai-security-and-safety/caiss-reports
- "NLP Security and Ethics, in the wild" (Lent et al., 2025)
- "Multilingual Jailbreak Challenges in Large Language Model" (Deng et al., 2024)
- NorMistral from the Language Technology Group at IFI: https://www.mn.uio.no/ifi/english/research/groups/ltg/llms-for-norwegian/
- Borealis from NB AI lab: https://ai.nb.no/borealis/