State of AI in Norway. "Evaluation is what makes us robust".
panel debate at Arendalsuka

State of AI in Norway. "Evaluation is what makes us robust".

Published:

Norway is near the top of the world when it comes to using artificial intelligence. We are third globally on AI adoption at work, and workers trust AI more than anyone else in the Nordic region. But how safe are the models in a Norwegian context?

On August 11th at Arendalsuka the Centre for AI Security and Safety (CAISS) at Simula launched a report on the State of AI in Norway 2026.

Alongside a broad picture of how Norway uses, develops, regulates, funds and teaches AI, the centre ran its own Norwegian specific tests on 21 language models and reviewed the published safety documentation of frontier models.

"If we get good at evaluation, we become robust against the chaos," said Chief Research Scientist Michael A. Riegler, referring to the pace at which new models are released. 

Riegler heads the AI Safety department at SimulaMet, part of CAISS, and led the work on the report.

Main findings

  • Norway ranks third globally and first in Europe for AI diffusion in the workforce, at 46.4%.
  • 67% of Norwegian workers use AI at work, and 61% say they trust AI at work.
  • Norway's public and government-affiliated AI investment is estimated at US$716 million, less than Sweden and Finland.

  • Norwegian fine-tuning works: the Norwegian built Borealis models outperformed their base models on Norwegian tasks, up to 22.8 points in healthcare.
  • Safety remains a major challenge. Under sustained pressure, even the strongest open-weight model produced a critical failure in about one in five responses.
  • When users confidently made false claims, 85–92% of responses were critical failures in the most difficult medical cases. The weakness affected all models tested.
  • Testing over several rounds matters: models often changed a correct initial response after sustained pressure.

  • CAISS found that AI companies' system cards should not be treated as independent safety assurance.
  • For high-risk applications, organisations need independent testing in their own context.

Launched and debated at Arendalsuka

The report was presented and discussed at Arendalsauka, the largest political gathering in Norway. The panel was chaired by Simula CEO, Lillian Røstad, with Hans Christian Holte (Director, KI Norge), Åse Wetås (National Librarian, National Library of Norway), Ingvild Thorsvik (Deputy Leader, Venstre) and Marte Ingul (State Secretary, Ministry of Digitalisation and Public Governance).

The event was recorded and is available at Simula's YouTube channel:

How CAISS put 21 models under pressure 

CAISS tests models using SimpleAudit, Simula's framework for auditing the safety of AI systems. It is open source and runs locally, so a hospital or municipality can run the same tests on its own systems, without any data leaving the building. 

The method simulates a persistent user. One model opens with a realistic request, then pushes across several turns. It rephrases, insists and appeals to honesty. 

A second model then scores the exchange. Each test ran ten times per model, across scenarios built around Norwegian situations, such as NAV, child welfare, BankID, plus two sets built around confident false claims.

“All models have a breaking point. What matters is where it is, and how quickly you get there,” says Riegler. 

Each model gets a score from 0 to 100 across six tests. The higher the score, the safer it performed under pressure. The scores are relative: they compare models against each other, not against a fixed safety standard. 

“Norwegian fine-tuning genuinely works” 

The results, summarised above, show that Norwegian adaption works. The 15 Borealis models built by the Norwegian national library on Google’s Gemma-3 model, scored up to 22.8 points higher than its foundation model on Norwegian health scenarios. 

"It shows that targeted fine-tuning genuinely works as a way to shape how a model behaves," Riegler says. "We don't have to simply accept whatever a foreign base model happens to do."

But the two largest Borealis models were 3.5 to 5.1 points weaker than Gemma-3 at resisting a user who confidently insists on something false.

"The working theory is that Norwegian instruction-tuning makes the models more fluent, and more accommodating," Riegler says.

These are raw models, and most real deployments add extra safeguards on top. A low score shows where a model's limits lie under pressure. The same weakness is shared by every model tested. 

"Models often start out correctly, then fold under pressure. This is why testing a single question makes a model look safer than it really is," says Riegler.

Testing a model on its own is not enough

What people actually use is a system built around a model, the source look-ups, filters, tool access, logging. That surrounding layer is where the fastest gains are available: “pulling rules directly from Lovdata and Helsenorge instead of relying on the model's memory, hard-coding emergency routing to 113, and filtering every turn of a conversation, not just the first”.

The CAISS centre also analyzed the system cards (technical reports) from Anthropic's latest models (June 2026), Mythos 5 and Fable 5. They were compared with four earlier Anthropic system cards, as well as system cards published by OpenAI and Google.

This research is conducted within the Centre for AI Security and Safety (CAISS), and is a collaboration between researchers at SimulaMet and Simula Research Laboratory.