Secure and Trustworthy Intelligent Systems
The Secure and Trustworthy Intelligent Systems (SENTRY) department studies the interplay between artificial intelligence and software & system engineering, with a focus on security. We investigate methods and techniques that provide evidence-based insights to support the development, assurance, and evolution of secure and trustworthy intelligent systems.

Software systems are increasingly built by combining traditional software systems with AI components. This combination is a source of great power, but it also leads to uncertainty, limited interpretability, and behavior that may shift over time. These systems are less predictable, harder to analyze, and vulnerable to new classes of failures and attacks. Traditional assurance techniques cannot offer credible guarantees under these new conditions.
We work in two directions:
- We use AI to make software systems more secure and resilient: detecting and repairing vulnerabilities, analyzing cyber threat intelligence, and creating autonomously self-healing systems.
- We apply software and security engineering to AI: red-teaming large language models, auditing and hardening AI components, and monitoring and controlling AI agents at runtime.
Our methods draw on automated program repair, repository and log mining, (statistical) program analysis, fuzzing, information-flow analysis, and empirical evaluation. We collaborate closely with industry to ensure that our research addresses real-world challenges and that new solutions are tested in realistic settings.
As AI becomes part of critical software, we investigate how it can be attacked, and how it fails, and we develop the techniques to keep such systems secure, trustworthy and under control.
Leon Moonen, head of the SENTRY department
SENTRY leads the AI Security research in the Centre for AI Security and Safety (CAISS), which evaluates the security and safety of AI systems as they become part of national infrastructure and advises Norwegian authorities on AI risk in critical systems. Our work in CAISS targets intentional harm and adversarial threats in AI systems.
Focus areas
AI Security
AI models are increasingly used as components in software that handles sensitive data and makes decisions that have societal impact. We study how these components can be attacked, misled, or exploited, for example through prompt injection, jailbreaks, data poisoning, and backdoors. Moreover, we develop frameworks and benchmarks for red-teaming foundation models and the systems built on them, aiming to evaluate their behavior under simulated but realistic adversarial attacks.
We adapt program-analysis techniques such as fuzzing and information-flow analysis to audit and harden AI components, and assess the risks of AI in high-stakes and regulated settings. Through CAISS, this work includes security auditing of language models used in Norway.
Monitoring and Control of Agentic AI
AI agents that plan, call tools, and act on their environment raise a new control problem: their behavior emerges at runtime and can change as they interact with other agents, the environment, or their human operators. We develop runtime monitoring, guardrails, and containment strategies for agentic systems, and study how to check that an agent stays within the limits it was given, including when the monitor is itself a learned component that can be targeted by an attacker. We also build agents ourselves, such as multi-agent systems that write, test, and repair code, which gives us first-hand insight into how they fail.
Cyber Security
Security professionals spend many hours inspecting code to find and repair vulnerabilities. We use large language models and other machine learning techniques to detect and automatically repair security vulnerabilities in source code, and to predict the impact of newly disclosed vulnerabilities.
A second line of work builds cyber threat intelligence knowledge graphs that connect software versions, vulnerabilities, threats, exploits, and incidents, including mappings of known exploited vulnerabilities to attacker techniques, to support proactive risk assessment.
Software Resilience
However thorough testing procedures are used, it is impossible to anticipate all possible failures in complex, highly interconnected software systems (and even more so if they contain probabilistic AI components). We investigate the use of adaptive, bio-inspired approaches to create autonomously self-healing systems. These are self-monitoring systems that can detect when they are not operating correctly and, without human intervention, make the necessary adjustments to restore normal operation. Specifically, we investigate the use of artificial immune systems and reinforcement learning on operational data to learn how to automatically improve a system’s resilience. We also analyze large-scale logging data, using language models for log parsing, to detect anomalies and support incident response.




