Skip to main content

Quantifying Agentic Misalignment in Multi-Agent Systems

As Large Language Model (LLM) agents gain enterprise tool-execution permissions, agentic misalignment becomes a measurable operational risk: autonomous systems may pursue sub-goals that violate implicit safety policies even when tools work as designed (Lotfi et al., 2026; Zhuang & Hadfield-Menell, 2020; Zscaler, 2026).

This study tests whether supervisor/sub-agent delegation creates a responsibility gap that increases unsafe behavior. Across 1,500 runs covering six models, eight adversarial scenarios, and five comparison groups, the results reject that hypothesis: delegated topologies scored 0% misaligned across all models, while failures concentrated in single-agent configurations, especially under prompt injection. Runtime telemetry captured tool calls and internal-monologue steps, but scoring reflects completed-action ground truth rather than validated real-time intervention (see Section 3).

These findings support behavior-first evaluation and show why runtime observability, tool-boundary controls, and scenario-specific validation are necessary complements to prompt-level safety instructions (Lotfi et al., 2026; Tarasenko & Voruganti, 2026).

SANS-Quantifying-Agentic-Misalignment-Multi-Agent-Systems-100126 (PDF, 1.18MB)

1 Oct 2026
ByVasu Koduru
Share
All papers are copyrighted

No re-posting of papers is permitted

Related Content

Invisible by Default: The Polyglot Smuggler

Research Paper

This paper demonstrates a novel data smuggling mechanism: Sub-Channel Steganography. By weaponizing uninspected MP4 subtitle tracks within a structurally compliant ISO base media file, adversaries can encapsulate malicious payloads and exfiltrate sensitive data.

  • 1 Oct 2026
  • John Porpora

Network Artifacts of Trusted Service Abuse

Research Paper

This study contrasts normal user traffic with known malicious activity to identify alternative indicators from four public C2 frameworks (DaaC2, DBC2, gdog, Callidus) operating on Dropbox, Discord, Gmail, and OneNote, using 51 packet captures (3 baselines, 6 automated browser tests and 6 framework tests per service).

  • 6 Aug 2026
  • Nicholas Vaniscak

Dual-Module QR Codes: Bypassing QR Code Scanners in Enterprise Email Gateways

Research Paper

By exploiting the decoding algorithm and leveraging existing QR data-layering techniques, this research provides the security industry with a working proof-of-concept to bypass email security gateway detection systems.

  • 6 Aug 2026
  • Steven Legere

Post-Exploitation: C2 Framework Effectiveness Against Advanced Audit Logging

Research Paper

This research paper examines the effectiveness of a sample of open-source Commandand-Control (C2) frameworks in evading advanced audit logging during postexploitation.

  • 20 Mar 2026
  • Benjamin Evans

Enhancing Security Operations with Google Threat Intelligence

Research Paper

This product review examines how Google Threat Intelligence's extensive data sources, real-time insights, and investigative capabilities can elevate SecOps workflows and strengthen an organization’s defensive posture.

  • 24 Nov 2025
  • Dave Shackleford

Interrogators: Attack Surface Mapping in an Agentic World

Research Paper

This research introduces the concept of AI agent interrogators and the open-source project Agent Interrogator, an opaque box interrogation framework designed to map the attack surface of agentic systems.

  • 23 Oct 2025
  • Michael Samson

The Mimic Octopus: Weaponizing File Corruption and Recoverability to Bypass Antivirus and Email Filtering

Research Paper

This paper investigates a novel tactic in phishing operations where threat actors intentionally corrupt document and archive files, such as DOCX, DOCM, PDF, and ZIP , to evade antivirus (AV) and email filtering systems.

  • 3 Sep 2025
  • Justin Gazick

From Crash to Compromise: Unlocking the Potential of Windows Crash Dumps in Offensive Security

Research Paper

This research explores how offensive security practitioners can incorporate crash dump analysis into their workflows to extract sensitive data such as plaintext credentials, encryption keys, and files from memory.

  • 9 May 2025
  • SANS Institute

CloudFront Real-Time Logs Rate Sampling and Detection

Research Paper

As businesses aim to optimize their AWS CloudFront expenses, some disable CloudFront Real-Time logs....

  • 29 Jan 2024
  • Merly Mathis

The Evolution of the Digital Predator: Using AI to Evade Security Controls

Research Paper

Since the advent of the computer, there has been a never-ending game of cat and mouse between those...

  • 20 Dec 2023
  • Foster Nethercott

Who Needs a Pentest: Validating the Configuration of an EDR Solution Using the MITRE ATT&CK Framework

Research Paper

Is that EDR suite fully configured, and providing the expected protection? Do we have a scalable way...

  • 7 Nov 2023
  • Adam Fowler

Tearing up Smart Contract Botnets

Research Paper

The distributed resiliency of smart contracts on private blockchains is enticing to bot herders as a...

  • 22 Oct 2018
  • Jonathan Sweeny

Clickbait: Owning SSL via Heartbleed, POODLE, and Superfish

Research Paper

In the twilight of SSL's effectiveness as a method of secure communication,demonstration of...

  • 23 Dec 2015
  • SANS Institute