D07 — Tactical Operationalization and AI-Enabled Planning
Safety Filter Bypassing or Jailbreaking
Description
The person attempts to bypass AI safety systems to obtain prohibited or harmful information.
Rationale
Attempted bypass of AI safety systems is the clearest behavioral evidence that the person is actively seeking harmful content that the AI would not otherwise provide. This behavior requires intent — the person knows the information is restricted and is deliberately trying to circumvent that restriction. It is a significant escalation indicator regardless of whether the attempt succeeds.
Evidence Base
Guardrails Under Test: Terrorist Misuse of AI Models and the Open-Weight ‘Abliteration’ Problem
Hadley, 2026. Guardrails Under Test: Terrorist Misuse of AI Models and the Open-Weight ‘Abliteration’ Problem. CTC Sentinel, 1(1), pages 1-13. DOI not present.
Peer-reviewed study (original research)“God has helped us, and so will AI”: How the Terrorist Group Boko Haram Uses Frontier AI
Antonia Juelich, 2026. “God has helped us, and so will AI”: How the Terrorist Group Boko Haram Uses Frontier AI. Frontier AI Working Paper Series, No. 1/2026.
AI company reportA scoping review on the mental health harms of LLM-based chatbots
Diel A, Torous J, Cuijpers P, Kleesiek J, Nensa F, Weber N, Faust F, Lalgi TJ, Mellis FS, Teufel M, Bäuerle A, 2024. A scoping review on the mental health harms of LLM-based chatbots. npj Digital Medicine, 9(644). https://doi.org/10.1038/s41746-026-03054-x
Peer-reviewed study (original research)Intelligent Systems, Vulnerable Minds: A Framework for Radicalization to Violence in the Age of AI
Kunst et al., 2026. Intelligent Systems, Vulnerable Minds: A Framework for Radicalization to Violence in the Age of AI. Personality and Social Psychology Review, 30(3), 395-426. https://doi.org/10.1177/10888683261430089
Peer-reviewed study (original research)