MITRE ATLAS 2026 Top AI Detection: LLM Jailbreak and Guardrail Bypass Attempt [A

This rule detects attempts to bypass LLM safety guardrails using adversarial prompting techniques (e.g., 'DAN' persona, hypothetical scenarios) or by identifying patterns of repeated moderation flags and refusals. It specifically looks for user prompts containing known jailbreak templates or an excessive frequency of safety-related trigger events, while excluding authorized red-team and safety evaluation accounts.