AI Agent Bypassing Guardrail/Moderation Checks (T1562)
Detects unauthorized attempts to disable, reconfigure, or bypass AI safety guardrails and moderation services, particularly when followed by high-risk tool invocations such as file deletion, network egress, or credential access. This activity indicates an attacker attempting to impair an AI agent's defensive mechanisms to facilitate malicious actions.
Sigma

