Detections
Explore public detection logic contributed by the community across SIEM and rule languages.
4 detections
Filters
Last updated
All Time
Detection languages
2
1
1
Contributors
4
Categories
4
2
Platforms
4
Products / Services
4
MITRE Techniques
18,030
15,416
12,645
8,188
6,021
This rule detects attempts to bypass LLM safety guardrails using adversarial prompting techniques (e.g., 'DAN' persona, hypothetical scenarios) or by identifying patterns of repeated moderation flags and refusals. It specifically looks for user prompts containing known jailbreak templates or an excessive frequency of safety-related trigger events, while excluding authorized red-team and safety evaluation accounts.
Detects potential indirect prompt injection attacks against LLM agents. The rule monitors for ingested or retrieved content containing known injection patterns (e.g., hidden comments, 'system:' role spoofing) that is followed within 5 minutes by a privileged tool invocation by the agent, such as data exfiltration or system command execution.
Detects potential indirect prompt injection attacks against a Large Language Model (LLM) where untrusted context (e.g., RAG-retrieved documents, web content, or tool outputs) containing instruction-override or persona-switch keywords is ingested. This is followed by the LLM performing an anomalous privileged tool call within the same session, indicating a possible compromise of agent autonomy or unauthorized action execution.
Detects potential LLM jailbreak attempts by identifying a high frequency of guardrail-blocked responses followed by user prompts containing known jailbreak strategy fingerprints (such as roleplay, persona switching, or instruction overrides) within a 15-minute window.
