Detections

Explore public detection logic contributed by the community across SIEM and rule languages.

4 detections

This rule detects attempts to bypass LLM safety guardrails using adversarial prompting techniques (e.g., 'DAN' persona, hypothetical scenarios) or by identifying patterns of repeated moderation flags and refusals. It specifically looks for user prompts containing known jailbreak templates or an excessive frequency of safety-related trigger events, while excluding authorized red-team and safety evaluation accounts.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects potential indirect prompt injection attacks against LLM agents. The rule monitors for ingested or retrieved content containing known injection patterns (e.g., hidden comments, 'system:' role spoofing) that is followed within 5 minutes by a privileged tool invocation by the agent, such as data exfiltration or system command execution.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
10 days ago
000
Detects potential indirect prompt injection attacks against a Large Language Model (LLM) where untrusted context (e.g., RAG-retrieved documents, web content, or tool outputs) containing instruction-override or persona-switch keywords is ingested. This is followed by the LLM performing an anomalous privileged tool call within the same session, indicating a possible compromise of agent autonomy or unauthorized action execution.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
10 days ago
000
Detects potential LLM jailbreak attempts by identifying a high frequency of guardrail-blocked responses followed by user prompts containing known jailbreak strategy fingerprints (such as roleplay, persona switching, or instruction overrides) within a 15-minute window.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
10 days ago
000