Detections

Explore public detection logic contributed by the community across SIEM and rule languages.

60,142 detections

Detects malicious instruction-override payloads, hidden HTML comments, or obscured text (e.g., zero-font size) in third-party content ingested into LLM/RAG systems, commonly used for indirect prompt injection.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects anomalous, high-volume querying against machine learning inference endpoints. The rule identifies potential model extraction or data exfiltration attempts by flagging high counts of requests from a single user/API key, particularly during off-hours or with high variance in unique inputs, which may indicate automated efforts to reconstruct proprietary models or extract sensitive training data.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects the installation of machine learning or AI software dependencies (pip, conda, npm, poetry) that either pull from unauthorized registries or match known malicious/typosquatting naming patterns (e.g., LiteLLM supply-chain compromise). The rule also monitors for newly published packages that have been flagged as suspicious.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects potential ML model-stealing attacks against inference API endpoints. The rule identifies anomalous behavior characterized by an extremely high volume of diverse queries from a single API key over a 24-hour period. Such patterns are consistent with systematic decision-boundary probing or input-space coverage attempts used for training a replica model (distillation or knockoff attacks) without direct access to model weights.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects the download of a dataset from an unapproved source that lacks hash verification, followed shortly by the launch of a model training process on the same host. This behavior is indicative of potential supply chain compromise or poisoning of the training data used in machine learning workflows.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects potential data leakage from an LLM application where a user repeatedly prompts the model to divulge its internal system instructions, configuration, or sensitive information (e.g., API keys, PII) and the model subsequently outputs that sensitive data.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects attempts to perform context or memory poisoning against an AI agent by injecting imperative, instruction-like payloads into the agent's long-term memory or context store. The rule specifically monitors updates originating from untrusted channels, such as tool outputs, external documents, or web content, which contain patterns commonly associated with prompt injection attacks (e.g., overriding system instructions or role definitions).
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects an AI agent session that performs a data exfiltration action (e.g., sending an email, webhook, or file share) to an unauthorized or non-allowlisted destination shortly after the agent has ingested potentially untrusted external content. This behavior is indicative of an AI agent being manipulated to exfiltrate data after processing malicious or adversarial inputs.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects the ingestion of sensitive credentials (API keys, passwords, private keys) into a RAG (Retrieval-Augmented Generation) document store, followed by the retrieval of those credentials by an application agent, indicating potential credential harvesting within the LLM application pipeline.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects unauthorized modifications to AI Agent Model Context Protocol (MCP) tool manifests. This includes the injection of imperative instructions into tool descriptions, the addition of suspicious instruction-related fields to tool output schemas, or manifest updates initiated by unverified publishers.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects the installation of Model Context Protocol (MCP) servers or AI agent plugins that exhibit characteristics of supply chain risk, such as unverified or unknown publishers, requests for highly sensitive permissions (filesystem write, credential read, unrestricted network egress), or recent public listing, which aligns with AI model supply chain poisoning techniques (AML.T0104).
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects systematic and high-frequency probing of internet-exposed LLM chatbot or API gateway endpoints. The rule monitors for sequences of requests, use of known prompt-infiltration or jailbreak testing toolkits, and varied path access indicative of automated reconnaissance or adversarial prompt engineering attempts.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects attempts by AI agent code-interpreter or sandbox processes to break out of their containerized environment. The rule monitors for the execution of common container escape techniques, such as namespace manipulation (nsenter, unshare), mounting host filesystems, accessing sensitive host devices like /dev/kmsg or /dev/mem, or referencing host-level paths from within an isolated AI agent execution process.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects suspicious AI inference API activity indicative of credential compromise or misuse. The rule identifies three main signals: rapid authentication failures (credential stuffing), anomalous geolocation/ASN usage for successful requests, and successful API calls occurring after a key rotation event.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects high-volume API requests to an AI inference service that explicitly request gradient, confidence vector, or raw logit information. Such patterns are indicative of model inversion or training data reconstruction attacks, where an adversary attempts to extract private training data by exploiting detailed output from the model.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects attempts by users to extract the underlying system prompt from an LLM via prompt injection techniques or identifies successful leaks where the model's response matches a pre-calculated system prompt hash. Exposing system prompts can lead to the discovery of proprietary business logic, safety guidelines, and facilitate further adversarial jailbreaking.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects instances where AI agents are activated through external communication channels like webhooks, email, or chat, and the provided payload includes potentially malicious imperative instructions (e.g., 'delete', 'execute', 'send credentials'). The detection logic excludes activity originating from trusted, allow-listed internal sources to reduce false positives.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects adversarial attempts to infer whether specific records were used in an AI model's training set by analyzing raw prediction confidence scores and identifying suspicious query patterns such as near-duplicate submissions or systematic sweeps.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
This rule detects attempts to bypass LLM safety guardrails using adversarial prompting techniques (e.g., 'DAN' persona, hypothetical scenarios) or by identifying patterns of repeated moderation flags and refusals. It specifically looks for user prompts containing known jailbreak templates or an excessive frequency of safety-related trigger events, while excluding authorized red-team and safety evaluation accounts.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects when a victim user invokes an AI agent tool that has been potentially poisoned. This is identified by either a change in the tool definition hash sourced from unverified or untrusted locations, or by an invocation that results in unexpected, undocumented, or unauthorized secondary tool calls.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000
Detects evidence of indirect prompt injection attacks where an adversary introduces malicious instructions into data sources retrieved by an LLM or agent. The rule identifies common obfuscation and override techniques, including the use of zero-width characters, HTML comments for prompt breaking, base64 encoded instruction verbs, and specific override phrases within retrieved documents or tool outputs.
avatar
Ibrahim Saud@tektrix
avatar
Detections.ai Community
6 days ago
000