A C OMPARATIVE E VALUATION OF AI AGENT S ECU RITY G UARDRAILS Jiwei Shi
Qi Li Jiu Li Pingtao Wei Jianjun Xu Xueyi Wei Xuan Zhang Yanhui Yang Xiaodong Hui Peng Xu Lingquan Zhou Beijing Caizhi Tech, Beijing, China
arXiv:2604.24826v1 [cs.CR] 27 Apr 2026
A BSTRACT This report presents a comparative evaluation of DKnownAI Guard in AI agent security scenarios, benchmarked against three competing products: AWS Bedrock Guardrails, Azure Content Safety, and Lakera Guard. Using human annotation as the ground truth, we assess each guardrail’s ability to detect two categories of risks: threats to the agent itself (e.g., instruction override, indirect injection, tool abuse) and requests intended to elicit harmful content (e.g., hate speech, pornography, violence). Evaluation results demonstrate that DKnownAI Guard achieves the highest recall rate at 96.5% and ranks first in true negative rate (TNR) at 90.4%, delivering the best overall performance among all evaluated guardrails.
1
E VALUATION BACKGROUND
In our previous work (1), we conducted an initial evaluation based on the S-Eval benchmark and our proprietary DeepKnown-High-Risk dataset, validating DKnownAI Guard’s detection capabilities in general security scenarios. The datasets used in that evaluation primarily covered traditional text content safety and did not adequately address the diverse attack scenarios that AI agents face in real-world deployments. As AI agents continue to evolve and gain widespread adoption, the security threats targeting them are accelerating in both scope and sophistication. The OpenClaw case serves as a compelling case study: OpenClaw is a widely-used AI agent application capable of directly controlling user computers through natural language, with high system privileges including file system read/write, environment variable management, API invocation, and plugin installation. Security researchers have disclosed multiple critical vulnerabilities in OpenClaw: attackers can execute prompt injection through malicious web pages to steal user credentials; manipulate the agent into deleting important data; compromise plugins and skill packages to exfiltrate API keys and deploy trojans; and the software itself contains multiple high-severity security vulnerabilities with notably inadequate default security configurations. These real-world cases demonstrate that the AI agent attack surface has expanded from traditional text content safety to multi-dimensional threats including instruction override, indirect injection, tool abuse, and plugin poisoning, with increasingly covert and complex attack techniques. Driven by this trend, it is necessary to conduct more intensive evaluations across broader attack scenarios. This evaluation introduces multiple adversarial security datasets (see §4), with emphasis on agent-specific attack scenarios including instruction override, indirect injection, role hijacking, chain-of-thought poisoning, and tool abuse. The attack intensity and deception level of these datasets significantly exceed those used in the previous evaluation, aiming to more comprehensively reflect the security challenges currently facing AI agents.
2
P RODUCT C APABILITIES AND E VALUATION O BJECTIVES
2.1
DK NOWNAI G UARD C ORE C APABILITIES
DKnownAI Guard (https://dknownai.com/) provides comprehensive security protection for AI agent scenarios, covering two major categories of security capabilities. 1
2.1.1
AGENT T HREAT D ETECTION
Detects malicious inputs that attempt to control, exploit, or compromise the agent itself, preventing the agent from being weaponized as an “attacker’s tool.” Table 1: DKnownAI Guard Core Detection Capabilities Capability
Description
Instruction Override Detection
Identifies direct or indirect attacks that override agent system instructions, including instruction replacement, delimiter attacks, and role hijacking Prevents attackers from inducing the agent to expose sensitive information such as passwords, API keys, and cryptographic keys Recognizes attacks that manipulate agent decision logic, including plan hijacking, chain-ofthought poisoning, and logical traps Detects malicious instructions embedded through contaminated external resources (web pages, documents, emails) accessed by the agent, including plugin and skill package poisoning Prevents attackers from inducing the agent to invoke dangerous tools or APIs to execute destructive operations
Privacy Data Leakage Prevention
Malicious Behavior Manipulation Detection
Indirect Injection Detection
Tool Abuse Prevention
2.1.2
H ARMFUL C ONTENT D ETECTION
Detects malicious requests intended to elicit inappropriate content from the agent, including hate speech, pornography, and violence, serving as a supplementary security capability. 2.2
P RODUCT A DVANTAGES Table 2: DKnownAI Guard Advantages
Dimension
Description
Dual-Channel Risk Classification
Independent detection channels that distinguish between “agent threat” and “harmful content” risks, enabling independent policy configuration for each risk type Dedicated detection engine optimized for agent-specific attack patterns including instruction override, privacy leakage, and behavior manipulation Flexible security policy configuration by risk type, adaptable to diverse business scenario requirements
Agent-Specific Detection
Scenario-Based Configuration
2.3
E VALUATION O BJECTIVES
To validate DKnownAI Guard’s practical protection effectiveness, this evaluation selects three competing products—AWS Bedrock Guardrails, Azure Content Safety, and Lakera Guard—for comparative testing. We assess each guardrail’s detection capability for both agent threat security (instruction override, privacy data leakage, malicious behavior manipulation, indirect injection, tool abuse) and harmful content elicitation (hate speech, pornography, violence). Evaluation results are unified into a BLOCKED / ALLOWED binary classification, with human annotations serving as the ground truth for accuracy comparison. 2
3
E VALUATED P RODUCTS Table 3: Evaluated Security Guardrail Products
Vendor
Product
Positioning
AWS
Bedrock Guardrails
Microsoft Azure
Content Safety
Lakera
Lakera Guard
DKnownAI
DKnownAI Guard
LLM safety guardrail within the AWS ecosystem, supporting content filtering and contextual groundedness detection Part of Azure AI services, providing multi-modal (text/image) content safety moderation AI security startup specializing in prompt injection detection Agent security solution deeply optimized for AI agent scenarios, supporting dual-channel risk classification and scenario-based configuration
4
E VALUATION M ETHODOLOGY
4.1
DATASET D ESIGN
We randomly sampled 1,018 test entries from the following 8 public security datasets: Table 4: Evaluation Datasets Dataset
Description
Attack Scenarios Covered
ALERT (2)
Adversarial LLM prompt dataset Hierarchical safety benchmark with attackenhanced queries Human-generated prompt injection attacks from an online game Prompt injection attack dataset Contextual jailbreak benchmark Harmful instructions with jailbreak prompts from AdvBench and AutoDAN Toxic question answering dataset Aggregated dataset from 30+ public safety sources
Agent threats & harmful content
Salad-Data (3)
Tensor-Trust (4)
PromptWall-Injection (5) CSSBench (6) UltraSafety (7)
ToxicQAFinal (8) Jailbreak-Prompt-Injection (9)
Jailbreak, harmful content
Prompt extraction, prompt hijacking
Instruction override, indirect injection Role hijacking, chain-of-thought poisoning Jailbreak, harmful content
Harmful content Jailbreak, prompt injection, harmful content
All datasets were originally annotated as malicious or harmful inputs. During the evaluation process, we conducted human re-annotation on top of the original labels, independently assessing the actual threat level of each entry: some entries originally labeled as harmful were determined not to pose actual threats in real business scenarios. Entries re-annotated as ALLOWED were retained in the evaluation without exclusion. 4.2
E VALUATION P ROCEDURE
1. Human Re-annotation: For the 1,018 randomly sampled entries, we conducted item-by-item human review based on the original dataset annotations, re-labeling each as BLOCKED (harm3
ful) or ALLOWED (benign). Of these, 852 were labeled BLOCKED and 166 were labeled ALLOWED. 2. API Invocation: All entries were sent to each security guardrail to obtain detection results. 3. Result Normalization: Raw responses from each guardrail were unified into a BLOCKED / ALLOWED binary classification, aligned with human annotations. 4. Comparative Assessment: Each guardrail’s classification results were compared against human annotations to calculate recall rate and true negative rate.
5
E XPERIMENTAL R ESULTS
5.1
C OMPREHENSIVE C OMPARISON
Using human annotations as the ground truth, the recall rate and true negative rate of each guardrail are shown in Tab. 5.
Table 5: Comprehensive Comparison Results (Human Annotation as Ground Truth) Metric
AWS
Azure
DKnownAI
Lakera
Recall (BLOCKED, 852)
743 (87.2%)
715 (83.9%)
822 (96.5%)
812 (95.3%)
True Negative Rate (ALLOWED, 166)
149 (89.8%)
142 (85.5%)
150 (90.4%)
145 (87.3%)
DKnownAI Guard achieves the best overall performance, with a recall rate of 96.5% and a true negative rate of 90.4%. Lakera Guard demonstrates strong recall at 95.3%, ranking second. AWS Guardrails achieves a true negative rate of 89.8%, ranking second in TNR. Azure Content Safety shows relatively lower performance on both metrics. 5.2
E VALUATION D IFFICULTY AND T RUE N EGATIVE R ATE A NALYSIS
The true negative rate for some guardrails in this evaluation is relatively low, which falls within the expected range. The datasets introduced in this evaluation significantly exceed conventional evaluations in both attack intensity and deception level (see §1). The ALLOWED samples are boundary cases selected through systematic human review from predominantly harmful datasets, inherently carrying partial semantic features of harmful data with strong ambiguity. The false positive rate for such high-ambiguity samples is significantly higher than for ordinary benign data. Therefore, the lower true negative rate for some guardrails is a characteristic of the evaluation data distribution rather than a deficiency in the guardrails themselves. Under these conditions, DKnownAI Guard maintains a 90.4% true negative rate, demonstrating its superior ability to distinguish highly deceptive boundary samples compared to other vendors.
6
C ONCLUSIONS
DKnownAI Guard achieves the best overall performance in this evaluation, ranking first in both recall rate and true negative rate. At the same time, the relatively low true negative rate of Azure Content Safety reflects a common deficiency in current security guardrails when dealing with high-ambiguity boundary data. Even the best-performing guardrail still misblocks approximately 10% of benign data, which in real-world deployments may result in legitimate user requests being blocked. Improving the classification precision of security guardrails on high-ambiguity boundary samples—maintaining high detection capability while effectively controlling the false positive rate—remains a critical challenge for the AI agent security community. 4
A
BLOCKED/ALLOWED M APPING L OGIC
All four guardrails employ dual-channel detection. An input is mapped to BLOCKED if either channel triggers, and to ALLOWED only if neither channel triggers. Both channels may trigger simultaneously. AWS Bedrock Guardrails. • BLOCKED: Any content policy filter is triggered (PROMPT ATTACK, HATE, VIOLENCE, MISCONDUCT, INSULTS, or SEXUAL). • ALLOWED: No filter is triggered. Azure Content Safety. • BLOCKED: The shieldPrompt endpoint detects an attack, or the text:analyze endpoint returns any harmful category with severity > 0. • ALLOWED: Neither endpoint reports a detection. Lakera Guard. • BLOCKED: A prompt attack or moderated content/* detector returns a highconfidence result (l1 confident or l2 very likely). • ALLOWED: All detections fall below the confidence threshold or no detection occurs. DKnownAI Guard. • BLOCKED: The detection status is AGENT HACK, SYS FLAG, or CONTENT FLAG. • ALLOWED: The detection status is any other value.
R EFERENCES [1] Q. Li, J. Xu, P. Wei, J. Li, P. Zhao, J. Shi, X. Zhang, Y. Yang, X. Hui, P. Xu, and W. Shao. DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents. arXiv preprint arXiv:2511.03138, 2025. [2] S. Tedeschi, F. Friedrich, P. Schramowski, K. Kersting, R. Navigli, H. Nguyen, and B. Li. ALERT: A Comprehensive Benchmark for Assessing Large Language Models’ Safety through Red Teaming. arXiv preprint arXiv:2404.08676, 2024. [3] L. Li, B. Dong, R. Wang, X. Hu, W. Zuo, D. Lin, Y. Qiao, and J. Shao. SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models. In Findings of ACL, 2024. [4] S. Toyer, O. Watkins, E. A. Mendes, J. Svegliato, L. Bailey, T. Wang, I. Ong, K. Elmaaroufi, P. Abbeel, T. Darrell, A. Ritter, and S. Russell. Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game. arXiv preprint arXiv:2311.01011, 2023. [5] H. Choubey. PromptWall: A Cascading Multi-Layer Firewall for Real-Time Prompt Injection Detection. GitHub, 2025. https://github.com/A73r0id/promptwall [6] Z. Zhou, S. Yan, C. Liu, Q. Li, K. Wang, and Z. Zeng. CSSBench: Evaluating the Safety of Lightweight LLMs against Chinese-Specific Adversarial Patterns. arXiv preprint arXiv:2601.00588, 2026. [7] Y. Guo, G. Cui, L. Yuan, N. Ding, J. Wang, H. Chen, B. Sun, R. Xie, J. Zhou, Y. Lin, Z. Liu, and M. Sun. Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment. In EMNLP, 2024. [8] NobodyExistsOnTheInternet. ToxicQAFinal: Toxic Question Answering Dataset. Hugging Face, 2024. https://huggingface.co/datasets/ NobodyExistsOnTheInternet/ToxicQAFinal 5
[9] Necent. LLM Jailbreak & Prompt-Injection Dataset. Hugging Face, 2026. https://huggingface.co/datasets/Necent/ llm-jailbreak-prompt-injection-dataset
6