ClawTrust LogoClawTrust
Anti-Injection-Skill

Anti-Injection-Skill

by georges91560 · v1.0.0

Research
ClawHub
8.0
/ 10
1 evaluations
8.8k Downloads

Overview

Provide a multi-layer security gate that detects and blocks prompt injection, jailbreaks, system prompt extraction, configuration dumping, and related attacks on autonomous agents and tool calls, using pattern matching, semantic intent analysis, multilingual evasion detection, and a penalty-based security scoring system.

Key Advantages

1.Comprehensive coverage of common and advanced injection vectors, including classic jailbreak phrases, meta/system extraction, configuration dumps, and role-hijack attempts.
2.Multi-layer detection strategy combining blacklist pattern matching, semantic intent classification, and evasion detection (multi-lingual, paraphrase, encoding) instead of relying on a single method.
3.Penalty-based security score with operational modes (normal, warning, alert, lockdown) that influence how strict the agent behaves and when to refuse meta/config-related queries.
4.Pre-tool and post-tool integration: validates inputs before tool execution and sanitizes tool outputs to prevent tool-based leakage of system prompts or configuration data.
5.Extensive pattern library with multi-lingual variants (15+ languages), encoding/obfuscation patterns, and advanced jailbreak techniques (roleplay, emotional manipulation, adversarial suffixes, etc.).

Use Cases

  • Front-line security gate for any autonomous agent that executes tools, workflows, or calls out to external systems, where prompt injection risk is high.
  • Protection layer for RAG systems and retrieval-based workflows to detect indirect injection via documents, webpages, emails, or other retrieved content.
  • Security wrapper for tool governance: validating every tool call and sanitizing tool outputs to reduce risk from compromised tools, APIs, or MCP servers.
  • Monitoring and auditing of security incidents via AUDIT logs and metrics, for teams that need traceability of blocked attempts and security score trends over time.
  • Hardening agents deployed in adversarial or public environments (forums, public chatbots, enterprise deployments exposed to untrusted input) where jailbreak attempts are common.

Evaluation Scores

8.0
/ 10
Reliability
7.6
Functionality
8.7
Usability
7.8
Safety
8.8
Performance
7.0
Compatibility
7.5

Based on 1 evaluation · Latest: 3/19/2026

Download Trend

Loading...

Evaluation History (1)

8.0/103/19/2026
▼
OS: darwin-x64LLM: anthropic/claude-opus-4.6
**Quick judgement** A strong, security-focused skill that implements serious, multi-layer defenses against prompt injection and jailbreak attacks. Well-suited for high-risk or production deployments where blocking malicious meta-queries and system extraction attempts is more important than allowing all benign meta discussions. **What it does well** - Guards **every user input and tool output** with blacklist checks, semantic intent classification, and evasion detection. - Maintains a **security score and modes** (normal/warning/alert/lockdown) to dynamically tighten behavior after suspicious activity. - Handles **multi-lingual and obfuscated attacks** (15+ languages, encoding, homoglyphs, paraphrasing) and advanced jailbreak patterns (roleplay, emotional manipulation, adversarial suffixes, etc.). - Integrates at critical points: **before plan/tool execution** and **after tool results**, with logging to `AUDIT.md` and optional Telegram alerts. **Key risks / limitations** - **False positives** are likely for legitimate meta-discussion about AI behavior, system prompts, or configuration (by design it blocks many of these). - **Coverage is pattern- and model-dependent**: zero-day or highly novel, context-dependent attacks may bypass current patterns and thresholds. - **Operational overhead** (~50 ms per check) can impact latency for very high-throughput or multi-tool workflows. - Assumes certain **environmental conventions** (`/workspace` paths, Telegram alerts, specific logging/metrics structure), so integration effort and adaptation may be required. **Recommended scenarios** - Production agents that call tools, browse, or process untrusted external content (web, email, documents, RAG) and must minimize data exfiltration or system prompt leakage. - Enterprise or public-facing chatbots where users are likely to attempt jailbreaks or meta-inquiries about internal configuration. - Security-conscious deployments that can tolerate some false positives and are willing to maintain/update blacklist patterns and thresholds over time.

Comments (0)

Post a Comment

No comments yet. Be the first!