8.7
/ 10
1 evaluations
2.5k Downloads
Overview
Input Guard scans untrusted external text for prompt injection and related adversarial patterns, returning a severity score and structured findings so agents can safely decide whether to ignore, block, or escalate the content.
Key Advantages
1.Defensive-by-design prefilter for any untrusted text (web pages, social feeds, API responses) before the agent reasons over it.
2.Rich detection taxonomy (16 categories) covering instruction override, jailbreaks, data exfiltration, safety bypass, encoded payloads, and more.
3.Multi-language pattern support (English, Korean, Japanese, Chinese) for globally sourced content.
4.Multiple sensitivity levels (low/medium/high/paranoid) to tune false-positive vs. coverage trade-offs per workflow.
5.Optional LLM-based semantic analysis layer that augments regex patterns and catches more evasive, indirect, or storytelling-based attacks, with explicit merge logic and confidence-based downgrades (no
Use Cases
- Front-line filter for web_fetch or browser-based tools: scan page snapshots or HTML before summarization, extraction, or reasoning.
- Guarding social media integrations (X/Twitter, threads, comments) to prevent jailbreaks or instruction overrides embedded in posts.
- Scanning search results (Brave, SerpAPI, etc.) before feeding snippets into downstream tools or agents.
- Protecting against data exfiltration attempts when users or third-party APIs return content that tries to extract system prompts, secrets, or API keys.
- Bulk or real-time scanning of third-party API responses in workflows where external services can be influenced by adversaries (chatbots, content aggregators, monitoring tools).
Evaluation Scores
8.7
/ 10
Reliability
8.2
Functionality
9.0
Usability
8.8
Safety
9.3
Performance
8.0
Compatibility
8.5
Based on 1 evaluation · Latest: 3/20/2026
Download Trend
Loading...
Evaluation History (1)
8.7/103/20/2026▼
OS: linux-arm64LLM: google/gemini-2.5-flash
**Judgement:** Input Guard is a strong, opinionated defensive layer for OpenClaw-style agents that consume untrusted text. It is well-suited as a default prefilter in any workflow that fetches external content.
**What it does well**
- Detects a broad range of prompt-injection tactics (instruction override, role/system spoofing, jailbreaks, data exfiltration, encoded payloads, etc.).
- Works offline with pure-Python pattern matching and no external dependencies; can optionally add an LLM layer for deeper semantic checks.
- Provides clear severities (SAFE → CRITICAL), consistent exit codes, and JSON output that is easy to script around.
- Offers good operational guidance: integration templates, AGENTS.md snippet, alert formats, MoltThreats reporting hooks.
**Key risks / limitations**
- Pattern-based scanning will inevitably have both false negatives (novel or highly obfuscated attacks) and false positives (overly strict regex hits), especially at higher sensitivities.
- LLM-based analysis improves coverage but introduces latency, cost, and dependence on external APIs; it can still be manipulated and must not be treated as infallible.
- Severity thresholds (e.g., blocking on MEDIUM+) are policy choices; aggressive settings may disrupt workflows if not tuned and monitored.
**Recommended scenarios**
- As a **mandatory pre-processing step** for any skill that ingests external web pages, search results, or social media content.
- In **security-conscious or regulated environments** where prompt injection and data-exfiltration attempts must be surfaced and logged to humans.
- For **pipelines that call third-party APIs** whose outputs can be influenced by adversaries (user-content platforms, forums, social feeds).
- With LLM scanning enabled (`--llm` or `--llm-auto`) for **high-stakes or high-risk inputs**, while using pattern-only scanning for fast, bulk, or cost-sensitive use cases.
Comments (0)
No comments yet. Be the first!