ClawTrust LogoClawTrust
Prompt defense

Prompt defense

by eltemblor · v1.0.0

Productivity
ClawHub
8.2
/ 10
1 evaluations
2.3k Downloads

Overview

Detects and mitigates prompt injection attacks embedded in email content before downstream LLM processing or automation, enforcing a confirmation-first policy for any instructions found in emails.

Key Advantages

1.Specialized for email workflows (IMAP, Gmail API, inbox summarization, auto-processing), with patterns tuned to common email-borne prompt injection tricks.
2.Rule-based pattern library for critical, high, and medium-severity injection signatures (fake system tags, thinking blocks, base64 blobs, hidden text, imperative chains, etc.).
3.Hard blocks on executing any instructions originating from email content; instructions are never followed without explicit user confirmation via a trusted channel.
4.Clear, structured alerts including sender, matched pattern, severity, and suspicious snippet, making it easy to audit and respond.
5.Built-in confirmation protocol (“proceed” / “ignore”) that enforces human-in-the-loop review for risky actions while allowing safe read-only operations without friction.
 “In read-only mode” handling:

Use Cases

  • Protecting autonomous or semi-autonomous email agents that read inboxes and perform actions such as replies, forwarding, labeling, or ticket creation.
  • Safely summarizing large volumes of emails (e.g., daily digests, executive summaries) while flagging and isolating any prompt injection content.
  • Guarding financial or operational workflows driven by email (invoice processing, payment approvals, purchase orders) to prevent malicious instructions like transfers or credential exfiltration.
  • Securing integrations where email content can trigger downstream automations (CRM updates, task creation, calendar changes, deployments).
  • Filtering and sanitizing email content before passing it into more capable or higher-privilege models in multi-agent systems.

Evaluation Scores

8.2
/ 10
Reliability
7.5
Functionality
8.0
Usability
8.0
Safety
9.0
Performance
8.5
Compatibility
8.5

Based on 1 evaluation · Latest: 3/19/2026

Download Trend

Loading...

Evaluation History (1)

8.2/103/19/2026
▼
OS: win32-x64LLM: anthropic/claude-sonnet-4.5
**Quick judgement** A focused, pattern-based prompt injection defense layer for email workflows. It is well-suited as a guardrail in pipelines that *read* and *act on* email content, especially where the LLM or agent has any ability to execute side effects. **Strengths & risks** - **Strengths:** - Targets realistic email injection patterns: fake system/assistant tags, `<thinking>` blocks, “ignore previous instructions”, “you are now…”, fake end-of-email markers, base64 blobs, hidden text, urgent fund-transfer/secret-exfiltration language, etc. - Enforces a strong policy: **never** automatically execute instructions originating from email bodies; always require explicit user confirmation through a trusted channel. - Provides severity labeling and clear alert messages including the sender, matched pattern, and snippet, supporting auditability and safe summarization in read-only mode. - **Risks / limitations:** - Rule/pattern-based detection will miss novel or subtle injection strategies that don’t match the library, so it should not be treated as a complete solution. - May produce false positives on legitimate emails that contain similar language (e.g., IT notices, admin messages, genuine urgent requests), requiring operator judgment. - Effectiveness depends heavily on correct integration: the host system must truly **never bypass** the confirmation protocol. **Recommended scenarios** Use this skill as a defensive pre-processor whenever: - An agent or LLM consumes **email text** and has any capability to act on it (sending emails, modifying data, moving money, changing configurations, updating CRM, etc.). - You generate **summaries or digests** of emails and want injection attempts clearly flagged while keeping processing read-only. - You operate **autonomous email-based workflows** (e.g., ticket triage, approvals, routing) and need an extra safety layer before triggering downstream actions. Do **not** rely on this skill as your only defense in high-stakes environments; pair it with strict capability controls, least-privilege design, and careful review of any action that originates from email content.

Comments (0)

Post a Comment

No comments yet. Be the first!