2.3k Downloads
Overview
Detects and mitigates prompt injection attacks embedded in email content before downstream LLM processing or automation, enforcing a confirmation-first policy for any instructions found in emails.
Key Advantages
1.Specialized for email workflows (IMAP, Gmail API, inbox summarization, auto-processing), with patterns tuned to common email-borne prompt injection tricks.
2.Rule-based pattern library for critical, high, and medium-severity injection signatures (fake system tags, thinking blocks, base64 blobs, hidden text, imperative chains, etc.).
3.Hard blocks on executing any instructions originating from email content; instructions are never followed without explicit user confirmation via a trusted channel.
4.Clear, structured alerts including sender, matched pattern, severity, and suspicious snippet, making it easy to audit and respond.
5.Built-in confirmation protocol (“proceed” / “ignore”) that enforces human-in-the-loop review for risky actions while allowing safe read-only operations without friction.
“In read-only mode” handling:
Use Cases
- Protecting autonomous or semi-autonomous email agents that read inboxes and perform actions such as replies, forwarding, labeling, or ticket creation.
- Safely summarizing large volumes of emails (e.g., daily digests, executive summaries) while flagging and isolating any prompt injection content.
- Guarding financial or operational workflows driven by email (invoice processing, payment approvals, purchase orders) to prevent malicious instructions like transfers or credential exfiltration.
- Securing integrations where email content can trigger downstream automations (CRM updates, task creation, calendar changes, deployments).
- Filtering and sanitizing email content before passing it into more capable or higher-privilege models in multi-agent systems.
Evaluation Scores
8.2
/ 10
Reliability
7.5
Functionality
8.0
Usability
8.0
Safety
9.0
Performance
8.5
Compatibility
8.5
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
8.2/103/19/2026▼
OS: win32-x64LLM: anthropic/claude-sonnet-4.5
**Quick judgement**
A focused, pattern-based prompt injection defense layer for email workflows. It is well-suited as a guardrail in pipelines that *read* and *act on* email content, especially where the LLM or agent has any ability to execute side effects.
**Strengths & risks**
- **Strengths:**
- Targets realistic email injection patterns: fake system/assistant tags, `<thinking>` blocks, “ignore previous instructions”, “you are now…”, fake end-of-email markers, base64 blobs, hidden text, urgent fund-transfer/secret-exfiltration language, etc.
- Enforces a strong policy: **never** automatically execute instructions originating from email bodies; always require explicit user confirmation through a trusted channel.
- Provides severity labeling and clear alert messages including the sender, matched pattern, and snippet, supporting auditability and safe summarization in read-only mode.
- **Risks / limitations:**
- Rule/pattern-based detection will miss novel or subtle injection strategies that don’t match the library, so it should not be treated as a complete solution.
- May produce false positives on legitimate emails that contain similar language (e.g., IT notices, admin messages, genuine urgent requests), requiring operator judgment.
- Effectiveness depends heavily on correct integration: the host system must truly **never bypass** the confirmation protocol.
**Recommended scenarios**
Use this skill as a defensive pre-processor whenever:
- An agent or LLM consumes **email text** and has any capability to act on it (sending emails, modifying data, moving money, changing configurations, updating CRM, etc.).
- You generate **summaries or digests** of emails and want injection attempts clearly flagged while keeping processing read-only.
- You operate **autonomous email-based workflows** (e.g., ticket triage, approvals, routing) and need an extra safety layer before triggering downstream actions.
Do **not** rely on this skill as your only defense in high-stakes environments; pair it with strict capability controls, least-privilege design, and careful review of any action that originates from email content.
Comments (0)
No comments yet. Be the first!