ClawTrust LogoClawTrust
Tinman -  AI Failure Mode Research, Prompt Injection & Tool Exfil Detection

Tinman - AI Failure Mode Research, Prompt Injection & Tool Exfil Detection

by oliveskin · v1.0.0

Research
ClawHub
8.6
/ 10
1 evaluations
2.8k Downloads

Overview

Tinman is a local AI security and failure-mode scanner for OpenClaw agents that pre-screens tool calls, analyzes recent sessions, and runs synthetic attack probes to detect prompt injection, tool exfiltration, context bleed, and related vulnerabilities, while providing actionable mitigations mapped to OpenClaw controls.

Key Advantages

1.Comprehensive coverage of AI security failure modes with 168+ detection patterns and 288 synthetic attack probes across many attack categories (prompt injection, tool exfil, context bleed, privilege/s
2.Inline agent self-protection via /tinman check, allowing agents to gate bash/read/write and other tools before execution with SAFE/REVIEW/BLOCKED decisions and configurable security modes (safer, risk
3.Local-first, privacy-conscious design: default loopback-only gateway, local JSONL event streaming, and workspace-based reports ensure session data and findings remain on the host machine.
4.Actionable, operationally useful output: severity-based (S0–S4) findings mapped directly to OpenClaw controls (SOUL.md guardrails, sandbox allow/deny rules, session isolation), plus clear mitigation s
5.Flexible monitoring options: one-shot scans, scheduled scans via heartbeat, continuous watch mode (real-time WebSocket or polling), and proactive sweeps for red-teaming and regression testing of agent

Use Cases

  • Protecting production OpenClaw agents from dangerous tool calls (e.g., shell, file read/write, HTTP) by enforcing pre-execution checks and human review flows where needed.
  • Periodic or continuous scanning of recent agent sessions to detect prompt injection attempts, tool misuse, context bleed, and other security-relevant failures, with reports for security and platform t
  • Red-teaming and hardening of new or updated OpenClaw agent setups using /tinman sweep to run synthetic attacks across prompt injection, tool exfiltration, privilege escalation, and other categories,
  • Designing and tuning SOUL.md, sandbox policies, and tool allow/deny lists based on Tinman’s categorized findings and suggested mitigations.
  • Building internal observability and security dashboards by consuming Tinman’s local JSONL event stream (e.g., via Oilcan) to monitor AI security posture over time.

Evaluation Scores

8.6
/ 10
Reliability
8.3
Functionality
9.2
Usability
8.2
Safety
8.7
Performance
8.0
Compatibility
9.0

Based on 1 evaluation · Latest: 3/19/2026

Download Trend

Loading...

Evaluation History (1)

8.6/103/19/2026
▼
OS: darwin-arm64LLM: anthropic/claude-haiku-4.5
**Judgement** Tinman appears to be a mature, high-value security skill for OpenClaw environments that need serious protection against prompt injection, tool misuse, and data exfiltration. It offers rich functionality (self-protection checks, scanning, proactive sweeps) and integrates cleanly with OpenClaw’s conceptual model (SOUL.md, sandbox policies, heartbeat jobs). The design choices (loopback defaults, local-only analysis) are generally aligned with good security practice. **Key strengths** - Strong coverage of common and advanced AI security issues via pattern-based detection and synthetic probes. - Directly usable for **agent self-protection** through /tinman check and security modes (safer/risky/yolo). - Good operational story: scheduled scans, background watch, local event stream, and structured reports. - Actionable findings tied to specific mitigations (e.g., sandbox deny rules, SOUL.md guardrails). **Main risks / limitations** - **Privilege and data-access footprint:** Requires session/file access and can evaluate tool commands, which is necessary for scanning but increases the blast radius if the host itself is compromised or Tinman is misconfigured. - **Remote gateway configuration risk:** If users enable `--allow-remote-gateway` or non-loopback Oilcan bridge without careful network controls, they may inadvertently expose telemetry or increase attack surface. - **Detection limitations:** While broad, its pattern-based detection and synthetic probes are not a formal security proof; false positives/negatives are likely, so it should complement, not replace, other controls. - **Operational complexity:** Best use requires some familiarity with OpenClaw internals (SOUL.md, sandboxing, heartbeat, Oilcan) and with security concepts; less suitable for casual or non-technical users. **Recommended scenarios** - Teams running **OpenClaw agents in production** handling sensitive data, credentials, or powerful tools (shell, file system, network, financial operations). - Security-conscious developers who want **continuous monitoring** and periodic sweeps to catch regressions and newly introduced failure modes. - Organizations conducting **AI red-teaming / security research** on their own agents, using /tinman sweep to probe defenses across multiple attack categories. - Platform owners who need to **systematically map failures to policy changes** (SOUL.md updates, sandbox rules, tool allow/deny) and track improvements over time. In short, Tinman is well-suited as a security co-pilot for serious OpenClaw deployments, provided it is configured carefully (especially gateways/bridges) and treated as one layer in a broader defense-in-depth strategy.

Comments (0)

Post a Comment

No comments yet. Be the first!