2k Downloads
Overview
A meta-quality assurance skill that evaluates other Clawdbot/OpenClaw skills using a structured, multi-framework rubric and optional automated static checks (Python script), producing a scored, prioritized readiness assessment before publishing.
Key Advantages
1.Combines automated structural checks (syntax, file structure, metadata, dependency/credential scans) with a detailed manual rubric for holistic evaluation.
2.Grounded in recognized frameworks (ISO 25010, OpenSSF, Shneiderman, Tognazzini, Norman, agent-specific heuristics), giving evaluations a defensible, standardized basis.
3.Clear, repeatable process: run script, read SKILL.md, skim code, score 25 criteria, and write EVAL.md using a provided template.
4.Prioritization of issues by severity (P0/P1/P2) directly aligned with publish-readiness decisions and remediation planning.
5.Covers agent-specific concerns (trigger precision, composability, idempotency, escape hatches) that generic software QA frameworks usually miss.
Provides an interpretable 0–100 scoring model with well
Use Cases
- Pre-publication review of new skills to decide whether they’re ready for Clawhub or internal catalogs.
- Systematic audit of existing skills to identify technical debt, security risks, and usability issues for roadmap planning.
- Establishing a standard review rubric for a team so multiple reviewers can evaluate skills consistently over time.
- Automated pre-checks in a CI pipeline (running eval-skill.py in JSON mode) to catch structural and configuration issues early.
- Guiding less-experienced skill authors through what “good” looks like across functionality, reliability, security, and usability dimensions.
Evaluation Scores
8.7
/ 10
Reliability
8.0
Functionality
8.7
Usability
9.3
Safety
8.8
Performance
9.2
Compatibility
8.5
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
8.7/103/19/2026▼
OS: linux-x64LLM: z-ai/glm-4.5-air
**Quick judgement**
Strong, well-structured evaluator skill that appears publish-ready and highly useful as a standard QA tool for OpenClaw/Clawdbot skills. It offers a solid mix of automated checks and a comprehensive manual rubric grounded in reputable frameworks.
**Strengths**
- Multi-framework rubric (ISO 25010, OpenSSF, HCI + agent-specific heuristics) covering 25 criteria across 8 categories.
- Concrete scoring guidance (0–4 per criterion, 0–100 total) and clear interpretation bands (Excellent/Good/Acceptable/etc.).
- Automated checks for file structure, frontmatter, script syntax, dependency/credential issues, and env-var documentation.
- Well-documented workflow (quick start, stepwise evaluation process, EVAL.md template) that supports consistent use across reviewers.
**Key risks / limitations**
- Relies on a local Python environment (3.6+ with PyYAML), so automation depends on correct tooling setup and filesystem access.
- Manual scoring still depends on reviewer judgment; consistency across reviewers will vary unless teams train on the rubric.
- Security coverage is intentionally “baseline”; could create a false sense of completeness if users don’t also run deeper tools (e.g., SkillLens).
- Quality of outcomes depends on how thoroughly reviewers read SKILL.md and the scripts; the tool guides but does not enforce depth of review.
**Recommended scenarios**
- Use as the default gatekeeper before publishing community or internal skills to Clawhub.
- Integrate the automated eval-skill.py checks into CI to block structurally invalid or obviously unsafe submissions.
- Standardize team code reviews around its rubric so multiple reviewers can produce comparable evaluations.
- Pair with a dedicated security scanner (like SkillLens) for high-risk or sensitive skills where security is critical.
Comments (0)
No comments yet. Be the first!