2.5k Downloads
Overview
Generate short lip-synced VRM avatar videos from text or audio and send them as Telegram video notes via OpenClaw workflows.
Key Advantages
1.Turns plain TTS/audio responses into engaging avatar-based video messages with lip sync.
2.Tightly specified pipeline (TTS → avatarcam → ffmpeg → Telegram video note) with example commands and config.
3.Supports configurable avatar model and background (solid color or image) via TOOLS.md.
4.Cross-platform design with explicit instructions for macOS, Linux (headless via xvfb), Windows, and Docker.
5.Output is optimized for Telegram video notes (384x384, 30fps, H.264/AAC, circular format).
Use Cases
- Responding to user requests like “send a video message” or “reply with an avatar video” in chat workflows.
- Delivering TTS replies as video notes instead of audio in Telegram-based assistants.
- Personalized greeting or announcement messages from a branded avatar in community or support channels.
- Short instructional or reminder clips (≤60s) delivered as avatar videos instead of plain text/audio.
- Language-learning or coaching bots that want more engaging, face-like feedback videos.
Evaluation Scores
7.8
/ 10
Reliability
7.0
Functionality
8.0
Usability
8.5
Safety
8.0
Performance
7.0
Compatibility
7.5
Based on 1 evaluation · Latest: 3/20/2026
Download Trend
Loading...
Evaluation History (1)
7.8/103/20/2026▼
OS: darwin-arm64LLM: z-ai/glm-4.5-air
**Judgement:** A solid, well-scoped skill for converting text/audio into short VRM avatar video messages, particularly effective in Telegram environments where video notes increase engagement. Technically mature enough for production use if dependencies are correctly installed.
**Strengths**
- Clear end-to-end workflow: TTS → avatarcam → MP4 → Telegram video note.
- Configurable avatar and background via TOOLS.md.
- Cross-platform with explicit dependency instructions (ffmpeg, xvfb for Linux, Docker guidance).
- Output tuned for Telegram video notes (square, 30fps, circular format) and short clips.
**Key Risks / Limitations**
- **Infra fragility:** Requires ffmpeg and, on Linux/headless, xvfb; misconfigured environments will cause failures that the skill itself doesn’t appear to auto-recover from.
- **Latency & duration limits:** Processing is ~1.5× realtime and capped at 60 seconds, so not suitable for long messages or low-latency critical flows.
- **Channel specificity:** Optimized for Telegram video notes; using it in other channels may need extra adaptation (e.g., ignoring `asVideoNote` or different video specs).
- **Operational details:** Requires correct cleanup of temp files and correct message-tool wiring; mistakes can leak disk usage or send nothing while the agent still replies.
- **Content control:** The avatar will faithfully deliver whatever TTS/audio is provided; upstream safety filters must handle harmful, deceptive, or abusive content.
**Recommended Scenarios**
- Telegram-based assistants where users *explicitly ask* for “video messages”, “avatar videos”, or “video replies”.
- Branded bots or communities that benefit from a consistent avatar persona for announcements and short updates.
- Optional, higher-engagement mode for certain responses (greetings, celebrations, reminders) rather than all replies.
- Environments where you can reliably manage ffmpeg/xvfb and Docker dependencies and accept modest processing latency.
Use this skill when a user clearly wants a video/visual reply or when you have a designed experience around avatar-based messaging; avoid using it by default for all responses or for long-form content where its latency and 60s cap become constraining.
Comments (0)
No comments yet. Be the first!