2.4k Downloads
Overview
Provide a simple text-to-speech bridge that converts assistant-generated text into MP3 audio using the Hume AI TTS API (preferred) or a legacy OpenAI TTS script, and exposes the resulting file back to the user via MEDIA message.
Key Advantages
1.Straightforward, single-purpose TTS wrapper that fits well when a voice/audio response is requested.
2.Supports two backends (Hume AI preferred, OpenAI as legacy fallback), improving flexibility if one provider is unavailable or deprecated.
3.Outputs standard MP3 files and prints a MEDIA line with an absolute file path, which integrates cleanly with OpenClaw’s message/media tooling.
4.Configuration is environment-variable–driven (API keys via HUME_API_KEY, HUME_SECRET_KEY, OPENAI_API_KEY), avoiding hard‑coded secrets.
5.Existing usage (downloads) suggests basic maturity and that the core flow (text → MP3 → MEDIA link) is already exercised in practice.
Use Cases
- Responding with an audio version of a text answer when the user explicitly asks to hear the response or for a voice message.
- Creating short spoken summaries (e.g., summary of an article, email, or conversation) for users who prefer listening over reading.
- Generating brief greeting messages or announcements that can be downloaded or replayed by the user.
- Producing simple practice audio (e.g., for language learners) where the assistant’s written examples are rendered as speech.
- Providing accessibility-friendly alternatives to text content by offering an MP3 link alongside key responses.
Evaluation Scores
6.9
/ 10
Reliability
6.5
Functionality
6.5
Usability
7.0
Safety
7.0
Performance
7.5
Compatibility
7.5
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
6.9/103/19/2026▼
OS: darwin-x64LLM: google/gemini-3.1-pro-preview
**Judgment:** A solid, narrowly scoped TTS utility that is well-suited for converting assistant replies into downloadable MP3 audio when a user explicitly asks to “hear” something or receive a voice message. It appears technically simple and focused, with a preference for Hume AI and a legacy OpenAI path.
**Strengths:**
- Clear purpose: turn text into speech and surface an MP3 via a MEDIA path.
- Dual backend (Hume AI preferred, OpenAI legacy) offers resilience if one API is unavailable.
- Uses environment variables for API keys, which is operationally standard and avoids embedding secrets.
**Risks / Limitations:**
- No obvious support for advanced controls (voice selection beyond a fixed preferred voice, speed, language, emotion, or length limits), which restricts flexibility.
- Relies on external TTS APIs; failures, latency, or rate limits from providers can affect reliability.
- Any text passed to the skill will be sent to third-party services (Hume/OpenAI), so sensitive or highly private content should be avoided or redacted before use.
- No explicit in-skill safeguards against generating harmful or inappropriate spoken content; relies on the calling assistant to perform content filtering.
**Recommended Scenarios:**
- When a user explicitly requests a spoken or audio version of a response (e.g., “Say this to me out loud,” “Send that as an audio message,” “I want to hear this answer”).
- For brief, well-bounded messages and summaries where reliability and latency are acceptable and content is safe to send to external TTS APIs.
- Less ideal for large-scale batch generation, highly customized voice control, or sensitive legal/medical/personal data that should not leave the platform.
Comments (0)
No comments yet. Be the first!