1.8k Downloads
Overview
Shell-based wrapper around Inworld.ai’s Text-to-Speech API that converts text input into MP3 audio files, with options for voice selection, speaking rate, temperature, model choice, and streaming for long texts.
Key Advantages
1.Straightforward CLI interface for generating MP3 audio from arbitrary text using a single script call.
2.Supports Inworld-specific options such as voice ID, model selection, temperature, and streaming endpoint for long content (>4000 characters).
3.Streaming mode allows handling long-form narration without hitting payload limits, suitable for continuous or large text blocks.
4.Relies only on common Unix utilities (curl, jq, base64), making installation and debugging relatively transparent on Linux/macOS systems.
5.Clear setup instructions for API key configuration and environment variables, plus example commands for quick testing.
Use Cases
- Generating MP3 voice responses from model output in an automated pipeline or agent workflow.
- Creating narrated versions of articles, blog posts, or documentation using Inworld voices.
- Producing character or NPC dialogue audio for games and interactive experiences that already leverage Inworld’s ecosystem.
- Rapid prototyping of voice-enabled agents or bots where text responses need to be turned into audio files on the fly.
- Batch-converting multiple text snippets or scripts into audio assets for content production or internal demos.
Evaluation Scores
7.3
/ 10
Reliability
7.0
Functionality
7.8
Usability
7.2
Safety
6.5
Performance
8.0
Compatibility
7.5
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
7.3/103/19/2026▼
OS: darwin-x64LLM: openai/gpt-5-nano
**Judgement:** A solid, Unix-oriented CLI wrapper around Inworld.ai’s TTS API. It’s well-suited for workflows that want MP3 output from text with configurable voices and support for long-form streaming, but it depends entirely on an external cloud service and a correctly configured API key.
**Strengths & Recommended Scenarios**
- Good fit when you:
- Already use Inworld and want to reuse its TTS voices.
- Need simple text→MP3 generation from scripts or automation.
- Require streaming support for long texts without manually chunking.
- Simple, explicit CLI usage (`./scripts/tts.sh "text" out.mp3 [options]`) and standard dependencies (curl, jq, base64) make it practical for Linux/macOS server environments, CI, or local tooling.
**Key Risks & Limitations**
- **Privacy & data exposure:** All text is sent to Inworld’s servers; not appropriate for highly sensitive or regulated content unless that’s acceptable within your compliance constraints.
- **Operational dependence:** Functionality depends on network reliability, Inworld API uptime, and correct API key permissions. The script itself appears light on robust error handling and retry logic.
- **Platform constraints:** Designed around a Unix shell environment; Windows users without WSL or similar will face friction.
**When to Prefer / Avoid**
- **Prefer this skill** when you need quick, scriptable TTS with Inworld voices and can accept cloud processing.
- **Avoid or supplement it** if you require offline TTS, strict data residency/privacy guarantees, advanced audio controls beyond the provided flags, or first-class support on non-Unix platforms.
Comments (0)
No comments yet. Be the first!