7.9
/ 10
1 evaluations
3k Downloads
Overview
Provide fully local text-to-speech (TTS) for German and English, generating Telegram-compatible voice note audio from arbitrary text without requiring any cloud services or API keys.
Key Advantages
1.100% offline operation using sherpa-onnx and Piper voices (no network latency, strong privacy).
2.No API keys, accounts, or external cloud dependencies required once installed.
3.Outputs Telegram-ready voice notes using the [[audio_as_voice]] tag and MEDIA: path format.
4.Supports German (thorsten) and English (ryan) voices with automatic language detection based on text content.
5.Configurable via environment variables (SHERPA_ONNX_DIR, PIPER_VOICES_DIR), allowing flexible deployment paths and additional voices if desired.
Use Cases
- Answering user requests for a spoken or voice reply within Telegram (or similar chat setups that honor [[audio_as_voice]]).
- Reading out long responses or summaries for users who prefer listening instead of reading.
- Supporting basic accessibility use cases for visually impaired users who need text content spoken aloud.
- Providing German and English practice audio for language learners using the included thorsten and ryan voices.
- Generating quick local TTS samples or announcements in bots and automation scripts without relying on external TTS APIs.
Evaluation Scores
7.9
/ 10
Reliability
7.5
Functionality
8.0
Usability
7.0
Safety
9.0
Performance
8.5
Compatibility
7.5
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
7.9/103/19/2026▼
OS: win32-x64LLM: moonshotai/kimi-k2.5
**Judgement:** A solid, privacy-friendly local TTS skill that is well-suited for Telegram voice replies in German and English, provided the host environment is correctly configured.
**What it does well**
- Converts arbitrary text to speech entirely offline using Piper via sherpa-onnx.
- Auto-detects language (German vs English) and picks an appropriate voice.
- Emits Telegram-compatible voice notes with the `[[audio_as_voice]]` tag and `MEDIA:` path, making integration straightforward for voice bubbles.
- Avoids cloud costs and latency; once installed, it should be fast and responsive on typical hardware.
**Main risks / limitations**
- **Environment complexity:** Requires sherpa-onnx, Piper voice models, and ffmpeg installed in specific locations with environment variables properly set; misconfiguration will cause failures.
- **Platform assumptions:** Documentation targets Linux (x64); other platforms may need manual adaptation and could be less reliable.
- **Limited out-of-the-box voices:** Only one German and one English voice are included by default; adding more voices requires manual script changes.
- **Language scope:** Auto-detection is designed around German vs English; content in other languages will likely fall back to English or behave suboptimally.
**Recommended scenarios**
- Bots or assistants running on self-hosted servers where privacy and offline capability are important.
- Telegram-based assistants that need to respond with voice notes when users ask to "hear" the answer or request a spoken reply.
- Workflows where network access is unreliable or restricted, but TTS is still required.
- Privacy-sensitive deployments (e.g., internal tools) that must avoid sending text to external TTS providers.
Comments (0)
No comments yet. Be the first!