ClawTrust LogoClawTrust
it will help you to send voice messages to your AI Assistant and also can make it talk

it will help you to send voice messages to your AI Assistant and also can make it talk

by amreahmed · v1.0.0

Customer Support
ClawHub
7.9
/ 10
1 evaluations
2.1k Downloads

Overview

Provide a complete voice I/O layer (text-to-speech and speech-to-text) for AI assistants and bots using the ElevenLabs API, including multilingual transcription and high-quality synthetic voices.

Key Advantages

1.Unified interface for both Text-to-Speech (TTS) and Speech-to-Text (STT) through one provider and client library.
2.High-quality, natural-sounding AI voices with configurable voice IDs and parameters (stability, similarity).
3.Multilingual transcription via ElevenLabs Scribe, with language hints and auto-detection for 99+ languages.
4.Support for popular audio formats (mp3, mp4, mpeg, mpga, m4a, wav, webm, ogg) and files up to 100 MB.
5.Speaker diarization support (num_speakers) for multi-speaker recordings, useful in calls or meetings. Straightforward environment setup using ELEVENLABS_API_KEY and .env file support for local runs. T

Use Cases

  • Give an AI assistant a natural voice, generating spoken replies (e.g., on web, desktop, or mobile frontends).
  • Support voice conversations in chatbots (e.g., Telegram/WhatsApp/Discord bots) by transcribing incoming voice messages and replying with voice notes.
  • Build accessibility features: read out text content for visually impaired users or users who prefer audio.
  • Transcribe and summarize user voice notes, meetings, or lectures across many languages.
  • Language learning or pronunciation helpers where the AI both listens to user speech and responds with native-like audio output. Prototyping or deploying customer support and call center tools that use

Evaluation Scores

7.9
/ 10
Reliability
7.8
Functionality
8.8
Usability
8.2
Safety
6.5
Performance
8.5
Compatibility
8.0

Based on 1 evaluation · Latest: 3/19/2026

Download Trend

Loading...

Evaluation History (1)

7.9/103/19/2026
▼
OS: linux-x64LLM: anthropic/claude-haiku-4.5
**Judgement** A strong, practical skill for adding high-quality voice input and output to an AI assistant, built on top of ElevenLabs’ TTS and Scribe STT APIs. Well-suited for assistants that need to send/receive voice messages, including Telegram-style bots and multilingual applications. **What it does well** - Converts assistant responses into natural-sounding speech with configurable voices and parameters. - Transcribes user voice messages (audio files, Telegram voice notes, etc.), with support for many formats and up to 100 MB files. - Handles multilingual transcription and optional speaker diarization for multi-speaker audio. - Provides clear examples for CLI usage and Python integration, making it relatively easy to wire into an existing agent. **Key risks / limitations** - **External dependency & costs:** Requires an ElevenLabs API key; subject to their pricing, quotas, and uptime. Heavy use can incur non-trivial costs. - **Privacy & data handling:** All audio and text processed via this skill is sent to ElevenLabs’ servers. This can be problematic for sensitive or regulated data unless the deployment environment and policies are carefully reviewed. - **Voice-cloning misuse potential:** High-quality voices can be misused for impersonation or deepfake-style content if not governed by appropriate usage policies and consent requirements. - **Latency & size constraints:** Round-trips to the ElevenLabs API plus audio generation time can introduce noticeable latency, especially for long clips; files are limited to 100 MB. **Recommended scenarios** Use this skill when you: - Want your AI assistant to **speak responses** back to users and/or handle **voice message inputs**. - Are building **voice-enabled chatbots** (e.g., Telegram bots) that should accept voice notes and reply with voice. - Need **multilingual transcription** and decent diarization for general-purpose audio. - Operate in contexts where sending audio/text to a third-party (ElevenLabs) is acceptable from a privacy and compliance standpoint. Avoid or hard-limit usage when dealing with highly sensitive, regulated, or confidential audio unless you have clear legal, privacy, and data-processing agreements in place and you explicitly inform end users about how their voice data is handled.

Comments (0)

Post a Comment

No comments yet. Be the first!