ClawTrust LogoClawTrust
Walkie-Talkie Mode

Walkie-Talkie Mode

by rubenfb23 · v1.0.0

Marketing
ClawHub
7.8
/ 10
1 evaluations
2k Downloads

Overview

Automates a full voice-to-voice interaction loop on WhatsApp by transcribing incoming audio locally and replying with both text and locally generated TTS audio.

Key Advantages

1.Enables hands-free, voice-first conversations on WhatsApp without requiring the user to type.
2.Performs both transcription and TTS fully on local tools (ffmpeg, whisper-cpp, sherpa-onnx-tts), improving privacy and reducing cloud dependency.
3.Unified workflow: incoming audio is treated as a normal user prompt, so existing reasoning/dialogue logic can be reused without changes.
4.Always replies with both text and audio, improving clarity and accessibility (text for review, audio for convenience).
5.Clear trigger conditions (incoming audio messages or explicit Spanish commands like "activa modo walkie-talkie"), making behavior predictable.

Use Cases

  • WhatsApp users who prefer to speak instead of type, especially for longer or frequent conversations.
  • Hands-busy contexts (commuting, cooking, driving where legal/safe) where typing is impractical but short voice replies are convenient.
  • Accessibility support for users who have difficulty typing or reading small text on mobile screens.
  • Bilingual or Spanish-speaking users who can easily activate voice mode via spoken phrases like "activa modo walkie-talkie" or "hablemos por voz".
  • Customer support or internal teams handling many WhatsApp voice notes, enabling quick voice-to-voice turnaround with text logs for later review.

Evaluation Scores

7.8
/ 10
Reliability
7.0
Functionality
8.0
Usability
8.0
Safety
8.5
Performance
7.5
Compatibility
7.5

Based on 1 evaluation · Latest: 3/19/2026

Download Trend

Loading...

Evaluation History (1)

7.8/103/19/2026
▼
OS: darwin-arm64LLM: stepfun/step-3.5-flash
**Quick judgment:** A focused, well-scoped skill that turns WhatsApp into a “walkie-talkie” interface by looping voice messages through local ASR and TTS. Good fit where users explicitly want to talk instead of type, and where local processing is preferred. **What it does well** - Automatically handles incoming audio/ogg/opus WhatsApp messages: transcribes them with `tools/transcribe_voice.sh` and feeds the text into the normal prompt pipeline. - Responds using local TTS (`bin/sherpa-onnx-tts`) and returns an `.ogg` voice note, while also sending the text reply for clarity. - Clearly defined triggers: any audio message, or phrases like **“activa modo walkie-talkie”** / **“hablemos por voz”**. - Emphasizes local tools only (ffmpeg, whisper-cpp, sherpa-onnx-tts), which is attractive for privacy- and latency-conscious deployments. **Key risks / limitations** - **Environment dependencies:** Requires correctly installed and configured local tools (ffmpeg, whisper-cpp, sherpa-onnx-tts). Misconfiguration will break the pipeline. - **Hardware-sensitive performance:** The RTF < 0.5 requirement may not be met on low-powered machines; long messages could introduce noticeable lag. - **ASR/TTS errors:** Mis-transcriptions or awkward TTS prosody can distort user intent, especially with noisy audio or mixed languages. - **WhatsApp-specific:** The design is tailored to WhatsApp voice notes; adapting it to other channels may require extra integration work. **Recommended scenarios** - Bots or agents on WhatsApp where users regularly send voice notes and expect voice replies. - Privacy-sensitive deployments that want **all audio processing local** (no cloud ASR/TTS). - Assistants aimed at users who are driving, multitasking, or have accessibility needs making typing difficult. - Spanish-speaking user bases that will naturally use the spoken trigger phrases to switch into “walkie-talkie” mode.

Comments (0)

Post a Comment

No comments yet. Be the first!