1.7k Downloads
Overview
Provides a WebSocket-based bridge between an OpenClaw agent and a live voice session, allowing the agent to send and receive transcribed user speech in real time via a bundled Python client script.
Key Advantages
1.Simple, well-documented CLI interface (`send`, `recv`, `listen`, `agent`) for voice message handling.
2.Supports both turn-based (`recv`) and streaming-style (`listen`) conversational flows with configurable timeouts.
3.Includes an `agent` bridge mode that automatically loops between user voice input and `openclaw agent` responses.
4.Clear JSON message schema with distinct types (`message`, `echo`, `pong`) and guidance on what to ignore.
5.Configurable WebSocket endpoint, enabling deployment against different voice server hosts/ports.
Use Cases
- Turn an existing text-based OpenClaw agent into a voice-enabled assistant with minimal changes.
- Run live, real-time voice conversations where user speech is transcribed to text and answered by the agent.
- Prototype or demo voice assistants for customer support, internal tools, or interactive kiosks.
- Collect user utterances via voice for testing and refinement of conversation flows using `listen`.
- Bridge between local development agents and a remote voice interface server over WebSockets.
Evaluation Scores
8.4
/ 10
Reliability
7.6
Functionality
8.8
Usability
8.5
Safety
9.3
Performance
7.8
Compatibility
8.2
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
8.4/103/19/2026▼
OS: darwin-arm64LLM: minimax/minimax-m2.5
**Judgement:** ClawVoice is a focused, practical skill for adding real-time voice I/O to OpenClaw agents via a WebSocket-backed Python client. It appears mature enough for development and demo use, and likely suitable for many production voice-interaction scenarios where basic reliability and responsiveness are acceptable.
**What it does well:**
- Bridges a live voice session to the agent using simple CLI commands (`send`, `recv`, `listen`, `agent`).
- Supports both single-turn and streaming, multi-turn listening with configurable timeouts.
- Provides an `agent` loop mode that continuously forwards user messages to `openclaw agent --agent main` and returns stdout to the user.
- Uses a straightforward JSON protocol with clear message types and behavior guidelines.
**Key risks / limitations:**
- Each interaction is mediated via a Python script (`uv run python client.py`), which may introduce overhead and complicate very high-throughput or low-latency workloads.
- Reliability depends on the external WebSocket voice server; the documentation does not describe reconnection or robust error recovery beyond simple timeouts and exit codes.
- Blocking behavior (`recv`, `listen`, `agent`) requires careful orchestration to avoid deadlocks or unintended parallelism in more complex agent workflows.
- No inherent guardrails on conversational content; safety depends almost entirely on the agent logic using the skill.
**Recommended scenarios:**
- Quickly enabling voice for an existing text-based OpenClaw agent (especially for demos, pilots, and internal tools).
- Real-time conversational assistants where moderate latency is acceptable and traffic volume is not extreme.
- Prototyping or experimenting with voice UX flows using the `listen` and `agent` modes to simulate live sessions.
- Development environments where a local WebSocket voice server (default `ws://localhost:3111/connect`) is available and controllable.
Comments (0)
No comments yet. Be the first!