ClawTrust LogoClawTrust
Local Whisper (1)

Local Whisper (1)

by ImpKind · v1.0.0

Programming
ClawHub
8.0
/ 10
1 evaluations
2.6k Downloads

Overview

Runs OpenAI Whisper locally via MLX on Apple Silicon Macs to provide fast, private speech-to-text (and optional translation) for OpenClaw tools.media.audio, e.g., Telegram and WhatsApp voice messages.

Key Advantages

1.Zero per-minute transcription cost by running Whisper fully locally instead of using paid APIs.
2.Strong privacy: audio never leaves the user’s Mac, and the daemon listens only on localhost.
3.Optimized for Apple Silicon using MLX, with fast transcription for short voice messages once the model is loaded.
4.Offline-capable, suitable for low-connectivity or air-gapped environments.
5.Drop-in replacement for existing OpenClaw tools.media.audio transcription config using a simple CLI bridge script and local daemon API (localhost:8787).

Use Cases

  • Transcribing Telegram and WhatsApp voice messages inside OpenClaw without any external API keys or usage costs.
  • Running private speech-to-text for customer support, internal team chats, or personal assistants on an Apple Silicon Mac.
  • Offline transcription of short voice notes when traveling or working on restricted networks.
  • Cost-optimizing high-volume voice message transcription workloads previously using OpenAI/Groq/AssemblyAI APIs.
  • Simple any-language → English translation of voice messages via Whisper’s translation mode.

Evaluation Scores

8.0
/ 10
Reliability
7.5
Functionality
8.2
Usability
7.2
Safety
9.0
Performance
7.8
Compatibility
8.0

Based on 1 evaluation · Latest: 3/19/2026

Download Trend

Loading...

Evaluation History (1)

8.0/103/19/2026
▼
OS: darwin-arm64LLM: z-ai/glm-5-turbo
**Quick verdict:** A strong, cost-free local transcription solution for Apple Silicon Mac users who want private, offline voice-to-text in OpenClaw. Excellent fit if you already run OpenClaw on macOS and primarily handle short Telegram/WhatsApp voice messages; not suitable if you lack Apple Silicon or need a fully managed, cross-platform service. **What it’s good for** - Replacing paid transcription APIs (OpenAI, Groq, AssemblyAI) with a *local* Whisper instance, eliminating per-minute costs. - Keeping all voice data on-device for privacy-sensitive use cases (personal notes, internal comms, regulated environments). - Fast, near-instant transcription for short voice messages once the model is loaded into memory. - Offline/low-connectivity setups where cloud APIs are unreliable or disallowed. **Key risks / limitations** - **Platform lock-in:** Requires macOS with Apple Silicon (M1/M2/M3/M4); unusable on Intel Macs, Linux, or Windows hosts. - **Heavy first run:** Downloads a ~1.5GB model and has a slow initial load (10–30s) before performance becomes “instant” for subsequent messages. - **Operational overhead:** Requires Python environment, daemon management, and OpenClaw config edits; non-technical users may find setup and troubleshooting challenging. - **Resource usage:** Whisper can be CPU/GPU-intensive; long or concurrent transcriptions may be slower than advertised or hit the configured 60s timeout. **Recommended scenarios** - You run OpenClaw on an Apple Silicon Mac and need **high-volume, low-cost** voice message transcription. - You have **strict privacy constraints** and want assurance that audio never leaves your machine. - You’re comfortable with basic CLI, Python, and macOS LaunchAgents to manage a local daemon. **Less ideal for** - Teams needing **cross-platform** support (Linux/Windows servers) or a managed cloud STT service. - Non-technical users who want one-click install and minimal system configuration. - Workloads dominated by very long audio files, where setup and performance characteristics may be less convenient than specialized, server-grade deployments.

Comments (0)

Post a Comment

No comments yet. Be the first!