2.1k Downloads
Overview
Automatically transcribe incoming Telegram voice messages (.ogg Opus) to text using a locally installed OpenAI Whisper tiny model, then reply with the transcription and securely delete the audio and intermediate files.
Key Advantages
1.Fully local transcription with no external APIs or credentials required, preserving privacy and avoiding rate limits.
2.Optimized for Telegram’s .ogg Opus voice messages and integrates directly with the OpenClaw media inbound directory.
3.Uses the lightweight Whisper tiny model for fast turnaround on modest hardware (sub‑second after cache on 1 vCPU / 4GB RAM).
4.Built‑in flow to auto‑reply with the transcribed text and remove original media and temporary files for lower storage footprint.
5.Supports Russian and English well, with option to use language auto‑detection or upgrade to larger Whisper models for better accuracy.
Use Cases
- Auto‑transcribing Telegram voice messages in an OpenClaw‑backed Telegram bot or assistant, replying with text in chat.
- Creating a lightweight, privacy‑preserving voicemail‑to‑text system for Russian and English speakers using Telegram voice notes.
- Assisting users who prefer speaking over typing by turning their Telegram voice inputs into text commands or logs.
- Building monitoring or logging tools that capture and archive the text content of voice messages for later search and analysis.
- Prototyping voice‑driven workflows where quick, approximate transcription is enough and low resource usage is important.
Evaluation Scores
7.8
/ 10
Reliability
7.2
Functionality
8.3
Usability
7.5
Safety
8.0
Performance
8.4
Compatibility
7.8
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
7.8/103/19/2026▼
OS: linux-x64LLM: z-ai/glm-5
**Quick verdict:** A focused, privacy‑friendly Telegram voice transcriber that runs Whisper locally and auto‑replies with text. Well‑suited for Telegram bots/assistants that need quick Russian/English transcription without external APIs.
**What it does well**
- Transcribes `.ogg` Opus Telegram voice messages using local Whisper **tiny** (no API keys, no network needed after initial model download).
- Integrates naturally with OpenClaw by watching `/root/.openclaw/media/inbound/` and replying via `message action=send`.
- Auto‑deletes both the original voice file and temporary Whisper output, keeping storage usage low and improving privacy.
- Good performance on modest hardware (after first run cache, ~<1s per clip on 1 vCPU/4GB as claimed).
**Main limitations & risks**
- Uses the **tiny** Whisper model by default: accuracy is acceptable but not state‑of‑the‑art, especially for noisy audio, accents, or long messages.
- Primarily tuned for **Russian and English**; other languages may be less accurate without adjusting model/language parameters.
- Installation requires `ffmpeg` and `openai-whisper` (with `--break-system-packages`), which can conflict with system package management on some environments.
- The provided `rm PATH /tmp/whisper/*` cleanup pattern must be used carefully to avoid accidental deletion if variables or paths are misconfigured; operators should validate paths and possibly narrow globs.
- Error handling and edge cases (corrupted files, concurrent writes, very large messages) are not deeply documented and may need additional wrapper logic for production.
**Best‑fit scenarios**
- Telegram‑based assistants that need **fast, cheap, offline** voice‑to‑text for Russian/English users.
- Privacy‑sensitive deployments where sending audio to third‑party APIs is not acceptable.
- Lightweight servers (e.g., single‑vCPU, 4GB RAM) where full‑size ASR models would be too heavy but approximate transcription is good enough.
**Use with caution if** you need high‑accuracy transcripts across many languages, strong guarantees around system package integrity, or robust fault tolerance for high‑volume production workloads without additional engineering around this skill.
Comments (0)
No comments yet. Be the first!