7.8
/ 10
1 evaluations
4.3k Downloads
Overview
Command-line tool to transcribe local audio files via OpenAI’s speech-to-text API, with support for custom vocabulary hints, text replacement post-processing, and caching based on file hash.
Key Advantages
1.Simple CLI workflow for turning voice memos and other audio files into text.
2.Supports many common audio formats (mp3, mp4, mpeg, mpga, m4a, wav, webm, ogg, opus).
3.Custom vocabulary file (vocab.txt) helps recognize domain-specific names and jargon.
4.Deterministic text cleanup via replacements.txt for known recurring transcription errors.
5.Content-addressable caching by SHA256 of the audio file avoids re-transcribing unchanged files, saving time and API cost.','Easy integration into shell pipelines (e.g., piping output to clipboard or其他
Use Cases
- Transcribing WhatsApp or other messaging app voice memos into text for quick reading or replying.
- Converting short meeting snippets, standups, or ad-hoc spoken notes into written summaries or to-dos.
- Processing podcast or interview clips into text for drafting posts, highlights, or search indexing.
- Piping transcriptions into other CLI tools or scripts (e.g., pbcopy, grep, or custom analyzers) for downstream automation.
- Maintaining high-quality transcriptions in domains with unusual names/terms by iteratively refining vocab.txt and replacements.txt.
Evaluation Scores
7.8
/ 10
Reliability
7.6
Functionality
8.2
Usability
8.0
Safety
7.0
Performance
8.5
Compatibility
7.5
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
7.8/103/19/2026▼
OS: darwin-arm64LLM: z-ai/glm-5
**Judgment:** A focused, practical CLI skill for English speech-to-text on local audio files, well-suited for power users who live in the terminal and already use OpenAI’s APIs.
**What it does well:**
- Quickly transcribes a wide range of audio formats using an OpenAI transcription model.
- Integrates cleanly into shell workflows (`transcribe <file>`, pipe to other tools, etc.).
- Offers two mechanisms to improve accuracy over time: a vocabulary hints file (`vocab.txt`) and deterministic text replacements (`replacements.txt`).
- Caches by audio file hash to avoid paying and waiting for repeated transcriptions of the same content.
**Key risks / limitations:**
- **Privacy & data handling:** All audio is sent to OpenAI’s servers; not appropriate for highly sensitive or regulated audio unless that’s acceptable under your data policies.
- **Environment assumptions:** Requires `uv` and a working OpenAI API key configured in a `.env` file; the example path is user-specific and may confuse less-experienced users.
- **Language constraint:** Assumes English; there’s no language detection or explicit multilingual handling.
- **Network/API dependency:** Reliability and latency depend on OpenAI’s API availability and your network; there’s no mention of robust error handling or retries.
**Recommended scenarios:**
- You frequently receive short audio messages (e.g., WhatsApp voice notes) and want fast, scriptable transcription on your machine.
- You’re comfortable with CLI tooling and want to chain transcription into broader automation (copy to clipboard, search, summarize, etc.).
- You need to iteratively improve transcription for specific jargon, product names, or personal names via custom vocab and replacements.
**Less ideal for:**
- Non-technical users who aren’t comfortable with terminal tools or environment configuration.
- Use cases that demand strict on-device processing or strong guarantees that audio never leaves your infrastructure.
- Multilingual or language-detection-heavy workflows.
Comments (0)
No comments yet. Be the first!