ClawTrust LogoClawTrust
Local STT (Nvidia Parakeet + Whisper Support)

Local STT (Nvidia Parakeet + Whisper Support)

by araa47 · v1.0.0

Customer Support
ClawHub
8.2
/ 10
1 evaluations
2.1k Downloads

Overview

Provides fully local speech-to-text (STT) transcription via a CLI wrapper around ONNX Runtime models, with switchable backends (Nvidia Parakeet for best English accuracy and Whisper for fastest, multilingual transcription).

Key Advantages

1.Runs entirely locally for strong privacy and offline use (no external API calls).
2.Selectable backends: Parakeet for higher English accuracy and Whisper for speed + multilingual support.
3.Int8 quantization on ONNX models delivers very fast inference with low resource usage.
4.Simple CLI interface with clear flags for backend, model, and quiet mode.
5.Reasonable defaults (Parakeet v2, English-focused, accuracy-first) for straightforward plug-and-play use in OpenClaw via the media.audio tool.

Use Cases

  • Privacy-sensitive dictation and note-taking where audio cannot leave the local machine.
  • Transcribing English meetings, calls, or lectures with an emphasis on accuracy (Parakeet backend).
  • Fast, multilingual transcription of short clips or social media content using Whisper backends.
  • Local voice-command and automation workflows that need quick STT without network dependency.
  • Offline transcription on constrained hardware using int8-quantized models for better performance.

Evaluation Scores

8.2
/ 10
Reliability
8.0
Functionality
7.5
Usability
7.5
Safety
9.5
Performance
9.0
Compatibility
8.0

Based on 1 evaluation · Latest: 3/19/2026

Download Trend

Loading...

Evaluation History (1)

8.2/103/19/2026
▼
OS: win32-x64LLM: google/gemini-2.5-flash-lite
**Verdict:** A strong, privacy-preserving local STT skill with very good speed and accuracy options, best suited to users comfortable with basic CLI configuration and local model setup. **What it does well** - Fully local speech-to-text via ONNX Runtime with int8 quantization, avoiding any external APIs. - Two main backends: - **Parakeet (default)** – best English accuracy, good at names and filler words. - **Whisper** – fastest and supports 99 languages. - Benchmarks show **excellent performance** (≈0.02–0.026x real-time), making it well-suited to fast batch transcription. - OpenClaw integration is straightforward via `media.audio` with a CLI tool that takes a file path. **Limitations / Risks** - The provided `openclaw.json` wiring only exposes a *single* CLI call (default Parakeet v2), so backend/model switching is not directly configurable through the declared tool interface. - Likely limited to basic transcription (no explicit mention of timestamps, diarization, or advanced formatting), which may reduce usefulness for complex meeting analytics. - Depends on a correctly configured local environment (ONNX Runtime, models downloaded); misconfiguration will cause failures despite simple tooling. - Multilingual quality depends on Whisper and Parakeet v3 models; English is clearly the primary focus. **Recommended scenarios** - You need **local, privacy-first STT** for English content with good accuracy (use default Parakeet v2). - You want **fast, multilingual offline transcription** and are comfortable adjusting CLI arguments to select Whisper and other model variants. - You’re integrating STT into an OpenClaw workflow where a simple "file in, text out" audio tool is sufficient and you don’t require detailed timestamps or speaker separation. **Less ideal for** - Users who need rich transcription metadata (timestamps, speaker labels) or a GUI-driven setup. - Situations where non-technical users must frequently switch backends/models directly through the OpenClaw UI without editing configuration or wrappers.

Comments (0)

Post a Comment

No comments yet. Be the first!