ClawTrust LogoClawTrust
Parakeet Stt

Parakeet Stt

by carlulsoe · v1.0.0

Programming
ClawHub
8.6
/ 10
1 evaluations
2k Downloads

Overview

Local speech-to-text (STT) server using NVIDIA Parakeet TDT 0.6B v3 via ONNX Runtime, exposing an OpenAI-compatible `/v1/audio/transcriptions` API for fast multi-language audio transcription on CPU.

Key Advantages

1.Runs entirely locally on CPU (no GPU or external cloud APIs required), improving privacy and easing deployment on commodity servers.
2.OpenAI-compatible transcription endpoint, making it a near drop-in replacement for existing Whisper/OpenAI STT integrations and the OpenAI Python SDK.
3.High claimed performance (~30x faster than realtime on CPU) with support for timestamps, word-level/segment metadata, and subtitle formats (SRT, VTT).
4.Automatic language detection across 25 supported languages, removing the need for manual language configuration in most workflows.
5.Provides multiple response formats (plain text, JSON, verbose JSON with segments, SRT, VTT) suitable for downstream processing and media workflows, plus a simple web UI for drag-and-drop transcription

Use Cases

  • Batch transcription of locally stored audio files (e.g., MP3, WAV) into plain text for search, summarization, or archiving.
  • Generating subtitles (SRT or VTT) for video content such as lectures, podcasts, tutorials, or internal training materials.
  • Building or replacing existing voice-processing pipelines that currently rely on OpenAI Whisper STT, by pointing clients to a local OpenAI-compatible endpoint.
  • Privacy-sensitive transcription scenarios where audio cannot leave a controlled environment (e.g., internal company meetings, research interviews).
  • Multi-language transcription projects across the 25 supported European languages with automatic language detection for mixed or unknown-language recordings.

Evaluation Scores

8.6
/ 10
Reliability
8.0
Functionality
8.6
Usability
8.5
Safety
9.1
Performance
8.9
Compatibility
9.0

Based on 1 evaluation · Latest: 3/19/2026

Download Trend

Loading...

Evaluation History (1)

8.6/103/19/2026
▼
OS: linux-x64LLM: anthropic/claude-opus-4.6
**Quick judgment** A strong local speech-to-text solution that closely mimics OpenAI’s transcription API while running entirely on CPU. It’s particularly compelling if you want Whisper-like accuracy and formats without sending audio to external cloud services. **What it does well** - Provides a local `/v1/audio/transcriptions` endpoint that works with the OpenAI Python SDK by just changing `base_url` and using a dummy API key. - Supports multiple response formats (text, JSON, verbose JSON with timestamps, SRT, VTT) and a browser-based drag-and-drop UI. - Claims very high CPU performance (~30x faster than realtime) and Whisper large-v3–comparable accuracy, with automatic language detection across 25 languages. **Key risks / limitations** - Accuracy claims versus Whisper are from the project description; there is no independent benchmarking here, so expect some variability by language, accent, and audio quality. - Language coverage is limited to the listed 25 mainly European languages; it may perform poorly or unpredictably on unsupported languages or heavy code-switching outside that set. - Being self-hosted, reliability and stability depend on your Docker/Python/ONNX environment and resource provisioning (CPU, memory, disk I/O). - CPU-only is convenient but still potentially CPU-intensive for large-scale or concurrent workloads; under-provisioned systems may see slowdowns or queueing. **Recommended scenarios** - You want a **drop-in, local replacement for OpenAI/Whisper transcription** with minimal code changes (e.g., just point your client to `PARAKEET_URL`). - You need **privacy-preserving transcription** of internal meetings, interviews, or proprietary content where cloud services are not acceptable. - You run **media workflows** that need timestamps and subtitle outputs (SRT/VTT) for podcasts, lectures, or training videos. - You have **CPU-only infrastructure** and still need high-throughput, multi-language STT without managing GPUs or paying per-minute cloud fees.

Comments (0)

Post a Comment

No comments yet. Be the first!