8.4
/ 10
1 evaluations
2.4k Downloads
Overview
Provide advanced, AssemblyAI-powered speech transcription and understanding via a Node-based CLI tailored for agent workflows, with rich exports (Markdown, JSON, subtitles) and post-processing (speaker mapping, translation, LLM-based extraction).
Key Advantages
1.Deep integration with AssemblyAI features, including model routing across universal-3-pro/universal-2, language detection, code switching, and Speech Understanding tasks (topics, entities, sentiment,等
2.Agent-first design: stable, machine-friendly outputs (Markdown + normalized agent JSON), manifest files, and non-interactive CLI suitable for automation pipelines.
3.Rich export formats for downstream use: paragraphs, sentences, SRT/VTT subtitles, raw JSON, agent JSON, and bundle manifests for multi-step workflows.
4.Flexible speaker handling: diarization, manual speaker/channel maps, AssemblyAI speaker identification, and deterministic display-name merging for consistent labels across Markdown/JSON.
5.Support for translation and multi-language workflows, including matching translated utterances back to original segments and explicit language/model selection when needed for accuracy.`,`Integrated Ls
Use Cases
- Transcribing meetings, interviews, podcasts, or calls into Markdown and agent JSON with diarization and named speakers for later summarization or analysis by other agents.
- Generating subtitles (SRT/VTT) and structured paragraph/sentence exports from audio/video for publishing, localization, or content repurposing workflows.
- Running Speech Understanding on existing transcripts to derive topics, entities, sentiment, and enriched transcript outputs without re-transcribing the source audio.
- Performing multi-language or unknown-language transcription with automatic model routing and language detection, especially when language may fall outside universal-3-pro’s core set.
- Post-processing existing AssemblyAI transcript JSON to apply new speaker maps, adjust formatting, or re-render bundles for new downstream consumers without incurring new transcription costs.
Evaluation Scores
8.4
/ 10
Reliability
8.0
Functionality
9.0
Usability
9.0
Safety
7.5
Performance
8.0
Compatibility
8.5
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
8.4/103/19/2026▼
OS: linux-arm64LLM: stepfun/step-3.5-flash
**Judgement**: Strong, agent-focused integration with AssemblyAI that is well-suited for complex transcription workflows, structured exports, and downstream automation. Best used when a user explicitly wants AssemblyAI or needs capabilities beyond generic speech-to-text.
**What it does well**
- Wraps AssemblyAI transcription, language detection, code switching, and Speech Understanding into a non-interactive Node CLI ideal for OpenClaw-style agents.
- Produces stable, machine-readable outputs (Markdown + normalized agent JSON + manifest) that are easy to chain into follow-on tasks.
- Supports diarization, speaker role/name mapping (manual + AssemblyAI identification), translation, subtitles (SRT/VTT), and LLM Gateway–based structured extraction.
- Provides clear recipes for common scenarios (unknown language, known language with universal-3-pro, meetings with speaker labels, translation, LLM extraction).
**Main risks / limitations**
- All audio and text are sent to AssemblyAI, which may be unsuitable for highly sensitive data or strict data-residency requirements (mitigated partly by EU endpoint support, but still a third-party processor).
- Dependent on AssemblyAI availability, latency, and pricing; failures or quota issues in the external API will directly affect this skill.
- Requires a Node environment and correct API key/region configuration; misconfiguration can cause subtle failures or inconsistent behavior.
- Rich option set (speaker maps, understanding, LLM Gateway, multiple exports) increases configuration complexity for simple use cases.
**Recommended scenarios**
- Users explicitly asking for AssemblyAI, or for features like diarization + speaker naming, translation, subtitles, or AssemblyAI’s Speech Understanding.
- Agent workflows that need robust, repeatable transcript bundles (Markdown + JSON + manifest) to feed downstream summarization, extraction, or reasoning agents.
- Multi-language or uncertain-language audio where automatic routing and language detection are valuable.
- Pipelines that already rely on AssemblyAI and want a stable, CLI-based integration with minimal external dependencies.
Comments (0)
No comments yet. Be the first!