WhisperFusion

by Collabora

Ultra-low-latency speech→LLM pipeline: WhisperLive + Mistral + TensorRT.

Looking at WhisperFusion? Try this first.

Drop your audio. Transcript in seconds. First transcript free, then $4 = 300 min

TL;DR

Best for voice-agent prototypes that need 500ms end-to-end speech→reply on a single GPU. Pricing: free.

Category
Open source
License
Stars
★ 1.6k
Last push
2024-07-31
Pricing
free
Platforms
Linux, Docker

What it is

Combines WhisperLive (ASR), Mistral-7B (LLM), and SileroVAD into one Triton-served pipeline. Demonstrates sub-second turn-taking on an RTX 4090. MIT.

Best for: Voice-agent prototypes that need 500ms end-to-end speech→reply on a single GPU.
Watch out for: Requires NVIDIA GPU + TensorRT-LLM; English-tuned; heavy Docker image.

Install / use

Features

Speaker diarizationNo
Word-level timestampsYes
Streaming / real-timeYes
Languages supported99
HIPAA eligibleNo

Links

GitHub repo ↗

WhisperFusion vs Whipscribe

FeatureWhisperFusionWhipscribe
CategoryOpen sourceTranscription APIs
Pricingfree$4–$24 one-time packs (credits never expire) · $2 single unlock · 60 min free on signup
Speaker diarizationNoYes
Word timestampsYesYes
StreamingYesNo
Languages9999
PlatformsLinux, DockerWeb, API, MCP

Alternatives to WhisperFusion

Whipscribe is a managed faster-whisper + whisperX service. If you want transcripts without running infrastructure, paste a URL or drop a file in the form below — you'll have a transcript in seconds.