NVIDIA FasterTransformer

by NVIDIA

Legacy NVIDIA inference engine — predecessor to TensorRT-LLM.

Looking at NVIDIA FasterTransformer? Try this first.

Drop your audio. Transcript in seconds. First transcript free, then $4 = 300 min

TL;DR

Best for reference implementations of fused Transformer kernels. Pricing: free.

Category
Open source
License
Apache-2.0
Stars
★ 6.4k
Last push
2024-03-27
Pricing
free
Platforms
Linux

What it is

NVIDIA's early high-performance Transformer inference codebase. Apache-2.0.

Best for: Reference implementations of fused Transformer kernels.
Watch out for: Largely superseded by TensorRT-LLM.

Install / use

Features

Speaker diarizationNo
Word-level timestampsYes
Streaming / real-timeNo
Languages supported60
HIPAA eligibleNo

Links

GitHub repo ↗

NVIDIA FasterTransformer vs Whipscribe

FeatureNVIDIA FasterTransformerWhipscribe
CategoryOpen sourceTranscription APIs
Pricingfree$4–$24 one-time packs (credits never expire) · $2 single unlock · 60 min free on signup
Speaker diarizationNoYes
Word timestampsYesYes
StreamingNoNo
Languages6099
PlatformsLinuxWeb, API, MCP

Alternatives to NVIDIA FasterTransformer

Whipscribe is a managed faster-whisper + whisperX service. If you want transcripts without running infrastructure, paste a URL or drop a file in the form below — you'll have a transcript in seconds.