Tools/SYSTRAN/faster-whisper

faster-whisper

Faster Whisper transcription with CTranslate2

25.7k ★+88/wkgrowthPythonMIT License →

The Lens

By Erik Loyd, SaaS CEO and former COO/CFO of an AWS Premier Partner.

Updated Sep 2026

faster-whisper runs OpenAI's Whisper speech-to-text models up to four times faster than the original implementation while using less memory, with the same accuracy. It reimplements inference on CTranslate2, and it has quietly become the engine inside most self-hosted transcription stacks. MIT licensed, completely free.

It's a Python library, not a service: pip install, pick a model size, feed it audio. A GPU gets you faster-than-realtime transcription; CPU works fine with the small and medium models. Wrapping it in an internal API for your team is an afternoon with FastAPI, and 8GB of VRAM comfortably runs the large model in int8.

Anyone paying AssemblyAI or Deepgram per minute for plain transcription should run the math. A cheap GPU instance chews through hours of audio for pennies, and openai/whisper accuracy is the same thing you're renting. The paid APIs keep winning on streaming, speaker diarization, and zero-ops.

The catch: the repo has gone quiet. The last release was v1.2.1 in late 2025 and there are hundreds of open issues. It still works, and CTranslate2 underneath it is stable, but nobody is shipping fixes right now. You also own the pipeline: chunking long files, retries, scaling workers, word-level timestamps. The per-minute APIs charge for exactly that boredom, and for being maintained.

Free vs Self-Hosted vs Paid

fully free

Free Tier (Everything)

The library and every Whisper model size (tiny through large-v3), MIT licensed. Quantization (int8) for smaller memory footprints, batching, VAD filtering, word timestamps. Nothing is gated.

Self-Hosted Setup

pip install faster-whisper, models download on first run. CPU handles small/medium models for casual use; a GPU with 8GB+ VRAM runs large-v3 faster than realtime. The engineering is around the library: job queues, chunking long audio, serving it to your team. Expect a day to a working internal transcription endpoint.

Paid Tier

None. The commercial comparison is hosted speech-to-text APIs: AssemblyAI charges $0.15/hr for Universal-2 and $0.21/hr for Universal-3.5 Pro on pre-recorded audio, and Deepgram's Nova-3 runs about $0.25-0.46/hr depending on plan and current promotional pricing.

The Math

At 100 hours of audio a month, hosted APIs run roughly $15-46/mo, cheap enough to not bother self-hosting. At thousands of hours, a $150-300/mo GPU instance (or spot capacity) transcribes it for a fraction of the API bill. The crossover is volume plus privacy: audio that can't leave your infrastructure makes the decision for you.

Verdict

Free and the de facto standard for self-hosted Whisper. Buy the API for streaming and diarization; run this for bulk and private transcription.

Free MIT-licensed Whisper inference. Beats per-minute APIs on cost at volume or when audio can't leave your infra.

What to do by team size

Solo
Runs on your own GPU or CPU for free, use it
Small team
An afternoon of setup replaces a per-minute API bill
Medium team
Great for bulk/private transcription, keep an API for streaming
Large team
The standard engine for in-house transcription pipelines
Self-hosting ops:moderate

Get tools like this every Wednesday

One featured tool, three on the radar. No fluff.

Similar Tools

Score
90/100 · A+
Adoption27/30
Maintenance25/25
Community13/20
License15/15
Analysis10/10

A low score is not a verdict on quality. Young and niche tools start low by design. How we calculate scores

Trust Signals

High adoption: 24,063 starsActive community: 1,963 forksCommunity discussions enabledOrganization account (7 public repos)

License: MIT License

Use freely, including commercial. Just keep the license.

Commercial use: ✓ Yes

About

Owner
SYSTRAN (Organization)
Stars
25,709
Forks
2,109

Explore Further

More tools in the directory