Key Features

Provides controllable verbatim and intended transcription modes.
Supports multilingual speech recognition across many Whisper-supported languages.
Captures fillers, repetitions, stutters, false starts, and vocal sounds.
Produces precise word-level timing for alignment-sensitive workflows.
Handles seamless longform transcription.
Includes Verbatimize workflows for improving legacy speech datasets.
Links to Nyra Labs' public CrisperWhisper repository and model resources.
Targets production-ready inference for real speech data.

The system distinguishes between verbatim transcripts that preserve fillers, repetitions, cut-offs, and vocal sounds, and intended transcripts that produce cleaner readable text. It also improves alignment by supervising Whisper-style cross-attention behavior to produce more precise word timings.


CrisperWhisper 2.0 is useful for clinical documentation, conversation analysis, speech datasets, subtitles, podcast workflows, and expressive speech synthesis pipelines. Its controllable style switch helps teams choose whether fidelity to spoken disfluency or readability is more important for a task.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!