Skip to main content

Overview

Destined Voice includes tools to evaluate Speech-to-Text (STT) providers. Test multiple providers against your audio and analyze accuracy with WER/CER metrics.

Supported Providers

Transcribing Audio

Send audio to multiple providers:

Calculating Accuracy

Compare transcriptions against ground truth:

Metrics Explained

Word Error Rate (WER)

Measures word-level accuracy:
  • Lower is better (0.0 = perfect, 1.0 = completely wrong)
  • Industry standard for STT evaluation

Character Error Rate (CER)

Measures character-level accuracy:
  • More granular than WER
  • Useful for detecting minor transcription errors

Demographic Bias Analysis

Analyze STT accuracy across demographics:

Best Practices

Test with audio at 16kHz or higher. Lower quality affects all providers equally.
Remove punctuation, lowercase text, and expand numbers for fair WER calculation.
STT accuracy varies by accent. Test with speakers matching your user base.
Some providers trade accuracy for speed. Choose based on your use case.

Provider Comparison (Typical Performance)

Actual performance varies by audio quality, accent, and domain vocabulary.