ASR definition
Speech-to-text, also called automatic speech recognition (ASR), is technology that converts spoken audio into written text. Modern systems use deep learning models, from open models such as OpenAI's Whisper and NVIDIA's Parakeet to cloud services from OpenAI, Google, AWS, Azure and Deepgram, to transcribe calls, meetings, voice notes and commands in real time or from recordings, across many languages and accents.
How speech recognition works
Audio is first converted into a representation the model can process, typically a spectrogram showing how frequencies change over time. An acoustic model, today usually a transformer or conformer network, maps these features to sounds and words, and a language model component predicts likely word sequences to resolve ambiguity, such as their versus there. End-to-end models like Whisper learn the whole mapping from audio to text directly from large amounts of transcribed audio.
Post-processing then adds punctuation, capitalization and number formatting, turning twenty five dollars into $25, and optionally speaker diarization, which labels who spoke when. Timestamps for each word let applications link text back to the exact moment in the recording.
Real-time vs batch transcription
Batch transcription processes complete recordings, such as recorded calls, podcasts or meetings, and can use larger models and more context for higher accuracy. Real-time or streaming transcription returns partial results within a fraction of a second as someone speaks, which is essential for live captions, voice assistants and voice AI agents, where any delay makes a conversation feel unnatural.
Streaming systems must balance speed and accuracy: early partial words may be revised as more audio arrives. They also need endpointing, detecting when a speaker has finished, which is surprisingly hard in noisy environments or with people who pause mid-sentence.
Accuracy, languages and accents
Accuracy is usually measured with word error rate (WER): the share of words substituted, inserted or deleted compared with a human transcript. It varies widely with audio quality, background noise, microphone distance, accents, crosstalk and vocabulary. Phone audio is harder than studio audio, and domain terms such as drug names or product codes are common failure points.
Improve results with custom vocabulary or phrase boosting, models tuned for your domain or language, better audio capture, and testing on recordings from your real users. Code-switching, such as Hindi mixed with English or Arabic with English, needs models and evaluation designed for it, since general benchmarks rarely reflect it.
Privacy shapes the architecture too. Recordings of calls and consultations are sensitive, so many teams prefer providers with regional hosting and no data retention, or self-host open models, and redact card numbers and other sensitive details from transcripts automatically.
Tools and business uses
Options range from open models you host yourself to managed APIs that add streaming, diarization and domain-specific models, usually priced per minute of audio. Common choices today, and the business uses they typically power across industries, include:
- Open models: Whisper and its faster variants, plus newer models such as NVIDIA Parakeet and Canary, for self-hosted transcription
- Cloud APIs: Google Cloud Speech-to-Text, Amazon Transcribe, Azure AI Speech, Deepgram and AssemblyAI
- Contact centers: transcribing and analyzing calls for quality, compliance and sentiment
- Meetings and media: notes, summaries, captions and searchable archives
- Healthcare: clinical documentation from doctor and patient conversations
- Voice interfaces: commands, dictation and voice agents