Embodied AI Glossary中文

Automatic Speech Recognition

语音识别ASRAdvanced

Automatically converts spoken words into text — the first step in a robot understanding a spoken command.

Automatic speech recognition converts a speech signal into text, also called speech-to-text (STT). The standard evaluation metric is word error rate (WER): the number of substitutions plus deletions plus insertions, divided by the number of reference words; for Chinese, this is usually computed per character and called character error rate (CER). The leading recent model is OpenAI’s Whisper, released in 2022, trained on 680,000 hours of weakly supervised, multilingual audio; it approaches the performance of supervised methods on several standard benchmarks without fine-tuning, and both the model and the inference code are open source. In embodied AI systems, speech recognition is usually the entry point of the interaction pipeline: a microphone array picks up audio, noise reduction cleans it up, ASR converts it to text, that text goes to a large language model or a VLA to generate a plan and actions, and finally text-to-speech (TTS) is used to reply. Some natively multimodal models take audio directly as input without converting to text first. On robots, ASR also has to cope with motor noise and far-field pickup.

ExampleA user says “hand me the red cup on the table”; an ASR model like Whisper converts it to text first, which is then passed to a VLA model as a language instruction to execute.

Also called
ASR, Speech-to-Text, STT
Related
Microphone Array · Large Language Model · Instruction Following · Human-Robot Interaction · Native Multimodal · Vision-Language-Action Model
Sources
Robust Speech Recognition via Large-Scale Weak Supervision (Whisper, arXiv:2212.04356)
Wikipedia: Speech recognition

See it in the full glossary →