MultimodalCross-modal capabilities: image, video and speech

语音识别(STT)

Technology that converts speech into text.

STT serves meeting transcription, subtitles and voice input; mainstream solutions support real-time streaming and speaker diarization, billed by audio duration.

Related terms