You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python实现类谷歌的实时流式快速语音识别?

实现基于Google语音识别的实时流式语音转文本

你可以利用Google Cloud Speech-to-Text的流式识别API来实现边说话边显示识别内容的效果,该服务提供每月60分钟的免费额度,足够个人长期使用,且识别准确率和你之前用的recognize_google一致(底层均为Google的识别引擎)。

实现步骤与代码

  1. 安装依赖库:
pip install google-cloud-speech pyaudio
  1. 前往Google Cloud控制台创建项目,启用Speech-to-Text API,下载服务账号密钥JSON文件,并设置环境变量:
# Windows系统
set GOOGLE_APPLICATION_CREDENTIALS=path\to\your\key.json

# Linux/macOS系统
export GOOGLE_APPLICATION_CREDENTIALS="path/to/your/key.json"
  1. 实时流式识别代码:
import pyaudio
from google.cloud import speech_v1p1beta1 as speech

# 匹配麦克风输入的音频参数
RATE = 16000
CHUNK = int(RATE / 10)  # 按100ms的切片发送音频

def listen_stream():
    client = speech.SpeechClient()

    # 基础识别配置
    config = speech.RecognitionConfig(
        encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16,
        sample_rate_hertz=RATE,
        language_code="en-in",
        enable_automatic_punctuation=True,
    )
    # 流式识别配置,启用中间结果返回
    streaming_config = speech.StreamingRecognitionConfig(
        config=config,
        interim_results=True
    )

    # 初始化麦克风音频流
    p = pyaudio.PyAudio()
    stream = p.open(format=pyaudio.paInt16,
                    channels=1,
                    rate=RATE,
                    input=True,
                    frames_per_buffer=CHUNK)

    print("Listening... (实时识别中,按Ctrl+C停止)")

    # 生成音频块的生成器,持续向API发送数据
    def audio_generator():
        while True:
            data = stream.read(CHUNK)
            yield speech.StreamingRecognizeRequest(audio_content=data)

    # 处理流式识别响应
    responses = client.streaming_recognize(streaming_config, audio_generator())

    try:
        for response in responses:
            if not response.results:
                continue
            result = response.results[0]
            if not result.alternatives:
                continue
            # 获取最新的识别文本(包括未确定的中间结果)
            transcript = result.alternatives[0].transcript
            # 单行刷新显示,模拟实时更新效果
            print(f"\rYou Said: {transcript}", end="")
    except KeyboardInterrupt:
        print("\nStopped listening.")
    finally:
        stream.stop_stream()
        stream.close()
        p.terminate()

if __name__ == "__main__":
    listen_stream()

关键说明

  • interim_results=True:这是实现边说话边显示的核心配置,API会实时返回尚未完全确定的识别文本。
  • 按100ms切片发送音频,平衡延迟与识别准确率,保证响应速度。
  • 用\r和end=""实现单行内容刷新,和谷歌语音识别的实时显示逻辑一致。

内容的提问来源于stack exchange,提问作者Aryav Verma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 15:20:32