You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python结合Google Speech-to-Text API实现实时麦克风音频转文字?

要实现实时边录音边转文本的效果,你需要用到Google Speech-to-Text的**流式识别(Streaming Recognition)**功能——这和你现在用的单次录音识别逻辑完全不同,它能把麦克风的音频流持续发送给API,实时返回识别结果。下面是具体的实现方案:

第一步:安装必要依赖

首先得安装Google Cloud的语音客户端库和音频处理库:

pip install google-cloud-speech pyaudio

第二步:配置Google Cloud凭据

因为要调用Google的API,你需要先完成以下操作:

  1. 创建一个Google Cloud项目,启用Speech-to-Text API
  2. 下载服务账号密钥JSON文件
  3. 设置环境变量指向这个文件:
# Windows系统
set GOOGLE_APPLICATION_CREDENTIALS=path/to/your/service-account-key.json

# macOS/Linux系统
export GOOGLE_APPLICATION_CREDENTIALS=path/to/your/service-account-key.json

第三步:实时流式识别代码示例

下面的代码会持续监听麦克风,把音频分块发送给API,实时输出中间识别结果和最终结果:

import pyaudio
from google.cloud import speech_v1p1beta1 as speech
import sys

# 音频参数配置(必须和Google API要求的格式一致)
RATE = 16000
CHUNK = int(RATE / 10)  # 每次读取100ms的音频块

def listen_print_loop(responses):
    """处理API返回的流式响应,实现实时文本更新"""
    num_chars_printed = 0
    for response in responses:
        if not response.results:
            continue

        # 获取第一个识别结果(API可能返回多个候选)
        result = response.results[0]
        if not result.alternatives:
            continue

        # 提取识别出的文本
        transcript = result.alternatives[0].transcript

        # 清除之前的临时文本(中间结果会不断更新,需要覆盖显示)
        overwrite_chars = ' ' * (num_chars_printed - len(transcript))

        if not result.is_final:
            # 中间结果:实时在同一行更新显示
            sys.stdout.write(transcript + overwrite_chars + '\r')
            sys.stdout.flush()
            num_chars_printed = len(transcript)
        else:
            # 最终结果:换行显示完整文本
            print(transcript + overwrite_chars)
            num_chars_printed = 0

def main():
    # 初始化Speech客户端
    client = speech.SpeechClient()

    # 基础识别配置
    config = speech.RecognitionConfig(
        encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16,
        sample_rate_hertz=RATE,
        language_code='hi-IN',  # 印地语,可按需修改为其他语言(如'en-US')
        enable_automatic_punctuation=True,  # 自动添加标点符号
    )

    # 流式识别配置:开启中间结果返回
    streaming_config = speech.StreamingRecognitionConfig(
        config=config,
        interim_results=True,
    )

    # 初始化麦克风音频流
    p = pyaudio.PyAudio()
    stream = p.open(format=pyaudio.paInt16,
                    channels=1,
                    rate=RATE,
                    input=True,
                    frames_per_buffer=CHUNK)

    print("开始说话吧!按Ctrl+C停止识别...")

    # 生成持续读取麦克风数据的生成器
    audio_generator = (stream.read(CHUNK) for _ in iter(lambda: True, None))

    # 包装成流式请求
    requests = (speech.StreamingRecognizeRequest(audio_content=content)
                for content in audio_generator)

    # 发送请求并处理响应
    responses = client.streaming_recognize(streaming_config, requests)

    try:
        listen_print_loop(responses)
    except KeyboardInterrupt:
        print("\n已停止识别")
        stream.stop_stream()
        stream.close()
        p.terminate()

if __name__ == '__main__':
    main()

关键逻辑说明

  • 流式识别核心:interim_results=True 是实现实时更新的关键,它会让API返回还在动态更新的中间识别结果,而不是只返回最终的完整文本。
  • 音频流处理:用PyAudio持续读取麦克风的小块音频,包装成流式请求发送给API,避免一次性录制完整音频再识别。
  • 结果显示优化:通过覆盖同一行文本的方式,模拟实时字幕的效果,让用户看到动态更新的识别内容。

注意事项

  • 确保网络稳定,流式识别需要持续和Google API通信。
  • 如果不需要印地语,直接修改language_code参数即可(比如改成'en-US'支持英语)。
  • 第一次运行可能有几秒初始化延迟,之后就能实时响应语音输入。

内容的提问来源于stack exchange,提问作者Akhil Sahu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:36:25