使用Uberi语音转文本代码调用GCP API遇版本错误求助
解决
googleapiclient.errors.UnknownApiNameOrVersion: name: speech version: v1beta1 错误 我之前碰到过完全一样的问题,来给你拆解下解决方案:
错误根源
你用的Uberi的SpeechRecognition库,内部调用的GCP语音转文本API版本是v1beta1——这个旧的测试版本早就被GCP弃用了,现在官方只支持稳定版的v1,所以才会抛出这个版本不匹配的错误。
快速修复现有代码(继续用SpeechRecognition库)
如果你不想换库,只需要修改SpeechRecognition的源代码就能搞定:
- 找到你Python环境里
speech_recognition包的安装路径(比如Lib/site-packages/speech_recognition/__init__.py,根据你的系统可能略有不同) - 打开
__init__.py,搜索v1beta1,找到类似这样的代码行:service = build("speech", "v1beta1", credentials=credentials) - 把
v1beta1改成v1,保存文件后重启你的脚本就可以正常调用了。
更靠谱的长期方案:直接用GCP官方客户端库
第三方库终究会有版本滞后的问题,直接用GCP官方的google-cloud-speech包才是更稳定的选择,这里给你一个完整的实时音频流识别实现:
第一步:安装依赖
先把需要的包装上:
pip install google-cloud-speech pyaudio
第二步:完整可运行代码
import pyaudio import threading import datetime from google.cloud import speech_v1 as speech from google.cloud.speech_v1 import types # 配置参数,根据你的需求修改 RATE = 16000 # 采样率,GCP推荐16000Hz CHUNK_SIZE = 1024 * 4 # 每次读取的音频块大小 CREDENTIALS_FILE = "your-service-account-key.json" # 替换成你的GCP凭证文件路径 def log_recognition_result(s): """记录识别结果到日志文件""" with open('recognition_log2.txt', 'a+', encoding='utf-8') as log_file: timestamp = datetime.datetime.now().strftime("[ %d-%b-%Y %H:%M:%S ] ") log_file.write(f"{timestamp}{s}\n") def start_streaming_recognition(): # 初始化GCP语音客户端 client = speech.SpeechClient.from_service_account_json(CREDENTIALS_FILE) # 配置识别参数 recognition_config = types.RecognitionConfig( encoding=types.RecognitionConfig.AudioEncoding.LINEAR16, sample_rate_hertz=RATE, language_code="zh-CN", # 换成你需要的语言,比如"en-US" enable_automatic_punctuation=True # 自动添加标点 ) streaming_config = types.StreamingRecognitionConfig( config=recognition_config, interim_results=False # 如果要实时看中间识别结果,改成True ) # 生成音频流的生成器函数 def audio_stream_generator(): audio_interface = pyaudio.PyAudio() stream = audio_interface.open( format=pyaudio.paInt16, channels=1, rate=RATE, input=True, frames_per_buffer=CHUNK_SIZE ) try: while True: chunk = stream.read(CHUNK_SIZE) yield types.StreamingRecognizeRequest(audio_content=chunk) finally: # 清理资源 stream.stop_stream() stream.close() audio_interface.terminate() # 发送流式请求并处理返回结果 responses = client.streaming_recognize(streaming_config, audio_stream_generator()) for response in responses: if not response.results: continue top_result = response.results[0] if not top_result.alternatives: continue transcript = top_result.alternatives[0].transcript print(f"识别到: {transcript}") log_recognition_result(transcript) if __name__ == "__main__": # 启动识别线程 recognition_thread = threading.Thread(target=start_streaming_recognition) recognition_thread.start() recognition_thread.join()
代码说明
- 用了GCP官方的
speech_v1稳定版API,不会出现版本兼容问题 - 支持实时流式识别,比你之前的缓冲分段方式更高效,延迟更低
- 凭证文件单独管理,比直接写在代码里更安全,也方便后续更换
- 可以灵活调整语言、是否返回中间结果、标点设置等参数
额外需要注意的点
- 确保你的GCP服务账号已经开通了Cloud Speech-to-Text API,并且给账号分配了
Speech User或者Speech Admin的角色 - 如果不想写死凭证文件路径,可以设置环境变量
GOOGLE_APPLICATION_CREDENTIALS指向你的凭证文件,代码里就不用指定路径了 - 要是遇到音频格式不兼容的问题,确认采样率、编码(LINEAR16)和声道数(单声道)和代码里的配置一致
内容的提问来源于stack exchange,提问作者Jlingz14
相关产品推荐
相关产品推荐

