You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Google Speech-to-Text long_running_recognize时speaker_tag全为0的问题

问题:使用Google Speech-to-Text长时识别API时说话人标签全部为0

问题描述

我参考Stack Overflow相关回答实现说话人分离,因音频时长超过1分钟,改用long_running_recognize方法替代recognize,主要调整点包括:

  • 将音频文件上传至云端并获取文件URI
  • 使用speech.RecognitionAudio(uri=uri)替代RecognitionAudio(content=content)
  • 使用client.long_running_recognize(config=config, audio=audio)替代client.recognize(config=config, audio=audio)

代码可正常运行,但返回结果中所有词的speaker_tag均为0,输出示例如下:

word: 'Алло', speaker_tag: 0
word: 'здравствуйте', speaker_tag: 0
word: 'Я', speaker_tag: 0
word: 'хочу', speaker_tag: 0
word: 'котёнок', speaker_tag: 0
word: 'Ты', speaker_tag: 0
word: 'очень', speaker_tag: 0
word: 'классная', speaker_tag: 0
word: 'Спасибо', speaker_tag: 0
word: 'приятно', speaker_tag: 0
word: 'что', speaker_tag: 0
word: 'вы', speaker_tag: 0
word: 'и', speaker_tag: 0
word: 'Хорошего', speaker_tag: 0
word: 'вам', speaker_tag: 0
word: 'дня', speaker_tag: 0
word: 'сегодня', speaker_tag: 0
word: 'Спасибо', speaker_tag: 0
word: 'до', speaker_tag: 0
word: 'свидания', speaker_tag: 0

实现代码

from pathlib import Path
from google.cloud import speech_v1p1beta1 as speech
from google.cloud import storage

def file_upload(client, file: Path, bucket_name: str = 'wav_files_ua_eu_standard'):
    bucket = client.get_bucket(bucket_name)
    blob = bucket.blob(file.name)
    blob.upload_from_filename(file)
    uri3 = 'gs://' + blob.id[:-(len(str(blob.generation)) + 1)]
    print(F"{uri3=}")
    return uri3


client = speech.SpeechClient()
client_bucket = storage.Client(project='my-project-id-is-hidden')
speech_file_name = R"C:\Users\vasil\OneDrive\wav_samples\wav_sample_phone_call.wav"
speech_file = Path(speech_file_name)

if speech_file.exists:
    uri = file_upload(client_bucket, speech_file)

audio = speech.RecognitionAudio(uri=uri) 

diarization_config = speech.SpeakerDiarizationConfig(
    enable_speaker_diarization=True,
    min_speaker_count=2,
    max_speaker_count=3,
)

config = speech.RecognitionConfig(
    encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16,
    sample_rate_hertz=8000,
    language_code="ru-RU",
    diarization_config=diarization_config,
)

print("Waiting for operation to complete...")
response = client.long_running_recognize(config=config, audio=audio)
   
words_info = result.results

# Printing out the output:
for word_info in words_info[0].alternatives[0].words:
    print(f"word: '{word_info.word}', speaker_tag: {word_info.speaker_tag}")

问题分析与解决方案

核心错误点

  1. 未正确获取异步操作结果:long_running_recognize返回的是异步操作对象,必须调用.result()方法等待任务完成并获取最终识别结果。
  2. 错误的结果引用:代码中直接使用未定义的result变量,且错误地取了第一个识别结果——Google Speech-to-Text的说话人分离数据只会包含在最后一个识别结果中。

修正后的关键代码片段

print("Waiting for operation to complete...")
# 发起异步长时识别请求
operation = client.long_running_recognize(config=config, audio=audio)
# 等待操作完成,设置超时时间(示例为5分钟)
response = operation.result(timeout=300)

# 说话人分离结果仅存在于最后一个result中
final_result = response.results[-1]
words_info = final_result.alternatives[0].words

# 打印带说话人标签的结果
for word_info in words_info:
    print(f"word: '{word_info.word}', speaker_tag: {word_info.speaker_tag}")

额外注意事项

  • 确保timeout值设置合理,根据音频时长调整,避免因任务未完成触发超时错误。
  • 验证音频文件的采样率、编码格式是否与RecognitionConfig中的配置完全匹配,参数不匹配可能导致识别或分离异常。
  • 若仍有问题,可检查Google Cloud Storage中音频文件的权限,确保Speech-to-Text服务账号有权限读取该文件。

内容的提问来源于Stack Exchange,提问作者Vasyl Kolomiets

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 04:54:59