如何在Google Speech v2中使用Class Tokens?及v1表现更优原因咨询
谷歌Speech-to-Text V2中Class Tokens的正确使用方法
问题背景
需要在谷歌Speech-to-Text V2版本中使用Class Tokens(如$TIME)实现语音转写,参考V1的实现逻辑能够成功将音频转写为格式化时间(如0720对应07:20),但自行编写的V2代码无法达到预期效果,且转写表现不如V1版本。
V1版本可正常运行的代码
from google.cloud import speech import os os.environ['GOOGLE_APPLICATION_CREDENTIALS']= 'auth.json' with open("file.wav", "rb") as f: audio_content = f.read() audio = speech.RecognitionAudio(content=audio_content) # 创建客户端实例 client = speech.SpeechClient() config = speech.RecognitionConfig( encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16, language_code="sv-SE", speech_contexts=[ # 在此处添加Class Tokens speech.SpeechContext( phrases=["$TIME"] # Class Token ) ] ) response = client.recognize(request={"config": config, "audio": audio}) for result in response.results: print(result.alternatives[0].transcript)
原V2版本问题代码
import os from google.cloud.speech_v2 import SpeechClient from google.cloud.speech_v2.types import cloud_speech PROJECT_ID = "test" os.environ['GOOGLE_APPLICATION_CREDENTIALS']= 'auth.json' with open("file.wav", "rb") as f: audio_content = f.read() client = SpeechClient() config = cloud_speech.RecognitionConfig( auto_decoding_config=cloud_speech.AutoDetectDecodingConfig(), language_codes=["sv-SE"], model="short", adaptation=cloud_speech.SpeechAdaptation( phrase_sets=[ cloud_speech.SpeechAdaptation.AdaptationPhraseSet( inline_phrase_set=cloud_speech.PhraseSet(phrases=[ { "value": "$TIME", "boost": 20 } ]) ) ] ) ) request = cloud_speech.RecognizeRequest( recognizer=f"projects/{PROJECT_ID}/locations/global/recognizers/_", config=config, content=audio_content, ) # 转写音频为文本 response = client.recognize(request=request) for result in response.results: print(result.alternatives[0].transcript)
核心问题分析
V2版本中Class Tokens的配置逻辑与V1存在本质差异:
- V1直接将
$TIME作为短语添加到speech_contexts中即可触发Class Token逻辑 - V2需要通过
class_reference字段明确指定Class Token,而非将其作为普通短语的value传入。原代码错误地将$TIME当作普通短语处理,导致API无法识别为特殊的Class Token,自然无法触发对应的格式化转写逻辑。
修正后的V2版本代码
import os from google.cloud.speech_v2 import SpeechClient from google.cloud.speech_v2.types import cloud_speech PROJECT_ID = "test" os.environ['GOOGLE_APPLICATION_CREDENTIALS']= 'auth.json' with open("file.wav", "rb") as f: audio_content = f.read() client = SpeechClient() config = cloud_speech.RecognitionConfig( auto_decoding_config=cloud_speech.AutoDetectDecodingConfig(), language_codes=["sv-SE"], model="short", adaptation=cloud_speech.SpeechAdaptation( phrase_sets=[ cloud_speech.SpeechAdaptation.AdaptationPhraseSet( inline_phrase_set=cloud_speech.PhraseSet( phrases=[ { "class_reference": {"class_id": "$TIME"}, "boost": 20 } ] ) ) ] ) ) request = cloud_speech.RecognizeRequest( recognizer=f"projects/{PROJECT_ID}/locations/global/recognizers/_", config=config, content=audio_content, ) response = client.recognize(request=request) for result in response.results: print(result.alternatives[0].transcript)
额外优化建议
- 模型选择:若
short模型表现不佳,可尝试使用latest_short模型,该模型包含更优的识别逻辑与更新的Class Token适配规则 - 音频编码配置:如果已知音频编码格式(如LINEAR16)和采样率,建议直接指定
encoding和sample_rate_hertz参数,而非依赖自动检测,避免检测误差影响转写效果 - Boost值调整:
boost=20是合理的权重值,但可根据实际测试结果微调,过高可能导致无关语音被误识别为时间格式,过低则无法体现Class Token的优先级
内容的提问来源于stack exchange,提问作者user3218338
相关产品推荐
相关产品推荐

