You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Google Speech v2中使用Class Tokens?及v1表现更优原因咨询

谷歌Speech-to-Text V2中Class Tokens的正确使用方法

问题背景

需要在谷歌Speech-to-Text V2版本中使用Class Tokens(如$TIME)实现语音转写,参考V1的实现逻辑能够成功将音频转写为格式化时间(如0720对应07:20),但自行编写的V2代码无法达到预期效果,且转写表现不如V1版本。

V1版本可正常运行的代码

from google.cloud import speech
import os

os.environ['GOOGLE_APPLICATION_CREDENTIALS']= 'auth.json'

with open("file.wav", "rb") as f:
    audio_content = f.read()
    audio = speech.RecognitionAudio(content=audio_content)

# 创建客户端实例
client = speech.SpeechClient()

config = speech.RecognitionConfig(
    encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16,
    language_code="sv-SE",
    speech_contexts=[  # 在此处添加Class Tokens
        speech.SpeechContext(
            phrases=["$TIME"]  # Class Token
        )
    ]
)

response = client.recognize(request={"config": config, "audio": audio})
for result in response.results:
    print(result.alternatives[0].transcript)

原V2版本问题代码

import os
from google.cloud.speech_v2 import SpeechClient
from google.cloud.speech_v2.types import cloud_speech

PROJECT_ID = "test"
os.environ['GOOGLE_APPLICATION_CREDENTIALS']= 'auth.json'

with open("file.wav", "rb") as f:
    audio_content = f.read()

client = SpeechClient()
config = cloud_speech.RecognitionConfig(
    auto_decoding_config=cloud_speech.AutoDetectDecodingConfig(),
    language_codes=["sv-SE"],
    model="short",
    adaptation=cloud_speech.SpeechAdaptation(
      phrase_sets=[
            cloud_speech.SpeechAdaptation.AdaptationPhraseSet(
                inline_phrase_set=cloud_speech.PhraseSet(phrases=[
                {
                    "value": "$TIME",
                    "boost": 20
                }
            ])
        )
      ]
    )
)

request = cloud_speech.RecognizeRequest(
    recognizer=f"projects/{PROJECT_ID}/locations/global/recognizers/_",
    config=config,
    content=audio_content,
)

# 转写音频为文本
response = client.recognize(request=request)
for result in response.results:
    print(result.alternatives[0].transcript)

核心问题分析

V2版本中Class Tokens的配置逻辑与V1存在本质差异:

  • V1直接将$TIME作为短语添加到speech_contexts中即可触发Class Token逻辑
  • V2需要通过class_reference字段明确指定Class Token,而非将其作为普通短语的value传入。原代码错误地将$TIME当作普通短语处理,导致API无法识别为特殊的Class Token,自然无法触发对应的格式化转写逻辑。

修正后的V2版本代码

import os
from google.cloud.speech_v2 import SpeechClient
from google.cloud.speech_v2.types import cloud_speech

PROJECT_ID = "test"
os.environ['GOOGLE_APPLICATION_CREDENTIALS']= 'auth.json'

with open("file.wav", "rb") as f:
    audio_content = f.read()

client = SpeechClient()
config = cloud_speech.RecognitionConfig(
    auto_decoding_config=cloud_speech.AutoDetectDecodingConfig(),
    language_codes=["sv-SE"],
    model="short",
    adaptation=cloud_speech.SpeechAdaptation(
        phrase_sets=[
            cloud_speech.SpeechAdaptation.AdaptationPhraseSet(
                inline_phrase_set=cloud_speech.PhraseSet(
                    phrases=[
                        {
                            "class_reference": {"class_id": "$TIME"},
                            "boost": 20
                        }
                    ]
                )
            )
        ]
    )
)

request = cloud_speech.RecognizeRequest(
    recognizer=f"projects/{PROJECT_ID}/locations/global/recognizers/_",
    config=config,
    content=audio_content,
)

response = client.recognize(request=request)
for result in response.results:
    print(result.alternatives[0].transcript)

额外优化建议

  • 模型选择:若short模型表现不佳,可尝试使用latest_short模型,该模型包含更优的识别逻辑与更新的Class Token适配规则
  • 音频编码配置:如果已知音频编码格式(如LINEAR16)和采样率,建议直接指定encoding和sample_rate_hertz参数,而非依赖自动检测,避免检测误差影响转写效果
  • Boost值调整:boost=20是合理的权重值,但可根据实际测试结果微调,过高可能导致无关语音被误识别为时间格式,过低则无法体现Class Token的优先级

内容的提问来源于stack exchange,提问作者user3218338

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 06:43:13