You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Watson语音转文本JSON DUMP中提取转录文本存入变量

从Watson Speech to Text结果提取转录文本,用于后续NLU服务

我帮你调整了代码,让你能轻松提取出转录文本并存成变量,方便传给Watson NLU。核心思路是在回调类里加一个变量来累积最终的转录内容,然后解析API返回的JSON结构,只提取确定的最终结果(避免中间的临时假设)。

首先看修改后的完整代码:

import json
from ibm_watson import SpeechToTextV1
from ibm_watson.websocket import RecognizeCallback, AudioSource
from os.path import join, dirname

class MyRecognizeCallback(RecognizeCallback):
    def __init__(self):
        super().__init__()  # 替代原来的初始化写法,更简洁规范
        self.full_transcript = ""  # 用来存储完整的最终转录文本

    def on_transcription(self, transcript):
        print(transcript)

    def on_connected(self):
        print('Connection was successful')

    def on_error(self, error):
        print('Error received: {}'.format(error))

    def on_inactivity_timeout(self, error):
        print('Inactivity timeout: {}'.format(error))

    def on_listening(self):
        print('Service is listening')

    def on_hypothesis(self, hypothesis):
        print(hypothesis)

    def on_data(self, data):
        # 解析返回的JSON数据
        results = data.get("results", [])
        for result in results:
            # 只处理标记为"final"的结果,这些是确定的转录内容
            if result.get("final"):
                alternatives = result.get("alternatives", [])
                if alternatives:
                    # 取置信度最高的第一个备选转录文本,拼接到完整内容里
                    self.full_transcript += alternatives[0].get("transcript", "") + " "
        # 可选:打印实时累积的转录内容
        print("当前已转录:", self.full_transcript.strip())

    def on_close(self):
        print("Connection closed")
        # 连接关闭时,full_transcript就是完整的转录文本了
        print("最终转录文本:", self.full_transcript.strip())

# 记得在这里初始化你的Speech to Text客户端(填入你的API密钥和服务地址)
# speech_to_text = SpeechToTextV1(
#     authenticator=IAMAuthenticator('你的API密钥')
# )
# speech_to_text.set_service_url('你的服务URL')

myRecognizeCallback = MyRecognizeCallback()
with open(join(dirname(__file__), './.', 'output.wav'), 'rb') as audio_file:
    audio_source = AudioSource(audio_file)
speech_to_text.recognize_using_websocket(
    audio=audio_source,
    content_type='audio/wav',
    recognize_callback=myRecognizeCallback,
    model='fr-FR_BroadbandModel')

# 连接关闭后,直接通过这个变量获取完整转录文本,传给NLU即可
final_transcript = myRecognizeCallback.full_transcript.strip()
print("准备传入NLU的文本:", final_transcript)
# 接下来就可以写调用Watson NLU的代码,用final_transcript作为输入

关键说明:

  • 我修复了你原来代码里的小语法错误(__init__方法少了右括号),并用更规范的super().__init__()初始化父类
  • 添加了self.full_transcript实例变量,用来累积所有最终的转录内容
  • 在on_data方法里,只处理final: true的结果——因为Speech to Text会返回中间的假设性结果(final: false),这些不需要存下来,只有final: true的是确定的转录
  • 从每个结果的alternatives[0]提取transcript,因为第一个备选是置信度最高的结果
  • 连接关闭后,myRecognizeCallback.full_transcript.strip()就是你需要的完整转录文本,直接传给NLU服务就行

内容的提问来源于stack exchange,提问作者Jeremy.l71

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:39:05