You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Azure认知服务语音转文本的转换结果保存至JSON文件?

解决Azure语音转文本结果保存至JSON文件的问题

嘿,我来帮你搞定把Azure语音转文本结果保存到JSON文件的问题!你的现有代码已经能完成语音识别,但缺少收集结果并写入文件的逻辑,我来一步步帮你调整:

核心思路说明

我们需要做这几件事:

  • 导入Python内置的json模块处理JSON文件操作
  • 创建一个列表来收集所有识别结果(因为连续识别会多次触发识别事件)
  • 修改识别事件的回调函数,把每次的结果存入列表
  • 在识别结束后,将列表内容写入JSON文件

修改后的完整代码

import azure.cognitiveservices.speech as speechsdk
import time
import json

# 替换为你的订阅密钥和服务区域
speech_key, service_region = "speech_key", "region"
speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region, speech_recognition_language="it-IT")
speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config)

# 创建列表存储所有识别结果,每个结果用字典保存更丰富的信息
recognized_results = []

def handle_recognized(evt):
    # 整理识别结果为字典,可按需添加更多元数据(比如时间戳、识别状态)
    result_item = {
        "recognized_text": evt.result.text,
        "recognition_status": evt.result.reason.name,
        "record_time": time.strftime("%Y-%m-%d %H:%M:%S")
    }
    recognized_results.append(result_item)
    print(f"\n识别到内容: {evt.result.text}")

# 绑定事件回调
speech_recognizer.session_started.connect(lambda evt: print('SESSION STARTED: {}'.format(evt)))
speech_recognizer.session_stopped.connect(lambda evt: print('\nSESSION STOPPED {}'.format(evt)))
speech_recognizer.recognized.connect(handle_recognized)

print('请说话(先触发一次单次识别,再进入10秒连续识别)\n\n')
# 处理单次识别的结果(按需保留,不需要可删除这部分)
single_recognition_result = speech_recognizer.recognize_once_async().get()
if single_recognition_result.reason == speechsdk.ResultReason.RecognizedSpeech:
    single_result_item = {
        "recognized_text": single_recognition_result.text,
        "recognition_status": single_recognition_result.reason.name,
        "record_time": time.strftime("%Y-%m-%d %H:%M:%S")
    }
    recognized_results.append(single_result_item)
    print(f"单次识别结果: {single_recognition_result.text}")
elif single_recognition_result.reason == speechsdk.ResultReason.NoMatch:
    print(f"未识别到有效语音: {single_recognition_result.no_match_details}")

# 启动连续识别
speech_recognizer.start_continuous_recognition()
print("进入10秒连续监听模式...")
time.sleep(10)
speech_recognizer.stop_continuous_recognition()

# 将结果写入JSON文件
output_filename = "speech_records.json"
with open(output_filename, "w", encoding="utf-8") as file:
    # ensure_ascii=False 确保意大利语特殊字符(如à、è)正常保存,indent=4让JSON格式更易读
    json.dump(recognized_results, file, ensure_ascii=False, indent=4)

print(f"所有识别结果已成功保存到文件: {output_filename}")

# 断开事件连接
speech_recognizer.session_started.disconnect_all()
speech_recognizer.recognized.disconnect_all()
speech_recognizer.session_stopped.disconnect_all()

关键细节解释

  1. 结果收集逻辑:用recognized_results列表存储每次识别的结果,每个结果封装成字典,包含识别文本、状态和时间戳,让JSON结构更清晰,方便后续解析。
  2. 事件回调优化:把原来的匿名lambda替换成命名函数handle_recognized,更便于扩展逻辑(比如增加错误处理)。
  3. JSON写入设置:
    • encoding="utf-8"确保多语言字符正常保存
    • ensure_ascii=False解决非ASCII字符(比如意大利语的特殊字母)被转义的问题
    • indent=4让生成的JSON文件有缩进,可读性更强
  4. 错误处理:对单次识别结果增加了状态判断,避免未识别到语音时出现异常。

运行代码后,你会在当前目录下看到speech_records.json文件,里面就是所有语音转文本的结果啦。

内容的提问来源于stack exchange,提问作者hojin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 12:12:31