You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure批量转录报错:下载录音URI时冲突,无法获取转录结果

问题

我正在尝试使用Microsoft Azure的批量转录功能,按照官方文档操作:

  1. 通过POST请求将音频发送至服务;
  2. 利用POST返回参数发起GET请求获取结果,但未得到有效内容。

具体代码如下:

import json
import requests

url_endpoint = "https://westeurope.api.cognitive.microsoft.com/speechtotext/v3.0/transcriptions/"
payload = json.dumps({
  "contentUrls": "<我的audio.mp3的URL>",
  "properties": {
    "wordLevelTimestampsEnabled": True
  },
  "locale": "en-US",
  "displayName": "Testing if works"
})

headers = {
  'Ocp-Apim-Subscription-Key': '我的订阅密钥',
  'Content-Type': 'application/json'
}

response = requests.request("POST", url_endpoint, headers=headers, data=payload)

通过POST响应构建GET请求:

response2 = requests.request("GET", json.loads(response.text)["links"]["files"], headers=headers, data={})

response2返回结果中包含contentUrl,手动访问该链接时出现错误:"errorMessage": "Error when downloading the recording URI. StatusCode: Conflict"。请问我哪里操作有误?如何获取转录结果?


问题分析与解决步骤

1. 核心错误点

  • contentUrls格式不符合要求:API要求该字段为字符串数组,你直接传入单个URL会导致服务解析异常;
  • 音频URL权限冲突:你提供的音频URL可能存在访问权限问题(比如Azure Blob未生成带读取权限的SAS链接、URL过期或权限不足),导致转录服务无法下载音频;
  • 请求时机错误:批量转录是异步任务,POST仅创建任务,你可能在任务未完成时就发起GET请求,此时返回的contentUrl无效。

2. 修正操作

(1)修正contentUrls格式

将payload中的contentUrls改为数组形式:

payload = json.dumps({
  "contentUrls": ["<我的audio.mp3的URL>"],  # 改为数组格式
  "properties": {
    "wordLevelTimestampsEnabled": True
  },
  "locale": "en-US",
  "displayName": "Testing if works"
})

(2)确保音频URL可被服务访问

如果音频存储在Azure Blob:

  • 生成带有读取权限的SAS链接,有效期覆盖转录任务处理时间;
  • 生产环境不建议将容器设为公开读取。
    如果是其他存储服务,确保URL无防盗链、权限限制,可直接公开访问。

(3)等待任务完成后再获取结果

批量转录需要时间处理音频,正确流程是:

  1. 发送POST请求创建任务,提取任务ID;
  2. 定期轮询任务状态,直到状态变为Succeeded或Failed;
  3. 任务成功后再获取结果文件。

示例轮询代码:

import time

# 从POST响应提取任务状态查询链接
transcription_data = json.loads(response.text)
task_id = transcription_data["self"].split("/")[-1]
status_url = f"https://westeurope.api.cognitive.microsoft.com/speechtotext/v3.0/transcriptions/{task_id}"

# 轮询任务状态
while True:
    status_response = requests.get(status_url, headers=headers)
    status_data = status_response.json()
    if status_data["status"] in ["Succeeded", "Failed"]:
        break
    print(f"任务状态:{status_data['status']},等待10秒后重试...")
    time.sleep(10)

# 获取转录结果
if status_data["status"] == "Succeeded":
    files_url = status_data["links"]["files"]
    files_response = requests.get(files_url, headers=headers)
    files_data = files_response.json()
    # 筛选出转录结果文件
    result_file = next(item for item in files_data if item["kind"] == "Transcription")
    # 直接访问contentUrl(无需携带订阅密钥)
    result_content = requests.get(result_file["contentUrl"])
    print("转录结果:", result_content.text)
else:
    print("转录任务失败:", status_data.get("error", "未知错误"))

(4)正确访问contentUrl

获取到的contentUrl已包含临时访问权限,直接发起GET请求即可,不需要携带Ocp-Apim-Subscription-Key头。


内容的提问来源于stack exchange,提问作者Enrique Benito Casado

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 20:24:30