Speechmatics Python代码调用API时无法使用GCS签名URL的fetch_data参数
问题:使用Google Cloud Storage签名URL调用Speechmatics转录API返回400错误
我尝试用带签名URL的GCS文件调用Speechmatics转录API做测试,根据文档要求用fetch_data参数提供文件URL,但调用submit_job接口时始终失败,代码如下:
import requests import logging from django.conf import settings logger = logging.getLogger('speechmatics') class SpeechMatics(): def submit_file(audio_url, webhook_url, lang): # Define request parameters url = 'https://asr.api.speechmatics.com/v2/jobs' headers = { 'Authorization': 'Bearer ' + settings.SPEECHMATICS_API_KEY, 'Content-Type': 'multipart/form-data' } logger.info("Submitting transcription job to Speechmatics API...") logger.info("audio_url type: {}".format(type(audio_url))) logger.info("API URL: {}".format(url)) logger.info("API headers: {}".format(headers)) payload = { "type": "transcription", "transcription_config": { "language": lang, "diarization": "speaker", }, "fetch_data": { "url": audio_url } } logger.info("API payload: {}".format(payload)) # Send the request try: response = requests.post(url, headers=headers, data=payload) # response.raise_for_status() # raise HTTPError for non-2xx response status codes logger.info("Speechmatics API response: {}".format(response.json())) logger.info("API response status code: {}".format(response.status_code)) logger.info("API response headers: {}".format(response.headers)) except requests.exceptions.RequestException as e: logger.error("Speechmatics API request failed: {}".format(e)) raise # re-raise the exception to be handled at a higher level
执行后收到400错误,返回信息如下:
[speechmatics.py:18] API headers: {'Authorization': 'Bearer TOKEN', 'Content-Type': 'multipart/form-data'} [speechmatics.py:29] API payload: {'type': 'transcription', 'transcription_config': {'language': 'en', 'diarization': 'speaker'}, 'fetch_data': {'url': 'https://storage.googleapis.com/staging-videoo-storage-7/8bdcf5aa-a998-440a-870d-7c31de591aca?...'}} [speechmatics.py:36] Speechmatics API response: {'code': 400, 'message': 'no multipart boundary param in Content-Type'} [speechmatics.py:37] API response status code: 400 [speechmatics.py:38] API response headers: {'Content-Length': '68', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=15724800; includeSubDomains', 'Request-Id': '6126c67c56f12b5042b7e4f78b4632aa', 'Access-Control-Allow-Origin': '*', 'Access-Control-Allow-Credentials': 'true', 'Access-Control-Allow-Methods': 'GET, PUT, POST, DELETE, PATCH, OPTIONS', 'Access-Control-Allow-Headers': 'DNT,Keep-Alive,User-Agent,X-Requested-With,If-Modified-Since,Cache-Control,Content-Type,Range,Authorization', 'Access-Control-Max-Age': '1728000', 'X-Cache': 'CONFIG_NOCACHE', 'X-Azure-Ref': '0AWAjZAAAAAB5KJjEElURTa7s+0gal3wqTUFOMzBFREdFMDMwNwBhN2JjOWQ4MC02YjBiLTQ1NWEtYjE3MS01NGJkZmNiYWE0YTk=', 'Date': 'Tue, 28 Mar 2023 21:45:36 GMT'}
解决方案
错误根源是错误设置了Content-Type: multipart/form-data,同时用data参数传递JSON格式请求体。当使用fetch_data让Speechmatics主动拉取文件时,请求体是纯JSON格式,不需要multipart类型。
修改代码的两个核心点:
- 移除手动设置的
Content-Type头(或改为application/json),requests库在使用json参数时会自动添加正确的Content-Type - 将
requests.post中的data参数替换为json参数,库会自动把字典序列化为JSON字符串
修改后的代码:
import requests import logging from django.conf import settings logger = logging.getLogger('speechmatics') class SpeechMatics(): def submit_file(audio_url, webhook_url, lang): # Define request parameters url = 'https://asr.api.speechmatics.com/v2/jobs' headers = { 'Authorization': 'Bearer ' + settings.SPEECHMATICS_API_KEY, # 移除Content-Type,或手动设置为'Content-Type': 'application/json' } logger.info("Submitting transcription job to Speechmatics API...") logger.info("audio_url type: {}".format(type(audio_url))) logger.info("API URL: {}".format(url)) logger.info("API headers: {}".format(headers)) payload = { "type": "transcription", "transcription_config": { "language": lang, "diarization": "speaker", }, "fetch_data": { "url": audio_url } } logger.info("API payload: {}".format(payload)) # Send the request try: # 使用json参数替代data参数,自动处理JSON序列化和Content-Type response = requests.post(url, headers=headers, json=payload) # response.raise_for_status() # raise HTTPError for non-2xx response status codes logger.info("Speechmatics API response: {}".format(response.json())) logger.info("API response status code: {}".format(response.status_code)) logger.info("API response headers: {}".format(response.headers)) except requests.exceptions.RequestException as e: logger.error("Speechmatics API request failed: {}".format(e)) raise # re-raise the exception to be handled at a higher level
额外注意:
multipart/form-data仅用于直接上传本地文件到API的场景,使用fetch_data时,API会自行从你提供的URL拉取文件,请求体只需包含配置信息的JSON- 确保GCS签名URL具备读权限,且在Speechmatics API请求时仍处于有效期内
内容的提问来源于stack exchange,提问作者london_utku
相关产品推荐
相关产品推荐

