Python调用Azure TTS合成ar-SY阿拉伯语语音报latin-1编码错误
报错原因
该编码错误由Python requests库默认编码规则与阿拉伯语字符集不兼容导致:
- 直接向
requests.post()传入字符串类型的请求体时,requests默认使用latin-1(ISO-8859-1)编码处理请求内容 latin-1编码仅支持西欧拉丁系字符,无法识别阿拉伯语这类非拉丁字符,因此在编码阿拉伯语文本时抛出异常,报错提示也明确给出了修复方向:需要将请求体按UTF-8编码后再发送。
修复方法
仅需两处调整即可正常完成阿拉伯语语音合成:
- 修改请求头的
Content-Type字段,补充UTF-8字符集声明,明确告知Azure服务端请求体使用的编码格式 - 将拼接完成的SSML字符串调用
.encode('utf-8')转为UTF-8编码的字节流,再作为请求体传入POST请求
修改后的可运行代码如下:
import requests input_text="هذا المحتوى مجاني، لذلك لدعم القناة لمزيد من المحتوى المجاني، يرجى الاشتراك، مثل، مشاركة، تعليق" def generate_speech(self,language_id, input_text, outfile, token): url = "https://{}.tts.speech.microsoft.com/cognitiveservices/v1".format(self.azure_location) print("input_text:"+input_text) header = { 'Authorization': 'Bearer '+str(token), 'Content-Type': 'application/ssml+xml; charset=utf-8', 'X-Microsoft-OutputFormat': 'audio-24khz-160kbitrate-mono-mp3' } data = "<speak version='1.0' xml:lang='ar-SY'>\ <voice xml:lang='ar-SY' xml:gender='Male' name='ar-SY-LaithNeural'>\ {}\ </voice>\ </speak>".format(input_text) try: response = requests.post(url, headers=header, data=data.encode('utf-8')) response.raise_for_status() with open(outfile, "wb") as file: file.write(response.content) print("生成完成,响应状态码:", response.status_code) response.close() except Exception as e: print("ERROR: ", e)
注意:如果需要切换其他阿拉伯语区域的发音人,仅需修改SSML中xml:lang和name属性为对应语音的参数即可,编码逻辑无需调整。
内容的提问来源于stack exchange,提问作者Babel8 Business
相关产品推荐
相关产品推荐

