如何让Azure语音合成SDK/API在Python容器中正常运行?
我正尝试让Azure语音合成SDK/API在Python容器中运行,在Python容器CLI中运行官方快速入门代码时,出现如下错误:
Speech synthesis canceled: CancellationReason.Error
Error details: USP error: timeout waiting for the first audio chunk
Did you set the speech resource key and region values?
已确认语音资源密钥和区域配置正确,本地主机CLI运行该代码完全正常,但容器里就触发上述错误。已知有OpenSSL限制问题,先后尝试了Ubuntu 22(内置OpenSSL 1.1.1.1)和Ubuntu 20容器,问题依旧。想知道这是容器端口问题还是Azure SDK/API的问题?
部署环境是Ubuntu 20虚拟机,用Docker Compose部署,相关配置文件如下:
Docker Compose配置
version: "3.8" services: # Define our individual services myapp: build: . container_name: app environment: - PYTHONUNBUFFERED=1 - PYTHONIOENCODING=UTF-8 expose: - 8888 networks: - my-network myserver: build: ./nginx container_name: myserver ports: - "80:80" - "443:443" - ./log/nginx:/var/log/nginx - ./nginx:/etc/nginx/conf.d networks: - my-network networks: my-network: driver: bridge
Dockerfile配置
# Use the Python3.10 container image FROM python:3.8.10 # Set the working directory to /app WORKDIR /app # Copy the current directory contents into the container at /app COPY . . ENV PYTHONFAULTHANDLER=1 \ PYTHONUNBUFFERED=1 \ PIP_NO_CACHE_DIR=off RUN apt update && apt install -y build-essential gcc ffmpeg libsm6 libxext6 wget nano unzip gdebi -y \ && pip install --upgrade pip \ && pip install -r requirements.txt RUN apt-get clean && rm -rf /var/lib/apt/lists/* # set display port to avoid crash ENV DISPLAY=:99 # Create a uwsgi log directory and files RUN mkdir /var/log/uwsgi RUN touch /var/log/uwsgi/uwsgi_access.log RUN touch /var/log/uwsgi/uwsgi_error.log # run the command to start uWSGI CMD ["uwsgi", "app.ini"]
Requirements.txt配置
-i https://pypi.python.org/simple uwsgi==2.0.22 werkzeug==2.3.6; python_version >= '3.8' azure-ai-textanalytics==5.2.0 azure-cognitiveservices-speech==1.31.0 azure-common==1.1.28 azure-core==1.28.0; python_version >= '3.7' flask==2.3.2 jinja2==3.1.2; python_version >= '3.7' requests==2.31.0; python_version >= '3.7' uwsgi==2.0.22
nginx.conf配置
server { listen 80; listen 443; server_name localhost; index index.php index.html index.htm index.nginx-debian.html; root /var/www/html; location / { try_files $uri $uri/ /index.php$is_args$args; } location ~ \.php$ { try_files $uri =404; fastcgi_split_path_info ^(.+\.php)(/.+)$; fastcgi_index index.php; include fastcgi_params; fastcgi_param SCRIPT_FILENAME $document_root$fastcgi_script_name; fastcgi_param PATH_INFO $fastcgi_path_info; fastcgi_read_timeout 300; } }
核心判断:大概率不是端口映射问题,而是容器网络或SDK依赖环境问题
排除端口问题:你的Docker Compose里
myapp只暴露了8888端口,但Azure语音合成SDK是主动向外发起HTTPS请求(默认用443端口),不需要容器对外暴露端口接收请求。本地能跑通说明密钥、区域没问题,容器里超时大概率是网络出不去或者环境缺失依赖。可能的原因及解决办法:
容器网络连通性问题:
进入容器内部,直接测试能否访问Azure语音服务的端点,比如执行curl https://<你的区域>.tts.speech.microsoft.com,如果连不通,说明容器网络有问题。检查Docker的bridge网络是否允许出站请求,或者虚拟机的防火墙是否拦截了容器的出站443流量。
也可以临时给myapp容器添加network_mode: host配置,如果能正常运行,说明自定义bridge网络存在配置问题。OpenSSL版本或SSL证书问题:
虽然你用了Ubuntu 20/22容器,但Python镜像自带的OpenSSL可能和系统的不一致。在容器里运行python -c "import ssl; print(ssl.OPENSSL_VERSION)"查看版本,Azure语音服务要求TLS 1.2+,如果版本过低会导致握手失败。
解决办法:升级容器内的OpenSSL,或者使用更新的Python镜像(比如python:3.10-slim-bookworm),新版本镜像默认支持更高的TLS版本。SDK依赖缺失:
Azure认知服务语音SDK依赖一些系统库,你的Dockerfile里安装了build-essential等,但可能缺了libssl-dev或libasound2(语音合成即使只生成音频文件也可能依赖音频基础库)。可以在apt install命令里添加libssl-dev libasound2试试。SDK版本问题:
你用的azure-cognitiveservices-speech==1.31.0是2023年的版本,尝试升级到最新稳定版(比如当前最新是1.40.0),新版本可能修复了容器环境下的网络超时问题。
额外排查步骤:
在容器里运行代码时,添加SDK的调试日志输出,能精准定位问题:import azure.cognitiveservices.speech as speechsdk speechsdk.set_log_level(speechsdk.LogLevel.Debug) # 后续合成代码不变查看详细日志,可判断是TLS握手失败、DNS解析问题还是其他网络错误。
内容的提问来源于stack exchange,提问作者S7bvwqX

