Linux Web应用安装FFmpeg后Azure语音SDK异常求助
Azure Cognitive Services Speech SDK与FFmpeg在Docker容器中冲突的问题排查与解决
问题场景
我正在构建一款Web应用,流程为接收任意格式音频,通过FFmpeg转换为.wav文件后,调用azure.cognitiveservices.speech的speechsdk进行转写,采用Docker容器部署。
遇到的异常情况:
- 在Linux Web应用容器中安装FFmpeg后,speechsdk无法正常工作;未安装FFmpeg时,或本地构建运行容器时,speechsdk均可正常运行。
- 已添加调试打印语句,能看到SpeechRecognizer类初始化,但缓冲逻辑与本地运行不同,程序无原因直接停止。
Dockerfile代码
# Use an official Python runtime as a parent image FROM python:3.11-slim # Version Run RUN echo "Version Run 1..." # Install ffmpeg RUN apt-get update && apt-get install -y ffmpeg && \ # Ensure ffmpeg is executable chmod a+rx /usr/bin/ffmpeg && \ # Clean up the apt cache by removing /var/lib/apt/lists saves space apt-get clean && rm -rf /var/lib/apt/lists/* # Set the working directory in the container WORKDIR /app # Copy the current directory contents into the container at /app COPY . /app # Install any needed packages specified in requirements.txt RUN pip install --no-cache-dir -r requirements.txt # Make port 80 available to the world outside this container EXPOSE 8000 # Define environment variable ENV NAME World # Run main.py when the container launches CMD ["streamlit", "run", "main.py", "--server.port", "8000", "--server.address", "0.0.0.0"]
Python转写代码
def transcribe_audio_continuous_old(temp_dir, audio_file, language): speech_key = azure_speech_key service_region = azure_speech_region time.sleep(5) print(f"DEBUG TIME BEFORE speechconfig") ran = generate_random_string(length=5) temp_file = f"transcript_key_{ran}.txt" output_text_file = os.path.join(temp_dir, temp_file) speech_recognition_language = set_language_to_speech_code(language) speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region) speech_config.speech_recognition_language = speech_recognition_language audio_input = speechsdk.AudioConfig(filename=os.path.join(temp_dir, audio_file)) speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_input, language=speech_recognition_language) done = False transcript_contents = "" time.sleep(5) print(f"DEBUG TIME AFTER speechconfig") print(f"DEBUG FIle about to be passed {audio_file}") try: with open(output_text_file, "w", encoding=encoding) as file: def recognized_callback(evt): print("Start continuous recognition callback.") print(f"Recognized: {evt.result.text}") file.write(evt.result.text + "\n") nonlocal transcript_contents transcript_contents += evt.result.text + "\n" def stop_cb(evt): print("Stopping continuous recognition callback.") print(f"Event type: {evt}") speech_recognizer.stop_continuous_recognition() nonlocal done done = True def canceled_cb(evt): print(f"Recognition canceled: {evt.reason}") if evt.reason == speechsdk.CancellationReason.Error: print(f"Cancellation error: {evt.error_details}") nonlocal done done = True speech_recognizer.recognized.connect(recognized_callback) speech_recognizer.session_stopped.connect(stop_cb) speech_recognizer.canceled.connect(canceled_cb) speech_recognizer.start_continuous_recognition() while not done: time.sleep(1) print("DEBUG LOOPING TRANSCRIPT") except Exception as e: print(f"An error occurred: {e}") print("DEBUG DONE TRANSCRIPT") return temp_file, transcript_contents
可能的冲突原因
- 系统依赖库覆盖:FFmpeg安装时会更新或安装音频相关系统库(如ALSA、libasound2),这些库可能与Speech SDK依赖的底层音频库版本不兼容,导致SDK静默崩溃。
- 动态链接库路径冲突:FFmpeg安装后修改了系统动态链接库路径,使得Speech SDK加载了错误版本的依赖库,引发运行异常。
解决方案
方案1:多阶段构建隔离FFmpeg环境
将FFmpeg安装与应用运行环境分离,仅复制FFmpeg二进制文件到应用镜像,避免引入多余依赖:
# 第一阶段:构建FFmpeg环境 FROM debian:bookworm-slim as ffmpeg-build RUN apt-get update && apt-get install -y ffmpeg && \ apt-get clean && rm -rf /var/lib/apt/lists/* # 第二阶段:应用运行环境 FROM python:3.11-slim WORKDIR /app # 从第一阶段复制FFmpeg二进制文件和依赖库 COPY --from=ffmpeg-build /usr/bin/ffmpeg /usr/bin/ffmpeg COPY --from=ffmpeg-build /usr/lib/x86_64-linux-gnu/ /usr/lib/x86_64-linux-gnu/ COPY . /app RUN pip install --no-cache-dir -r requirements.txt EXPOSE 8000 ENV NAME World CMD ["streamlit", "run", "main.py", "--server.port", "8000", "--server.address", "0.0.0.0"]
方案2:指定Speech SDK依赖库路径
设置环境变量,强制Speech SDK优先使用自身携带的依赖库:
在Dockerfile中添加:
# 根据实际Python版本和SDK安装路径调整 ENV LD_LIBRARY_PATH=/usr/local/lib/python3.11/site-packages/azure/cognitiveservices/speech/lib/x64:$LD_LIBRARY_PATH
方案3:检查并修复依赖版本
在容器中运行以下命令,排查Speech SDK依赖库的版本冲突:
ldd /usr/local/lib/python3.11/site-packages/azure/cognitiveservices/speech/lib/x64/libMicrosoft.CognitiveServices.Speech.core.so
若发现不兼容库,可手动安装指定版本,例如:
RUN apt-get update && apt-get install -y libasound2=1.2.8-1+b1 && \ apt-get clean && rm -rf /var/lib/apt/lists/*
方案4:调整音频配置
修改代码,明确指定音频输入模式,避免自动检测系统设备引发的冲突:
audio_input = speechsdk.AudioConfig(filename=os.path.join(temp_dir, audio_file), use_default_device=False)
调试建议
- 开启核心转储,分析程序崩溃原因:
ulimit -c unlimited streamlit run main.py --server.port 8000 --server.address 0.0.0.0 gdb python core.* - 增加系统环境日志,辅助排查:
import subprocess print("LD_LIBRARY_PATH:", os.environ.get("LD_LIBRARY_PATH")) print("FFmpeg version:", subprocess.check_output(["ffmpeg", "-version"]).decode()) print("Speech SDK version:", speechsdk.__version__)
内容的提问来源于stack exchange,提问作者Kakobo kakobo
相关产品推荐
相关产品推荐

