You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Linux Web应用安装FFmpeg后Azure语音SDK异常求助

Azure Cognitive Services Speech SDK与FFmpeg在Docker容器中冲突的问题排查与解决

问题场景

我正在构建一款Web应用,流程为接收任意格式音频,通过FFmpeg转换为.wav文件后,调用azure.cognitiveservices.speech的speechsdk进行转写,采用Docker容器部署。

遇到的异常情况:

  • 在Linux Web应用容器中安装FFmpeg后,speechsdk无法正常工作;未安装FFmpeg时,或本地构建运行容器时,speechsdk均可正常运行。
  • 已添加调试打印语句,能看到SpeechRecognizer类初始化,但缓冲逻辑与本地运行不同,程序无原因直接停止。

Dockerfile代码

# Use an official Python runtime as a parent image
FROM python:3.11-slim

# Version Run
RUN echo "Version Run 1..."

# Install ffmpeg
RUN apt-get update && apt-get install -y ffmpeg && \
    # Ensure ffmpeg is executable
    chmod a+rx /usr/bin/ffmpeg && \
    # Clean up the apt cache by removing /var/lib/apt/lists saves space
    apt-get clean && rm -rf /var/lib/apt/lists/*

# Set the working directory in the container
WORKDIR /app

# Copy the current directory contents into the container at /app
COPY . /app

# Install any needed packages specified in requirements.txt
RUN pip install --no-cache-dir -r requirements.txt

# Make port 80 available to the world outside this container
EXPOSE 8000

# Define environment variable
ENV NAME World

# Run main.py when the container launches
CMD ["streamlit", "run", "main.py", "--server.port", "8000", "--server.address", "0.0.0.0"]

Python转写代码

def transcribe_audio_continuous_old(temp_dir, audio_file, language):
    speech_key = azure_speech_key
    service_region = azure_speech_region

    time.sleep(5)
    print(f"DEBUG TIME BEFORE speechconfig")

    ran = generate_random_string(length=5)
    temp_file = f"transcript_key_{ran}.txt"
    output_text_file = os.path.join(temp_dir, temp_file)
    speech_recognition_language = set_language_to_speech_code(language)
    
    speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region)
    speech_config.speech_recognition_language = speech_recognition_language
    audio_input = speechsdk.AudioConfig(filename=os.path.join(temp_dir, audio_file))
        
    speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_input, language=speech_recognition_language)
    done = False
    transcript_contents = ""

    time.sleep(5)
    print(f"DEBUG TIME AFTER speechconfig")
    print(f"DEBUG FIle about to be passed {audio_file}")

    try:
        with open(output_text_file, "w", encoding=encoding) as file:
            def recognized_callback(evt):
                print("Start continuous recognition callback.")
                print(f"Recognized: {evt.result.text}")
                file.write(evt.result.text + "\n")
                nonlocal transcript_contents
                transcript_contents += evt.result.text + "\n"

            def stop_cb(evt):
                print("Stopping continuous recognition callback.")
                print(f"Event type: {evt}")
                speech_recognizer.stop_continuous_recognition()
                nonlocal done
                done = True
            
            def canceled_cb(evt):
                print(f"Recognition canceled: {evt.reason}")
                if evt.reason == speechsdk.CancellationReason.Error:
                    print(f"Cancellation error: {evt.error_details}")
                nonlocal done
                done = True

            speech_recognizer.recognized.connect(recognized_callback)
            speech_recognizer.session_stopped.connect(stop_cb)
            speech_recognizer.canceled.connect(canceled_cb)

            speech_recognizer.start_continuous_recognition()
            while not done:
                time.sleep(1)
                print("DEBUG LOOPING TRANSCRIPT")

    except Exception as e:
        print(f"An error occurred: {e}")

    print("DEBUG DONE TRANSCRIPT")

    return temp_file, transcript_contents

可能的冲突原因

  1. 系统依赖库覆盖:FFmpeg安装时会更新或安装音频相关系统库(如ALSA、libasound2),这些库可能与Speech SDK依赖的底层音频库版本不兼容,导致SDK静默崩溃。
  2. 动态链接库路径冲突:FFmpeg安装后修改了系统动态链接库路径,使得Speech SDK加载了错误版本的依赖库,引发运行异常。

解决方案

方案1:多阶段构建隔离FFmpeg环境

将FFmpeg安装与应用运行环境分离,仅复制FFmpeg二进制文件到应用镜像,避免引入多余依赖:

# 第一阶段:构建FFmpeg环境
FROM debian:bookworm-slim as ffmpeg-build
RUN apt-get update && apt-get install -y ffmpeg && \
    apt-get clean && rm -rf /var/lib/apt/lists/*

# 第二阶段:应用运行环境
FROM python:3.11-slim
WORKDIR /app

# 从第一阶段复制FFmpeg二进制文件和依赖库
COPY --from=ffmpeg-build /usr/bin/ffmpeg /usr/bin/ffmpeg
COPY --from=ffmpeg-build /usr/lib/x86_64-linux-gnu/ /usr/lib/x86_64-linux-gnu/

COPY . /app
RUN pip install --no-cache-dir -r requirements.txt

EXPOSE 8000
ENV NAME World
CMD ["streamlit", "run", "main.py", "--server.port", "8000", "--server.address", "0.0.0.0"]

方案2:指定Speech SDK依赖库路径

设置环境变量,强制Speech SDK优先使用自身携带的依赖库:
在Dockerfile中添加:

# 根据实际Python版本和SDK安装路径调整
ENV LD_LIBRARY_PATH=/usr/local/lib/python3.11/site-packages/azure/cognitiveservices/speech/lib/x64:$LD_LIBRARY_PATH

方案3:检查并修复依赖版本

在容器中运行以下命令,排查Speech SDK依赖库的版本冲突:

ldd /usr/local/lib/python3.11/site-packages/azure/cognitiveservices/speech/lib/x64/libMicrosoft.CognitiveServices.Speech.core.so

若发现不兼容库,可手动安装指定版本,例如:

RUN apt-get update && apt-get install -y libasound2=1.2.8-1+b1 && \
    apt-get clean && rm -rf /var/lib/apt/lists/*

方案4:调整音频配置

修改代码,明确指定音频输入模式,避免自动检测系统设备引发的冲突:

audio_input = speechsdk.AudioConfig(filename=os.path.join(temp_dir, audio_file), use_default_device=False)

调试建议

  • 开启核心转储,分析程序崩溃原因:
    ulimit -c unlimited
    streamlit run main.py --server.port 8000 --server.address 0.0.0.0
    gdb python core.*
    
  • 增加系统环境日志,辅助排查:
    import subprocess
    print("LD_LIBRARY_PATH:", os.environ.get("LD_LIBRARY_PATH"))
    print("FFmpeg version:", subprocess.check_output(["ffmpeg", "-version"]).decode())
    print("Speech SDK version:", speechsdk.__version__)
    

内容的提问来源于stack exchange,提问作者Kakobo kakobo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 21:13:27