You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

FastAPI处理音频文件时遇UnicodeDecodeError的排查与解决

FastAPI音频处理中的UnicodeDecodeError问题与解决方案

错误回溯

Traceback (most recent call last):
  File "C:\Users\sanja\AppData\Local\Programs\Python\Python310\lib\site-packages\uvicorn\protocols\http\httptools_impl.py", line 399, in run_asgi
    result = await app(  # type: ignore[func-returns-value]
  ...
  File "C:\Users\sanja\AppData\Local\Programs\Python\Python310\lib\site-packages\fastapi\encoders.py", line 303, in jsonable_encoder
    jsonable_encoder(
  File "C:\Users\sanja\AppData\Local\Programs\Python\Python310\lib\site-packages\fastapi\encoders.py", line 289, in jsonable_encoder
    encoded_value = jsonable_encoder(
  File "C:\Users\sanja\AppData\Local\Programs\Python\Python310\lib\site-packages\fastapi\encoders.py", line 318, in jsonable_encoder
    return ENCODERS_BY_TYPE[type(obj)](obj)
  File "C:\Users\sanja\AppData\Local\Programs\Python\Python310\lib\site-packages\fastapi\encoders.py", line 59, in <lambda>
    bytes: lambda o: o.decode(),
UnicodeDecodeError: 'utf-8' codec can't decode byte 0x9f in position 144: invalid start byte

代码示例

@app.post("/submit_response")
async def submit_response(session_id: str = Form(...), audio: UploadFile = File(...)):
    session = get_session(session_id)
    audio_path = f"{session_id}_response.wav"
    
    # Save the uploaded audio file
    with open(audio_path, "wb") as f:
        shutil.copyfileobj(audio.file, f)
    
    # Process the audio file here
    response_text = transcribe_audio(audio_path)  # Replace with actual transcription function

    # Update the current scenario conversation with the candidate's response
    session.current_scenario_conversation[-1] = (
        session.current_scenario_conversation[-1][0],
        session.current_scenario_conversation[-1][1],
        response_text
    )
    save_conversation_to_file(session.interview_filename, session.current_scenario_conversation[-1])

    # Retrieve the current trait and perform satisfaction check
    trait = TRAITS[session.current_trait_index]
    status, feedback = satisfaction_check(
        session.agents["satisfaction_check"],
        session.current_scenario_conversation[-1][1],
        response_text,
        trait['trait_name']
    )

    if status == "satisfied":
        # If satisfied, score the scenario and move to the next trait/scenario
        score = score_scenario(session.agents["scoring"], session.current_scenario_conversation, trait)
        save_conversation_to_file(session.interview_filename, ("Score", score))
        move_to_next_scenario(session)
        return {
            "message": "Moving to next scenario",
            "question": session.current_scenario_conversation[-1][1],
            "score": score  # Return the score
        }
    elif status == "insufficient":
        # If the response is insufficient, generate a follow-up question
        if len(session.current_scenario_conversation) >= 2:
            move_to_next_scenario(session)
            return {
                "message": "Moving to next scenario due to insufficient response",
                "question": session.current_scenario_conversation[-1][1]
            }
        else:
            follow_up_question = generate_follow_up(
                session.agents["follow_up"],
                session.candidate_name,
                session.current_scenario_conversation,
                len(session.current_scenario_conversation),
                insufficient=True
            )
            session.current_scenario_conversation.append(("Follow-Up", len(session.current_scenario_conversation), follow_up_question, ""))
            save_conversation_to_file(session.interview_filename, session.current_scenario_conversation[-1])
            return {
                "message": "Follow-up question for insufficient response",
                "question": follow_up_question
            }
    else:  # unsatisfied
        # Generate a follow-up question for unsatisfactory response
        follow_up_question = generate_follow_up(
            session.agents["follow_up"],
            session.candidate_name,
            session.current_scenario_conversation,
            len(session.current_scenario_conversation),
            insufficient=False
        )
        session.current_scenario_conversation.append(("Follow-Up", len(session.current_scenario_conversation), follow_up_question, ""))
        save_conversation_to_file(session.interview_filename, session.current_scenario_conversation[-1])

        if len(session.current_scenario_conversation) >= 3:
            move_to_next_scenario(session)
            return {
                "message": "Moving to next scenario after follow-up",
                "question": session.current_scenario_conversation[-1][1]
            }

        return {
            "message": "Follow-up question for unsatisfactory response",
            "question": follow_up_question
        }

def transcribe_audio(audio_path: str) -> str:
    """
    Transcribe audio file to text using speech_recognition.
    """
    recognizer = sr.Recognizer()
    try:
        with sr.AudioFile(audio_path) as source:
            audio_data = recognizer.record(source)
            text = recognizer.recognize_google(audio_data)
            return text
    except sr.UnknownValueError:
        return "Audio unintelligible"
    except sr.RequestError as e:
        return f"Could not request results; {e}"

背景信息

  • 使用FastAPI构建处理音频上传、语音转文字的API
  • 错误发生在应用尝试处理或编码音频处理返回的数据时
  • 音频文件以二进制格式上传和处理

技术问询

  1. 该场景下UnicodeDecodeError的成因是什么?
  2. 在FastAPI中处理二进制数据时如何解决此问题?
  3. 在FastAPI应用中集成语音转文字功能时,处理音频文件及其编码有哪些最佳实践?

解答

1. 错误成因

从回溯信息可知,错误出在FastAPI的jsonable_encoder尝试将bytes类型对象解码为UTF-8字符串时失败。核心原因是:返回的响应数据或session对象中包含未处理的二进制数据,FastAPI默认会对bytes类型调用.decode()方法转成字符串,但音频文件的二进制内容并非合法UTF-8编码,因此解码失败。

排查代码可知,大概率是session.current_scenario_conversation或返回值中不小心混入了二进制数据,最终在响应序列化阶段触发错误。

2. 解决方法

  • 清理会话数据:检查session.current_scenario_conversation及其他要返回的数据,确保所有元素都是字符串、数字等可JSON序列化类型,彻底移除二进制对象。比如仅保存音频文件路径到会话,而非二进制内容。
  • 自定义JSON编码器:若需处理二进制数据,可自定义编码器,对bytes类型用Base64编码替代直接解码:
    from fastapi.encoders import jsonable_encoder
    import base64
    
    def custom_json_encoder(obj):
        if isinstance(obj, bytes):
            return base64.b64encode(obj).decode("utf-8")
        return jsonable_encoder(obj)
    
    可在返回响应时手动调用该编码器,或在FastAPI实例初始化时指定json_encoder参数。
  • 提前过滤二进制数据:在数据存入会话或准备返回前,检查并过滤掉所有二进制类型的内容。

3. 最佳实践

  • 分离存储与处理:音频上传后仅保存文件路径到会话或数据库,不将二进制内容存入内存或会话对象,从根源避免序列化问题。
  • 验证音频格式:上传时通过UploadFile的content_type验证是否为合法音频格式(如audio/wav、audio/mpeg),提前过滤无效文件。
  • 增强错误处理:在语音转文字函数中补充更多异常捕获,比如文件损坏、格式不支持的情况,避免错误数据流入后续流程。
  • 异步处理音频:针对大体积音频,用异步任务(如Celery)处理转文字,避免阻塞API请求,同时返回任务ID让客户端轮询结果。
  • 清理临时文件:处理完音频后定时清理临时保存的音频文件,避免占用过多磁盘空间。

内容的提问来源于stack exchange,提问作者user26335862

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 20:57:02