You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

FastAPI上传.docx文件传入python-docx Document报错求助

解决FastAPI上传docx文件后python-docx处理的TypeError问题

问题原因

你遇到的TypeError: stat: path should be string, bytes, os.PathLike or integer, not BytesIO错误,本质是python-docx 0.8.11版本在处理BytesIO对象时,会优先尝试将其当作文件路径解析,而你的BytesIO对象不具备路径解析所需的属性,同时直接读取文件字节后包装的BytesIO指针默认停在末尾,也会导致python-docx无法读取有效内容。

正确实现代码

修正后的代码需要重置BytesIO指针到起始位置,同时补充缺失的导入与模型定义:

from docx import Document
from io import BytesIO
from http import HTTPStatus
from typing import Any, Dict

from fastapi import FastAPI, UploadFile
from fastapi.responses import JSONResponse
from pydantic import BaseModel

# 定义错误响应模型
class Error(BaseModel):
    detail: str

app = FastAPI()

@app.post("/upload", responses={HTTPStatus.BAD_REQUEST: {"model": Error}})
async def upload_document(file: UploadFile) -> Dict[str, Any]:
    # 读取文件字节内容
    file_content = await file.read()
    # 将字节包装为BytesIO并重置指针到起始位置
    doc_buffer = BytesIO(file_content)
    doc_buffer.seek(0)
    # 初始化Document对象
    doc = Document(doc_buffer)
    
    # 示例:获取文档段落数量(可替换为你的业务逻辑)
    paragraph_count = len(doc.paragraphs)
    
    return JSONResponse(
        status_code=HTTPStatus.CREATED,
        content={
            "file_name": file.filename,
            "paragraph_count": paragraph_count
        },
    )

其他写法的错误分析

  • 写法1:doc = Document(BytesIO(await file.read()).read())
    BytesIO.read()返回的是原始字节数据,而Document构造函数要求传入支持seek、read方法的类文件对象,直接传字节会触发AttributeError: 'bytes' object has no attribute 'seek'。

  • 写法2:doc = Document(base64.decodebytes(BytesIO(await file.read())))
    base64.decodebytes()需要传入字节数据,你传入的BytesIO对象类型不匹配,导致TypeError: expected bytes-like object, not BytesIO;且上传的docx文件本身不是base64编码格式,这一步完全多余。

额外建议

如果环境允许,建议将python-docx升级到0.10.x及以上版本,新版本对类文件对象的支持更完善,无需手动调用seek(0)即可正常处理BytesIO对象。

内容的提问来源于stack exchange,提问作者Wojciech Ziarnik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 05:27:22