You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Dash无本地目录读写权限时流式传文件到python-pptx的问题

问题根因

dcc.Upload 组件返回的contents值并非原始文件二进制流,而是带MIME声明前缀的base64数据URI,格式为:
data:application/vnd.openxmlformats-officedocument.presentationml.presentation;base64,<文件base64编码内容>
直接将该字符串传入Presentation()时,python-pptx会将其识别为本地文件路径尝试读取,必然触发PackageNotFoundError。

修复方案

全程基于内存流处理,无任何本地磁盘读写操作,完全满足部署约束:

  1. 拆分数据URI,剔除前缀,提取纯base64编码段
  2. 将base64段解码为原始二进制字节,封装为BytesIO内存流传入Presentation()即可

首先补充缺失的依赖导入:

import base64
from io import BytesIO

替换原有回调函数为以下实现:

@app.callback(
    Output("shape-list", "children"),
    [Input("upload-data", "filename"), Input("upload-data", "contents")],
)
def update_output(uploaded_filenames, uploaded_file_contents):
    shape_text = []
    if uploaded_filenames is not None and uploaded_file_contents is not None:
        for name, data in zip(uploaded_filenames, uploaded_file_contents):
            # 按第一个逗号拆分,分离前缀和base64内容
            _, base64_content = data.split(',', 1)
            # 解码为二进制字节,封装为内存字节流
            ppt_stream = BytesIO(base64.b64decode(base64_content))
            prs = Presentation(ppt_stream)
            # 遍历所有幻灯片、过滤无文本属性的形状(图片/图形等)避免报错
            for slide in prs.slides:
                shape_text.extend([
                    shape.text for shape in slide.shapes 
                    if hasattr(shape, "text")
                ])
    
    return [html.Li(txt) for txt in shape_text]
避坑说明
  • 必须使用BytesIO处理二进制流,StringIO仅适用于文本内容,无法解析pptx格式
  • 拆分字符串时必须指定maxsplit=1(即代码里split的第二个参数1),避免文件内容中存在逗号时拆分异常
  • 读取形状文本时增加hasattr(shape, "text")判断,跳过图片、自由形状、表格单元格容器等无直接text属性的元素,避免运行时报错

内容的提问来源于stack exchange,提问作者road_rash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.02 22:24:31