Dash无本地目录读写权限时流式传文件到python-pptx的问题
问题根因
dcc.Upload 组件返回的contents值并非原始文件二进制流,而是带MIME声明前缀的base64数据URI,格式为:data:application/vnd.openxmlformats-officedocument.presentationml.presentation;base64,<文件base64编码内容>
直接将该字符串传入Presentation()时,python-pptx会将其识别为本地文件路径尝试读取,必然触发PackageNotFoundError。
修复方案
全程基于内存流处理,无任何本地磁盘读写操作,完全满足部署约束:
- 拆分数据URI,剔除前缀,提取纯base64编码段
- 将base64段解码为原始二进制字节,封装为
BytesIO内存流传入Presentation()即可
首先补充缺失的依赖导入:
import base64 from io import BytesIO
替换原有回调函数为以下实现:
@app.callback( Output("shape-list", "children"), [Input("upload-data", "filename"), Input("upload-data", "contents")], ) def update_output(uploaded_filenames, uploaded_file_contents): shape_text = [] if uploaded_filenames is not None and uploaded_file_contents is not None: for name, data in zip(uploaded_filenames, uploaded_file_contents): # 按第一个逗号拆分,分离前缀和base64内容 _, base64_content = data.split(',', 1) # 解码为二进制字节,封装为内存字节流 ppt_stream = BytesIO(base64.b64decode(base64_content)) prs = Presentation(ppt_stream) # 遍历所有幻灯片、过滤无文本属性的形状(图片/图形等)避免报错 for slide in prs.slides: shape_text.extend([ shape.text for shape in slide.shapes if hasattr(shape, "text") ]) return [html.Li(txt) for txt in shape_text]
避坑说明
- 必须使用
BytesIO处理二进制流,StringIO仅适用于文本内容,无法解析pptx格式 - 拆分字符串时必须指定
maxsplit=1(即代码里split的第二个参数1),避免文件内容中存在逗号时拆分异常 - 读取形状文本时增加
hasattr(shape, "text")判断,跳过图片、自由形状、表格单元格容器等无直接text属性的元素,避免运行时报错
内容的提问来源于stack exchange,提问作者road_rash
相关产品推荐
相关产品推荐

