You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Streamlit批量解析PDF转DataFrame时文件名赋值异常排查

问题背景

开发Streamlit PDF批量转Excel工具时,需求是给解析拼接后生成的汇总DataFrame添加PO列,存储每一行数据对应的源PDF文件名。实际运行时出现匹配错误:测试上传2个文件时,PO列所有行都填充为同一个文件名,无法实现每份解析数据与自身源文件名的对应匹配。
问题复现代码如下:

import pandas as pd
import numpy as np
import streamlit as st
from tabula.io import read_pdf
import os
import glob

# Title
st.title('PDF to Excel')

files = st.file_uploader('Upload all orders',accept_multiple_files=True)

if files:
    for file in files:
        file.seek(0)

    files_read = [read_pdf(file,pages='all')[0] for file in files]
    
    st.write(type(files_read))

    df = pd.concat(files_read)

    df["PO"] = file.name

问题运行效果参考:
PO列填充错误运行效果

故障原因

核心错误为PO列的赋值时机和逻辑错误:

  • Python中for循环、列表推导的迭代变量不会在迭代结束后销毁,会保留最后一次迭代的对象值。代码中执行完列表推导读取所有文件后,变量file指向的是最后一个被遍历的上传文件。
  • 代码是在所有文件的解析结果完成拼接、生成汇总DataFrame后,才统一给整列赋值file.name,此时只能拿到单个文件名,导致所有行的PO值完全相同,无法和源文件一一对应。
修复方案

调整赋值逻辑:在读取单个PDF生成对应DataFrame时,就给当前单文件的解析结果添加对应文件名的PO列,之后再做全量拼接,即可保证每行数据和源文件名的匹配关系。
修正后完整代码如下:

import pandas as pd
import numpy as np
import streamlit as st
from tabula.io import read_pdf
import os
import glob

st.title('PDF to Excel')

files = st.file_uploader('Upload all orders', accept_multiple_files=True)

if files:
    files_read = []
    for file in files:
        file.seek(0)
        # 读取单个PDF的表格数据
        single_pdf_df = read_pdf(file, pages='all')[0]
        # 给当前文件的解析结果绑定对应文件名
        single_pdf_df["PO"] = file.name
        files_read.append(single_pdf_df)
    
    # 拼接所有带文件名标记的DataFrame,重置行索引避免索引重复
    df = pd.concat(files_read, ignore_index=True)
    
    # 后续可添加表格展示、导出Excel等逻辑
    st.dataframe(df)

内容的提问来源于stack exchange,提问作者Pete3p0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 05:06:21