You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDFplumber处理Excel导出PDF异常:提取后写入Excel含二进制内容

问题:PDF表格提取与数据映射异常处理

背景

我有两个从Excel导出的外观完全一致的PDF文件,使用以下pdfplumber代码提取表格数据:

all_data = []

with pdfplumber.open(pdf_path) as pdf:
    for page in pdf.pages:
        table = page.extract_table()
        if table:
            filtered_table = table[5:]  # 跳过表头
            header = [clean_text(h) for h in filtered_table[0]]  # 清理表头文本
            data = filtered_table[1:]

            # 移除表格中的空行
            data = [row for row in data if any(cell and cell.strip() for cell in row)]

            # 将当前页的数据追加到总列表
            all_data.extend(data)
            
if not all_data:
    return

df = pd.DataFrame(all_data, columns=header)
df.dropna(how='all', inplace=True)

这段代码仅对其中一个PDF有效,在VS Code中检查后发现两个PDF的头部结构存在差异。将第二个PDF提取的数据写入Excel时,内容中夹杂大量二进制代码。

排查更新

后续排查确认:pdfplumber读取的原始内容本身是正确的,但执行以下映射函数后出现异常:

def apply_mapping(text):
    for key, value in data_mapping.items():
        if key in text:
            return value
    return text

for col in df.columns:
    df[col] = df[col].apply(lambda x: apply_mapping(clean_text(str(x))) if x is not None else "")

执行时收到FutureWarning提示:

FutureWarning: Series.__getitem__ treating keys as positions is deprecated. In a future version, integer keys will always be treated as labels (consistent with DataFrame behavior). To access a value by position, use ser.iloc[pos]
df[col] = df[col].apply(lambda x: apply_mapping(clean_text(str(x))) if x is not None else "")

需要解决的问题

  • 实现两个PDF文件的正确读取
  • 解决写入Excel时内容夹杂二进制代码的问题
  • 消除上述FutureWarning警告

内容的提问来源于stack exchange,提问作者ruggero luca salvatore

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 19:12:32