PDFplumber处理Excel导出PDF异常:提取后写入Excel含二进制内容
问题:PDF表格提取与数据映射异常处理
背景
我有两个从Excel导出的外观完全一致的PDF文件,使用以下pdfplumber代码提取表格数据:
all_data = [] with pdfplumber.open(pdf_path) as pdf: for page in pdf.pages: table = page.extract_table() if table: filtered_table = table[5:] # 跳过表头 header = [clean_text(h) for h in filtered_table[0]] # 清理表头文本 data = filtered_table[1:] # 移除表格中的空行 data = [row for row in data if any(cell and cell.strip() for cell in row)] # 将当前页的数据追加到总列表 all_data.extend(data) if not all_data: return df = pd.DataFrame(all_data, columns=header) df.dropna(how='all', inplace=True)
这段代码仅对其中一个PDF有效,在VS Code中检查后发现两个PDF的头部结构存在差异。将第二个PDF提取的数据写入Excel时,内容中夹杂大量二进制代码。
排查更新
后续排查确认:pdfplumber读取的原始内容本身是正确的,但执行以下映射函数后出现异常:
def apply_mapping(text): for key, value in data_mapping.items(): if key in text: return value return text for col in df.columns: df[col] = df[col].apply(lambda x: apply_mapping(clean_text(str(x))) if x is not None else "")
执行时收到FutureWarning提示:
FutureWarning:
Series.__getitem__treating keys as positions is deprecated. In a future version, integer keys will always be treated as labels (consistent withDataFramebehavior). To access a value by position, useser.iloc[pos]
df[col] = df[col].apply(lambda x: apply_mapping(clean_text(str(x))) if x is not None else "")
需要解决的问题
- 实现两个PDF文件的正确读取
- 解决写入Excel时内容夹杂二进制代码的问题
- 消除上述FutureWarning警告
内容的提问来源于stack exchange,提问作者ruggero luca salvatore
相关产品推荐
相关产品推荐

