You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Pandas读取PDF转Excel非结构化表格报错并提取考试信息导出新表

问题原因

你之前的代码报KeyError有两个核心原因:

  1. pd.read_excel默认会把Excel第一行作为DataFrame的列名,你的原始表格是非结构化的,没有规范表头,自然不存在名为A2、B2的列
  2. iterrows返回的row对象是按列名取值,不能直接用Excel单元格坐标取数,按坐标取数需要用iloc方法

修复方案

读文件时先取消默认表头,把所有内容都当做数据读取,列自动用0、1、2……编号(对应Excel的A、B、C……列),行号用0、1、2……编号(对应Excel的1、2、3……行),然后遍历单元格匹配目标字段提取值即可。

完整可运行代码

import pandas as pd

# 读取Excel时不指定表头,所有内容按原始位置读取
df = pd.read_excel(r'Copy of pdf-v1.xlsx', header=None)

# 定义原始字段和目标表头的对应关系
field_map = {
    'name': 'Name',
    'Exam ID': 'Exam Id',
    'ID': 'ID',
    'Test Address': 'Address',
    'theory test time': 'Theory Test time',
    'Skill test time': 'Skill test Time',
    'Skill test Address': 'Skill test Address'
}

result = []  # 存储所有提取完成的人员信息
current_person = {}  # 存储正在收集的单个人的信息

# 遍历所有行
for _, row in df.iterrows():
    # 遍历当前行的每一列
    for col_idx in range(len(row)):
        cell_content = str(row[col_idx]).strip()
        # 匹配到目标字段时,取右侧相邻单元格的值
        if cell_content in field_map:
            value = str(row[col_idx + 1]).strip() if col_idx + 1 < len(row) else ''
            current_person[field_map[cell_content]] = value
            # 单个人的所有字段收集完成后,存入结果列表,重置临时存储对象
            if len(current_person) == len(field_map):
                result.append(current_person)
                current_person = {}

# 最后如果有未存入结果的零散数据,也补入
if current_person:
    result.append(current_person)

# 按要求的列顺序生成DataFrame,导出到新Excel
output_df = pd.DataFrame(result, columns=['Name', 'ID', 'Exam Id', 'Theory Test time', 'Address', 'Skill test Time', 'Skill test Address'])
output_df.to_excel('提取后的考试信息.xlsx', index=False)

适配说明

如果你的原始数据是单个人的信息分散在多行,只需要调整收集逻辑:匹配到name字段时,就认为是新一条记录的开始,先把之前收集完成的current_person存入结果列表,再重置临时存储对象即可。

内容的提问来源于stack exchange,提问作者Sherry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 22:27:00