如何将tabula-py读取的非结构化Pandas DataFrame转为结构化格式?
重构单列PDF提取结果为结构化DataFrame
核心思路
由于提取后的DataFrame是单列的标识-值对,且地址存在多行内容,核心逻辑是逐行遍历识别标识字段,收集对应的值(包括多行地址),最后拆分地址组件并构建结构化数据。
代码实现
假设你通过tabula读取后的单列DataFrame列名为content,以下是可直接复用的实现:
1. 导入依赖
import pandas as pd
2. 加载原始数据
替换为你实际从tabula获取的DataFrame即可:
# 模拟tabula读取的单列结果(实际使用时替换为你的数据) raw_data = { "content": [ "Name: John Doe", "Address:", "123 Main St", "Springfield, IL 62704", "Date: 2024-05-20", "Name: Jane Smith", "Address:", "456 Oak Ave", "Chicago, IL 60601", "Date: 2024-05-21" ] } raw_df = pd.DataFrame(raw_data)
3. 核心转换逻辑
structured_entries = [] current_entry = {} collecting_address = False address_parts = [] for line in raw_df["content"]: line = line.strip() if not line: continue # 跳过空行 # 识别Name字段,开启新条目 if line.lower().startswith("name:"): # 先处理上一个未完成的条目 if current_entry: # 填充地址信息 if address_parts: current_entry["address_street"] = address_parts[0] if len(address_parts) >= 2: city_section, state_zip = address_parts[1].split(", ", 1) current_entry["address_city"] = city_section current_entry["address_state"] = state_zip.split()[0] structured_entries.append(current_entry) # 初始化新条目 current_entry = {"name": line.replace("Name:", "").strip()} collecting_address = False address_parts = [] # 触发地址收集模式 elif line.lower().startswith("address:"): collecting_address = True # 收集地址行 elif collecting_address: address_parts.append(line) # 识别Date字段,完成当前条目 elif line.lower().startswith("date:"): current_entry["date"] = line.replace("Date:", "").strip() # 填充地址信息 if address_parts: current_entry["address_street"] = address_parts[0] if len(address_parts) >= 2: city_section, state_zip = address_parts[1].split(", ", 1) current_entry["address_city"] = city_section current_entry["address_state"] = state_zip.split()[0] structured_entries.append(current_entry) # 重置状态 current_entry = {} collecting_address = False address_parts = [] # 处理最后一个可能未提交的条目(若PDF结尾无Date字段) if current_entry: if address_parts: current_entry["address_street"] = address_parts[0] if len(address_parts) >= 2: city_section, state_zip = address_parts[1].split(", ", 1) current_entry["address_city"] = city_section current_entry["address_state"] = state_zip.split()[0] structured_entries.append(current_entry) # 生成最终结构化DataFrame final_df = pd.DataFrame(structured_entries) print(final_df)
适配调整说明
- 地址格式兼容:如果地址包含更多行(如公寓号)或分隔方式不同,可修改地址解析逻辑,比如用正则表达式
r'[A-Z]{2}'匹配州缩写来定位address_state字段。 - 标识格式兼容:若PDF中标识无冒号(如
Name John Doe)或大小写混乱,可将判断条件改为line.lower().startswith("name"),再用字符串分割提取值。 - 多PDF批量处理:将转换逻辑封装成函数,遍历每个PDF的DataFrame(或tabula返回的DataFrame列表),处理后用
pd.concat()合并所有结果。
内容的提问来源于stack exchange,提问作者cekar
相关产品推荐
相关产品推荐

