You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将tabula-py读取的非结构化Pandas DataFrame转为结构化格式?

重构单列PDF提取结果为结构化DataFrame

核心思路

由于提取后的DataFrame是单列的标识-值对,且地址存在多行内容,核心逻辑是逐行遍历识别标识字段,收集对应的值(包括多行地址),最后拆分地址组件并构建结构化数据。

代码实现

假设你通过tabula读取后的单列DataFrame列名为content,以下是可直接复用的实现:

1. 导入依赖

import pandas as pd

2. 加载原始数据

替换为你实际从tabula获取的DataFrame即可:

# 模拟tabula读取的单列结果(实际使用时替换为你的数据)
raw_data = {
    "content": [
        "Name: John Doe",
        "Address:",
        "123 Main St",
        "Springfield, IL 62704",
        "Date: 2024-05-20",
        "Name: Jane Smith",
        "Address:",
        "456 Oak Ave",
        "Chicago, IL 60601",
        "Date: 2024-05-21"
    ]
}
raw_df = pd.DataFrame(raw_data)

3. 核心转换逻辑

structured_entries = []
current_entry = {}
collecting_address = False
address_parts = []

for line in raw_df["content"]:
    line = line.strip()
    if not line:
        continue  # 跳过空行
    
    # 识别Name字段,开启新条目
    if line.lower().startswith("name:"):
        # 先处理上一个未完成的条目
        if current_entry:
            # 填充地址信息
            if address_parts:
                current_entry["address_street"] = address_parts[0]
                if len(address_parts) >= 2:
                    city_section, state_zip = address_parts[1].split(", ", 1)
                    current_entry["address_city"] = city_section
                    current_entry["address_state"] = state_zip.split()[0]
            structured_entries.append(current_entry)
        
        # 初始化新条目
        current_entry = {"name": line.replace("Name:", "").strip()}
        collecting_address = False
        address_parts = []
    
    # 触发地址收集模式
    elif line.lower().startswith("address:"):
        collecting_address = True
    
    # 收集地址行
    elif collecting_address:
        address_parts.append(line)
    
    # 识别Date字段,完成当前条目
    elif line.lower().startswith("date:"):
        current_entry["date"] = line.replace("Date:", "").strip()
        # 填充地址信息
        if address_parts:
            current_entry["address_street"] = address_parts[0]
            if len(address_parts) >= 2:
                city_section, state_zip = address_parts[1].split(", ", 1)
                current_entry["address_city"] = city_section
                current_entry["address_state"] = state_zip.split()[0]
        structured_entries.append(current_entry)
        # 重置状态
        current_entry = {}
        collecting_address = False
        address_parts = []

# 处理最后一个可能未提交的条目(若PDF结尾无Date字段)
if current_entry:
    if address_parts:
        current_entry["address_street"] = address_parts[0]
        if len(address_parts) >= 2:
            city_section, state_zip = address_parts[1].split(", ", 1)
            current_entry["address_city"] = city_section
            current_entry["address_state"] = state_zip.split()[0]
    structured_entries.append(current_entry)

# 生成最终结构化DataFrame
final_df = pd.DataFrame(structured_entries)
print(final_df)

适配调整说明

  • 地址格式兼容:如果地址包含更多行(如公寓号)或分隔方式不同,可修改地址解析逻辑,比如用正则表达式r'[A-Z]{2}'匹配州缩写来定位address_state字段。
  • 标识格式兼容:若PDF中标识无冒号(如Name John Doe)或大小写混乱,可将判断条件改为line.lower().startswith("name"),再用字符串分割提取值。
  • 多PDF批量处理:将转换逻辑封装成函数,遍历每个PDF的DataFrame(或tabula返回的DataFrame列表),处理后用pd.concat()合并所有结果。

内容的提问来源于stack exchange,提问作者cekar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 18:42:38