You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何清理pdfplumber解析的印度人口预测PDF表格并转为CSV?

问题背景

我正在开展印度各地区人口预测项目,因印度自2011年未开展人口普查,需采用人口预测数据。项目需分析地区儿童相关变量,因此需要6岁以下儿童的预测数据。目前数据存储在一份1300多页的PDF文件中,每页包含两个表格。我尝试使用pypdf、pdfplumber等库解析数据,其中pdfplumber可读取文本但输出结果混乱,存在大量None和"\n",希望得到与PDF中格式一致或近似的表格,最终转换为CSV格式。我使用的代码如下:

import pdfplumber    

page_index = 52      # 目标表格所在的页码(索引从0开始)
with pdfplumber.open(r"filepath.pdf") as pdf:
    first_page = pdf.pages[page_index]
    tables = first_page.extract_tables()    
    for table in tables:
        print(table)

解析结果存在大量无效值与换行,无法直接使用。该PDF为印度国际人口科学研究所发布的《印度各地区2012-2031年按五岁年龄组和性别划分的年度人口预测》报告,表格从第45页开始。

解决方案

1. 优化pdfplumber的表格识别参数

默认的extract_tables参数可能无法适配该PDF的复杂表格结构,通过自定义表格识别策略,能大幅提升提取精度:

import pdfplumber
import csv

page_index = 52  # 目标页码(索引从0开始)
with pdfplumber.open("filepath.pdf") as pdf:
    page = pdf.pages[page_index]
    # 针对该PDF的表格特征调整识别参数
    tables = page.extract_tables({
        "vertical_strategy": "lines",  # 基于PDF中的线条识别垂直单元格边界
        "horizontal_strategy": "lines", # 基于线条识别水平单元格边界
        "snap_tolerance": 3,  # 允许线条轻微偏移的像素值
        "join_tolerance": 3,  # 合并相近短线条的像素值
        "edge_min_length": 10, # 过滤过短的无效线条
        "min_words_vertical": 3, # 垂直方向最少单词数,确保识别的是表格区域
        "min_words_horizontal": 3
    })
    
    # 清洗提取到的数据:移除None、替换换行符、去除多余空格
    cleaned_tables = []
    for table in tables:
        cleaned_table = []
        for row in table:
            cleaned_row = []
            for cell in row:
                if cell is not None:
                    processed_cell = cell.replace("\n", " ").strip()
                    cleaned_row.append(processed_cell)
                else:
                    cleaned_row.append("")
            cleaned_table.append(cleaned_row)
        cleaned_tables.append(cleaned_table)
    
    # 将每页的两个表格分别保存为CSV文件
    for idx, table in enumerate(cleaned_tables, 1):
        with open(f"page_{page_index+1}_table_{idx}.csv", "w", newline="", encoding="utf-8") as f:
            writer = csv.writer(f)
            writer.writerows(table)

2. 手动定位表格区域

如果自动识别仍不准确,可通过手动裁剪页面区域提取每页的两个表格(先导出页面截图确认坐标):

import pdfplumber
import csv

page_index = 52
with pdfplumber.open("filepath.pdf") as pdf:
    page = pdf.pages[page_index]
    # 导出页面截图用于确认坐标:page.to_image().save("page_52_preview.png")
    # 根据截图设置两个表格的裁剪区域(left, top, right, bottom)
    table1_area = (20, 120, page.width - 20, page.height / 2 - 30)
    table2_area = (20, page.height / 2 + 20, page.width - 20, page.height - 50)
    
    # 逐个提取并清洗表格
    for idx, area in enumerate([table1_area, table2_area], 1):
        cropped_page = page.crop(area)
        table = cropped_page.extract_table({
            "vertical_strategy": "lines",
            "horizontal_strategy": "lines"
        })
        cleaned_table = []
        for row in table:
            cleaned_row = [cell.replace("\n", " ").strip() if cell else "" for cell in row]
            cleaned_table.append(cleaned_row)
        # 保存CSV
        with open(f"page_{page_index+1}_table_{idx}.csv", "w", newline="", encoding="utf-8") as f:
            writer = csv.writer(f)
            writer.writerows(cleaned_table)

3. 处理合并单元格

该PDF表格存在大量合并单元格(如地区名称列),提取后会出现空值,需手动填充:

# 在清洗表格后添加以下代码处理第一列的合并单元格
for table in cleaned_tables:
    prev_district = ""
    for row in table:
        if row[0]:  # 如果当前行第一列有内容,更新基准值
            prev_district = row[0]
        else:  # 否则填充上一行的地区名称
            row[0] = prev_district

4. 备选工具:camelot-py

如果pdfplumber效果仍不理想,可尝试专门的PDF表格提取工具camelot-py,它对带线条的表格支持更好:

# 安装依赖
pip install camelot-py[cv]
import camelot

# 注意:camelot的页码从1开始,对应原PDF的第52页需传入53
page_num = 53
# 使用lattice模式(适合有明确线条的表格)提取
tables = camelot.read_pdf("filepath.pdf", pages=str(page_num), flavor="lattice")
# 保存每个表格为CSV
for idx, table in enumerate(tables, 1):
    table.to_csv(f"page_{page_num}_table_{idx}.csv")

内容的提问来源于stack exchange,提问作者simrpal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 13:02:01