如何清理pdfplumber解析的印度人口预测PDF表格并转为CSV?
问题背景
我正在开展印度各地区人口预测项目,因印度自2011年未开展人口普查,需采用人口预测数据。项目需分析地区儿童相关变量,因此需要6岁以下儿童的预测数据。目前数据存储在一份1300多页的PDF文件中,每页包含两个表格。我尝试使用pypdf、pdfplumber等库解析数据,其中pdfplumber可读取文本但输出结果混乱,存在大量None和"\n",希望得到与PDF中格式一致或近似的表格,最终转换为CSV格式。我使用的代码如下:
import pdfplumber page_index = 52 # 目标表格所在的页码(索引从0开始) with pdfplumber.open(r"filepath.pdf") as pdf: first_page = pdf.pages[page_index] tables = first_page.extract_tables() for table in tables: print(table)
解析结果存在大量无效值与换行,无法直接使用。该PDF为印度国际人口科学研究所发布的《印度各地区2012-2031年按五岁年龄组和性别划分的年度人口预测》报告,表格从第45页开始。
解决方案
1. 优化pdfplumber的表格识别参数
默认的extract_tables参数可能无法适配该PDF的复杂表格结构,通过自定义表格识别策略,能大幅提升提取精度:
import pdfplumber import csv page_index = 52 # 目标页码(索引从0开始) with pdfplumber.open("filepath.pdf") as pdf: page = pdf.pages[page_index] # 针对该PDF的表格特征调整识别参数 tables = page.extract_tables({ "vertical_strategy": "lines", # 基于PDF中的线条识别垂直单元格边界 "horizontal_strategy": "lines", # 基于线条识别水平单元格边界 "snap_tolerance": 3, # 允许线条轻微偏移的像素值 "join_tolerance": 3, # 合并相近短线条的像素值 "edge_min_length": 10, # 过滤过短的无效线条 "min_words_vertical": 3, # 垂直方向最少单词数,确保识别的是表格区域 "min_words_horizontal": 3 }) # 清洗提取到的数据:移除None、替换换行符、去除多余空格 cleaned_tables = [] for table in tables: cleaned_table = [] for row in table: cleaned_row = [] for cell in row: if cell is not None: processed_cell = cell.replace("\n", " ").strip() cleaned_row.append(processed_cell) else: cleaned_row.append("") cleaned_table.append(cleaned_row) cleaned_tables.append(cleaned_table) # 将每页的两个表格分别保存为CSV文件 for idx, table in enumerate(cleaned_tables, 1): with open(f"page_{page_index+1}_table_{idx}.csv", "w", newline="", encoding="utf-8") as f: writer = csv.writer(f) writer.writerows(table)
2. 手动定位表格区域
如果自动识别仍不准确,可通过手动裁剪页面区域提取每页的两个表格(先导出页面截图确认坐标):
import pdfplumber import csv page_index = 52 with pdfplumber.open("filepath.pdf") as pdf: page = pdf.pages[page_index] # 导出页面截图用于确认坐标:page.to_image().save("page_52_preview.png") # 根据截图设置两个表格的裁剪区域(left, top, right, bottom) table1_area = (20, 120, page.width - 20, page.height / 2 - 30) table2_area = (20, page.height / 2 + 20, page.width - 20, page.height - 50) # 逐个提取并清洗表格 for idx, area in enumerate([table1_area, table2_area], 1): cropped_page = page.crop(area) table = cropped_page.extract_table({ "vertical_strategy": "lines", "horizontal_strategy": "lines" }) cleaned_table = [] for row in table: cleaned_row = [cell.replace("\n", " ").strip() if cell else "" for cell in row] cleaned_table.append(cleaned_row) # 保存CSV with open(f"page_{page_index+1}_table_{idx}.csv", "w", newline="", encoding="utf-8") as f: writer = csv.writer(f) writer.writerows(cleaned_table)
3. 处理合并单元格
该PDF表格存在大量合并单元格(如地区名称列),提取后会出现空值,需手动填充:
# 在清洗表格后添加以下代码处理第一列的合并单元格 for table in cleaned_tables: prev_district = "" for row in table: if row[0]: # 如果当前行第一列有内容,更新基准值 prev_district = row[0] else: # 否则填充上一行的地区名称 row[0] = prev_district
4. 备选工具:camelot-py
如果pdfplumber效果仍不理想,可尝试专门的PDF表格提取工具camelot-py,它对带线条的表格支持更好:
# 安装依赖 pip install camelot-py[cv]
import camelot # 注意:camelot的页码从1开始,对应原PDF的第52页需传入53 page_num = 53 # 使用lattice模式(适合有明确线条的表格)提取 tables = camelot.read_pdf("filepath.pdf", pages=str(page_num), flavor="lattice") # 保存每个表格为CSV for idx, table in enumerate(tables, 1): table.to_csv(f"page_{page_num}_table_{idx}.csv")
内容的提问来源于stack exchange,提问作者simrpal
相关产品推荐
相关产品推荐

