You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何清洗无明确分隔符的多页表格类混乱PDF数据?

解决PDF跨页表格拆分与数字分隔问题

一、提取无页码的纯文本内容

用pdfplumber可精准定位并排除页脚页码区域,提取每页有效内容:

import pdfplumber

def extract_text_without_page_numbers(pdf_path):
    all_text = ""
    with pdfplumber.open(pdf_path) as pdf:
        for page in pdf.pages:
            # 假设页码在页脚底部10%区域,可根据实际调整坐标范围
            page_height = page.height
            cropped_page = page.crop((0, 0, page.width, page_height * 0.9))
            all_text += cropped_page.extract_text() + "\n"
    return all_text

# 导出无页码文本
raw_text = extract_text_without_page_numbers("your_file.pdf")
with open("no_page_numbers.txt", "w", encoding="utf-8") as f:
    f.write(raw_text)

如果页码位置不固定,也可以直接用正则匹配删除常见页码格式:

import re

cleaned_text = re.sub(r"(第\d+页|\d+/\d+|\d+$)", "", raw_text, flags=re.MULTILINE)

二、按标题拆分单个表格

假设每个表格有明确的专属标题(如“销售数据统计表”“月度库存表”),用正则分割文本并保存为独立文件:

# 匹配表格标题的正则,需根据实际标题格式调整
table_title_pattern = re.compile(r"(销售数据统计表|月度库存表|[\u4e00-\u9fa5]+数据表)")
# 分割文本为标题+内容的块
table_blocks = re.split(table_title_pattern, cleaned_text)

# 配对标题与内容,逐个保存表格
for i in range(1, len(table_blocks), 2):
    title = table_blocks[i]
    content = table_blocks[i+1].strip()
    with open(f"{title}.txt", "w", encoding="utf-8") as f:
        f.write(f"{title}\n{content}")

三、按小数点后位数拆分连写数字

如果数字是固定小数位数(比如两位),用正则提取所有符合格式的数字,完成拆分:

def split_connected_numbers(text, decimal_places=2):
    # 匹配带指定小数位数的数字
    num_pattern = re.compile(r"\d+\.\d{" + str(decimal_places) + "}")
    numbers = num_pattern.findall(text)
    # 按行输出拆分后的数字,也可根据需求调整排版
    return "\n".join(numbers)

# 批量处理每个表格文件
import os
for filename in os.listdir("./"):
    if filename.endswith(".txt") and not filename.startswith("拆分后_"):
        with open(filename, "r", encoding="utf-8") as f:
            content = f.read()
        split_content = split_connected_numbers(content)
        with open(f"拆分后_{filename}", "w", encoding="utf-8") as f:
            f.write(split_content)

补充说明

  • 若表格标题格式不统一,可先整理明确的标题列表,再逐个匹配分割
  • 若小数位数不固定,可将正则调整为r"\d+\.\d+"匹配所有带小数的数字,再结合业务规则筛选
  • 拆分后的数据可导入pandas生成DataFrame,进一步整理为规范的表格格式

内容的提问来源于stack exchange,提问作者Chloe Chan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 16:54:58