You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pdfplumber提取PDF表格:如何合并跨行多行单元格?

解决PDF提取表格中跨行多行单元格的合并问题

我用pdfplumber提取PDF表格数据时,遇到跨行单元格被拆分成多行的问题。比如某单元格内容被拆成两行:

AÉRO-SPORT DU GRAND-DUCHÉ
DE LUXEMBOURG A.S.B.L.
需要合并成单行:AÉRO-SPORT DU GRAND-DUCHÉ DE LUXEMBOURG A.S.B.L.,且要对所有列应用这个处理。

当前使用的脚本如下:

import pdfplumber
import pandas as pd

pdf_path = "C:/Users/xxxx/Downloads/relev-aronefs-11-12-2023.pdf"
all_rows = []
explicit_vertical_lines = [45, 100, 250, 410, 480, 640, 780]

with pdfplumber.open(pdf_path) as pdf:
    for page_num in range(len(pdf.pages)):
        page = pdf.pages[page_num]

        if page_num == 0:
            top_region = 50 
        else:  # For all other pages
            top_region = 0

        cropped_page = page.crop((0, top_region, page.width, page.height))

        table = cropped_page.extract_table({
            "vertical_strategy": "explicit",
            "explicit_vertical_lines": explicit_vertical_lines,
            "horizontal_strategy": "text",
        })

        if table:
            if page_num == 0:
                all_rows += table[1:] 
            else:
                all_rows += table

df = pd.DataFrame(all_rows)

df.columns = ['Immat', 'Constructeur', 'Type d’aéronef', 'SN aéronef', 'Propriétaire', 'Exploitant']

output_csv_path = "C:/Users/xxxx/Downloads/extracted_table_data.csv"
df.to_csv(output_csv_path, index=False)

print(df.head())

解决方案

核心思路是:识别拆分的跨行行(通常前导唯一标识列为空),将其内容合并到上一行对应列,最后清理无效行。

修改后的完整脚本:

import pdfplumber
import pandas as pd

pdf_path = "C:/Users/xxxx/Downloads/relev-aronefs-11-12-2023.pdf"
all_rows = []
explicit_vertical_lines = [45, 100, 250, 410, 480, 640, 780]

with pdfplumber.open(pdf_path) as pdf:
    for page_num in range(len(pdf.pages)):
        page = pdf.pages[page_num]

        if page_num == 0:
            top_region = 50 
        else:
            top_region = 0

        cropped_page = page.crop((0, top_region, page.width, page.height))

        table = cropped_page.extract_table({
            "vertical_strategy": "explicit",
            "explicit_vertical_lines": explicit_vertical_lines,
            "horizontal_strategy": "text",
        })

        if table:
            if page_num == 0:
                all_rows += table[1:] 
            else:
                all_rows += table

# 处理跨行合并的单元格
merged_rows = []
for row in all_rows:
    # 以第一列(Immat)为空判断是否为跨行拆分行(该列无跨行合并)
    if not row[0] and merged_rows:
        # 遍历列,将当前行内容追加到上一行对应列
        for col_idx in range(len(row)):
            if row[col_idx]:
                if merged_rows[-1][col_idx]:
                    merged_rows[-1][col_idx] += " " + row[col_idx]
                else:
                    merged_rows[-1][col_idx] = row[col_idx]
    else:
        # 深拷贝行,避免后续修改影响已存入的数据
        merged_rows.append(row.copy())

df = pd.DataFrame(merged_rows)
df.columns = ['Immat', 'Constructeur', 'Type d’aéronef', 'SN aéronef', 'Propriétaire', 'Exploitant']

output_csv_path = "C:/Users/xxxx/Downloads/extracted_table_data.csv"
df.to_csv(output_csv_path, index=False)

print(df.head())

关键逻辑说明

  1. 拆分行判断:利用表格中Immat列作为唯一标识(无跨行合并),若该列为空,则判定当前行是上一行的跨行拆分内容。
  2. 内容合并:遍历拆分行的每一列,将非空内容追加到上一行对应列,用空格分隔原分行文本。
  3. 深拷贝行:确保每一行的修改不会影响已存入合并列表的行数据。

注意事项

如果你的表格中第一列也存在跨行情况,可以调整判断条件(比如检查多个关键列是否为空),或者根据实际表格结构修改拆分行的识别逻辑。

内容的提问来源于stack exchange,提问作者Mark k

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 23:42:54