You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现PDF表格转HTML:rowspan/colspan计算问题及Camelot咨询

问题解答:Camelot能否解决PDF转HTML表格的rowspan/colspan问题

是的,Camelot 完全可以解决PDF表格转HTML时的rowspan和colspan计算问题。它在提取PDF表格时会保留单元格的原始合并信息,无需手动判断,能精准还原原表格的结构。

原代码的问题分析

你当前使用的pdfplumber方案存在两个核心缺陷:

  • pdfplumber的extract_tables()返回的是扁平化的文本矩阵,丢失了单元格合并的元数据(比如哪些单元格是合并的、合并范围多大)
  • 手动计算rowspan/colspan的逻辑有误:比如colspan的循环条件range(j + 1, j + colspan)初始值colspan=1,循环根本不会执行,无法识别横向合并;rowspan的判断也仅依赖空单元格,无法区分是合并导致的空还是原表格本身的空单元格

Camelot解决方案实现

1. 安装依赖

首先安装Camelot及必要依赖:

pip install camelot-py[cv]
# 需额外安装Ghostscript(用于PDF渲染),根据系统自行安装

2. 提取表格并生成HTML

以下代码会利用Camelot提取表格的合并信息,直接生成带rowspan/colspan的HTML表格:

import camelot
from bs4 import BeautifulSoup

def camelot_table_to_html(table):
    # 获取表格的总行数和列数
    max_row = max(cell.row_end for cell in table.cells)
    max_col = max(cell.col_end for cell in table.cells)
    
    # 初始化单元格矩阵,标记已处理的单元格
    cell_matrix = [[None for _ in range(max_col)] for _ in range(max_row)]
    html_rows = []
    
    for cell in table.cells:
        # 转换为从0开始的索引
        row_start = cell.row_start - 1
        row_end = cell.row_end - 1
        col_start = cell.col_start - 1
        col_end = cell.col_end - 1
        
        # 如果当前单元格未被处理(避免合并单元格重复生成)
        if cell_matrix[row_start][col_start] is None:
            rowspan = row_end - row_start + 1
            colspan = col_end - col_start + 1
            cell_text = cell.text.strip()
            
            # 生成td标签
            td_tag = f'<td rowspan="{rowspan}" colspan="{colspan}">{cell_text}</td>'
            cell_matrix[row_start][col_start] = td_tag
            
            # 标记合并范围内的其他单元格为已处理
            for r in range(row_start, row_end + 1):
                for c in range(col_start, col_end + 1):
                    if r != row_start or c != col_start:
                        cell_matrix[r][c] = "processed"
    
    # 生成HTML行
    for row in cell_matrix:
        html_row = "  <tr>\n"
        for cell in row:
            if cell and cell != "processed":
                html_row += f"    {cell}\n"
        html_row += "  </tr>"
        html_rows.append(html_row)
    
    # 拼接完整表格
    return f"<table>\n{chr(10).join(html_rows)}\n</table>"

# 提取PDF中的表格
pdf_file = 'simple_rowspan_colspan.pdf'
tables = camelot.read_pdf(pdf_file, flavor='lattice')  # 针对带边框的表格用lattice,无边框用stream

# 生成所有表格的HTML
html_tables = []
for table in tables:
    html_table = camelot_table_to_html(table)
    html_tables.append(html_table)

# 合并并格式化HTML
output_html = "\n\n".join(html_tables)
soup = BeautifulSoup(output_html, 'html.parser')
pretty_html = soup.prettify()

# 保存结果
with open('output_table.html', 'w', encoding='utf-8') as f:
    f.write(pretty_html)

print(pretty_html)

关键说明

  • flavor参数:如果你的表格是带清晰边框的,用flavor='lattice';如果是无边框的纯文本表格,用flavor='stream'
  • Camelot的cell对象包含row_start/row_end/col_start/col_end属性,直接对应单元格的合并范围,无需手动计算
  • 代码中通过cell_matrix标记已处理的合并单元格,避免重复生成td标签

内容的提问来源于stack exchange,提问作者Jayesh Ranjan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 17:24:55