You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

寻求含彩色单元格的PDF转Excel解决方案:tabula-py存识别问题

解决方案:识别PDF彩色单元格并转换为Excel

针对你需要识别PDF表格彩色单元格并转Excel的需求,以下是几个本地专属方案:

1. 使用PyMuPDF(fitz)+ openpyxl

PyMuPDF可以精准提取PDF中每个文本块的位置、内容以及背景/字体颜色,openpyxl负责将数据写入Excel并还原单元格颜色。

核心步骤:

  • 遍历PDF页面,定位表格区域(可通过坐标手动指定或结合表格检测逻辑)
  • 提取每个单元格的文本内容和背景色RGB值
  • 创建Excel文件,写入内容并设置对应单元格的填充颜色

示例代码:

import fitz
from openpyxl import Workbook
from openpyxl.styles import PatternFill

# 打开PDF
doc = fitz.open("colored_table.pdf")
wb = Workbook()
ws = wb.active

# 假设表格在第一页,手动指定表格区域坐标(x0, y0, x1, y1)
table_rect = fitz.Rect(50, 100, 550, 600)
page = doc[0]

# 提取表格内的文本块
text_blocks = page.get_text("blocks", clip=table_rect)

# 按行排序文本块(y坐标升序,x坐标升序)
text_blocks.sort(key=lambda b: (b[1], b[0]))

row_idx = 1
col_idx = 1
prev_y = None

for block in text_blocks:
    x0, y0, x1, y1, text, _, _ = block
    # 检测是否换行(y坐标差超过阈值则视为新行)
    if prev_y is not None and abs(y0 - prev_y) > 5:
        row_idx += 1
        col_idx = 1
    prev_y = y0

    # 提取单元格背景色(遍历页面图形元素匹配位置)
    background_fill = None
    for shape in page.get_drawings():
        if shape["rect"].intersects(fitz.Rect(x0, y0, x1, y1)):
            # 将fitz颜色值转换为0-255范围的RGB
            rgb = [int(c * 255) for c in shape["fill"]]
            background_fill = PatternFill(start_color=f"{rgb[0]:02X}{rgb[1]:02X}{rgb[2]:02X}",
                                         end_color=f"{rgb[0]:02X}{rgb[1]:02X}{rgb[2]:02X}",
                                         fill_type="solid")
            break

    # 写入Excel并设置颜色
    cell = ws.cell(row=row_idx, column=col_idx, value=text.strip())
    if background_fill:
        cell.fill = background_fill
    col_idx += 1

wb.save("colored_table_output.xlsx")
doc.close()

2. 使用pdfplumber + openpyxl

pdfplumber对表格结构的识别更友好,同时能获取文本的样式属性(包括背景色),适合结构化较强的PDF表格。

示例代码:

import pdfplumber
from openpyxl import Workbook
from openpyxl.styles import PatternFill

with pdfplumber.open("colored_table.pdf") as pdf:
    page = pdf.pages[0]
    # 自动提取表格(可指定表格区域)
    table = page.extract_table()
    wb = Workbook()
    ws = wb.active

    # 遍历表格单元格,提取颜色信息
    for row_idx, row in enumerate(table, start=1):
        for col_idx, cell_text in enumerate(row, start=1):
            # 获取单元格区域
            cell_bbox = page.cells[row_idx-1][col_idx-1]["bbox"]
            # 提取单元格背景色
            fill_color = None
            for obj in page.objects["rect"]:
                if obj["bbox"] == cell_bbox and obj["fill"] is not None:
                    rgb = obj["fill"]
                    fill_color = PatternFill(start_color=f"{int(rgb[0]*255):02X}{int(rgb[1]*255):02X}{int(rgb[2]*255):02X}",
                                           end_color=f"{int(rgb[0]*255):02X}{int(rgb[1]*255):02X}{int(rgb[2]*255):02X}",
                                           fill_type="solid")
                    break
            # 写入单元格并设置颜色
            cell = ws.cell(row=row_idx, column=col_idx, value=cell_text)
            if fill_color:
                cell.fill = fill_color
    wb.save("colored_table_pdfplumber.xlsx")

3. 结合tabula-py + PyMuPDF(兼顾结构与颜色)

如果tabula-py能准确提取表格结构,只是缺颜色信息,可以用它先导出表格框架,再用PyMuPDF补充颜色数据:

  • 用tabula-py提取表格数据到DataFrame
  • 用PyMuPDF定位每个单元格的颜色
  • 将颜色信息映射到Excel单元格

关键代码片段:

import tabula
import fitz
from openpyxl import Workbook
from openpyxl.utils.dataframe import dataframe_to_rows

# 用tabula提取表格结构
df = tabula.read_pdf("colored_table.pdf", pages=1, output_format="dataframe")[0]

# 用PyMuPDF提取颜色信息(参考方案1的逻辑)
# ...(此处省略颜色提取代码)

# 将DataFrame写入Excel并设置颜色
wb = Workbook()
ws = wb.active
for r in dataframe_to_rows(df, index=False, header=True):
    ws.append(r)
# 遍历单元格设置颜色(需对应位置匹配)
# ...
wb.save("tabula_with_color.xlsx")

注意事项:

  • 以上方案需要手动调整坐标或匹配逻辑,根据你的PDF表格结构微调
  • 若PDF是扫描件(图片型),需要先结合OCR工具(如pytesseract)提取文本,再识别颜色
  • 所有方案均为本地工具,无需依赖在线服务

内容的提问来源于stack exchange,提问作者marskernel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 01:27:10