You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyMuPDF提取非可编辑PDF键值对及复选框数据遇问题求助

问题:PyMuPDF提取非可编辑PDF字段与复选框数据异常

遇到两个核心问题:

  1. 提取的字典中字段与值对应错误,出现先列所有键再单独列对应值的情况
  2. 无法读取任何复选框数据

尝试的代码

import fitz  
import pandas as pd
import re

# Function to clean text
def clean_text(text):
    return re.sub('\\s+', ' ', text).strip()

def is_field_name(line):
    # A field name is likely to end with a colon followed by an optional space
    return bool(re.match(r'.*:\\s*$', line))

# Function to determine if the next line is a checkbox indicator

def is_checkbox(line):
    # Looking for lines that have a checkbox indication, e.g., "\[X\] Yes" or "\[ \] No"
    return bool(re.match(r'\[(X| )\]\\s\*(Yes|No)', line))

pdf_document = fitz.open('PTO 2024.pdf')

page = pdf_document[0]
text = page.get_text("text")

pdf_document.close()

lines = text.split('\\n')

extracted_data = {}
current_field = None

# Process each line, assuming that fields are followed by their values or checkboxes
for i, line in enumerate(lines):
    line = clean_text(line)
    if is_field_name(line):
        # The current line is a field name
        current_field = line[:-1]  # Remove the colon at the end
    elif current_field:
        if is_checkbox(line):
            # Line is a checkbox indicator, e.g., "\[X\] Yes"
            extracted_data[current_field] = 'Checked' if 'Yes' in line else 'Unchecked'
            current_field = None  # Reset current field after capturing checkbox
        elif line:
            # Line has content and is not a checkbox, so it's a value for the current field
            extracted_data[current_field] = line
            current_field = None  # Reset current field after capturing value

df_extracted = pd.DataFrame(list(extracted_data.items()), columns=['Field', 'Value'])

print("Extracted Lines:")
for line in lines:
    print(line)

print("\nExtracted Data:")
print(df_extracted.head())

问题分析与解决方案

问题1:字段与值对应错误

原因

原代码假设字段名的下一行必然是对应值,但实际PDF文本提取时,字段与值之间可能存在空行,或者值跨多行,导致current_field未被及时赋值,后续新字段覆盖旧字段,最终出现键值错位。

修复方案

  • 过滤空行,避免空行干扰字段匹配逻辑
  • 允许字段值跨多行累积,直到遇到下一个字段名
  • 优化字段名的清理逻辑,兼容冒号后带空格的情况

问题2:无法读取复选框数据

原因

  1. 正则表达式错误:原正则r'\[(X| )\]\\s\*(Yes|No)'存在转义错误(\\s*应为\s*),且仅匹配Yes/No选项,兼容性差
  2. 非可编辑PDF的复选框可能是图形元素而非文本,get_text("text")无法提取这类内容

修复方案

  1. 修正正则表达式,匹配更通用的复选框格式(如[X]或[ ]开头的任意行)
  2. 添加图形复选框提取逻辑:通过get_drawings()获取页面图形,识别复选框形状(正方形),并匹配附近的字段文本

修改后的完整代码

import fitz  
import pandas as pd
import re

# 清理文本,合并多余空格但保留必要分隔
def clean_text(text):
    return re.sub(r'\s+', ' ', text).strip()

def is_field_name(line):
    # 匹配以冒号结尾的字段名,兼容冒号后带空格的情况
    return bool(re.match(r'.*:\s*$', line))

def is_checkbox(line):
    # 匹配[X]或[ ]开头的行,支持后续任意文本
    return bool(re.match(r'\[(X| )\]\s*', line))

pdf_document = fitz.open('PTO 2024.pdf')
page = pdf_document[0]

# 提取并预处理文本行,过滤空行
text = page.get_text("text")
lines = [clean_text(line) for line in text.split('\n') if clean_text(line)]

# 提取页面图形,用于识别非文本型复选框
drawings = page.get_drawings()
checkbox_shapes = []
for draw in drawings:
    # 筛选正方形图形(假设复选框为10-20px的正方形,可根据实际调整)
    rect = draw['rect']
    if abs(rect.width - rect.height) < 2 and 10 < rect.width < 20:
        checkbox_shapes.append(rect)

# 提取所有文本的位置信息,用于匹配图形复选框对应的字段
word_list = page.get_text("words")  # 格式:(x0, y0, x1, y1, text, block_no, line_no, word_no)

extracted_data = {}
current_field = None

# 处理文本型字段与复选框
for line in lines:
    if is_field_name(line):
        # 清理字段名,去掉末尾的冒号和空格
        current_field = line.rstrip(': ').strip()
    elif current_field:
        if is_checkbox(line):
            # 判断复选框状态
            extracted_data[current_field] = 'Checked' if '[X]' in line else 'Unchecked'
            # 提取复选框对应的选项文本(可选)
            option_text = line.split(']')[-1].strip()
            if option_text:
                extracted_data[f"{current_field} (选项)"] = option_text
            current_field = None
        else:
            # 处理多行字段值,累积内容
            if current_field in extracted_data:
                extracted_data[current_field] += f" {line}"
            else:
                extracted_data[current_field] = line

# 处理图形型复选框(文本提取不到的情况)
for rect in checkbox_shapes:
    # 找到复选框附近的文本(交集判断)
    nearby_words = [word for word in word_list if fitz.Rect(word[:4]).intersects(rect)]
    if nearby_words:
        # 合并附近文本作为字段名
        field_name = ' '.join([word[4] for word in nearby_words])
        # 判断复选框是否被勾选(通过填充颜色判断,白色为未勾选)
        is_checked = any(draw['fill'] != (1.0, 1.0, 1.0) for draw in drawings if draw['rect'] == rect)
        extracted_data[field_name] = 'Checked' if is_checked else 'Unchecked'

# 转换为DataFrame并输出
df_extracted = pd.DataFrame(list(extracted_data.items()), columns=['Field', 'Value'])

print("提取的文本行:")
for line in lines:
    print(line)

print("\n提取的数据:")
print(df_extracted)

pdf_document.close()

内容的提问来源于stack exchange,提问作者Mar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 08:49:54