You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于PDF中的粗体换行符拆分内容并提取数据至列表?

基于PDF粗体换行符拆分数据块的解决方案

嘿,我懂你现在的困扰——按行拆分PDF文本后,好多字段挤在一行里,处理起来特别闹心,想要用PDF里的粗体换行符当分隔符来切分数据块,这样提取数据肯定顺畅多了!

问题分析

你当前的代码是按固定行数(每8行)拆分PDF文本,这种方式太死板,一旦PDF格式稍有变动就会失效;而且正则匹配还存在错误(比如地址的正则和商号完全一样),最重要的是没利用PDF里的粗体格式信息,导致字段混排,提取效率低。你想要的是把内容按粗体换行符拆分成独立数据块,最终得到像这样的列表:

['DISTRICT ROW LLC', 'Premises No.: 0', 'License Key: 0', 'Date Entered:09/08/2021', 'Tradename: OLSEN RUN WINERY', 'Email Address: rachel@olsenrun.com', 'License Type/Action: F-COM/ N/O', ]

解决方案:用pdfplumber识别粗体拆分数据

我们可以用pdfplumber库,它能提取文本的字体样式(比如是否粗体),帮我们精准识别作为分隔符的粗体行,进而拆分数据块。

步骤1:安装依赖

首先安装pdfplumber:

pip install pdfplumber

步骤2:完整代码实现

import re
import pdfplumber
import csv

def is_bold(text_obj):
    # 判断文本对象是否为粗体(字体名含"Bold")
    return any("Bold" in font for font in text_obj.get("fontnames", []))

def extract_target_text_objects(page, start_index):
    # 从指定位置开始提取文本对象
    full_text = page.extract_text()
    filtered_objs = []
    for text_obj in page.extract_text_objects():
        text = text_obj.get_text()
        try:
            if full_text.index(text) >= start_index:
                filtered_objs.append(text_obj)
        except ValueError:
            continue
    return filtered_objs

def split_into_data_blocks(text_objects):
    # 按粗体行拆分数据块
    data_blocks = []
    current_block = []
    for text_obj in text_objects:
        text = text_obj.get_text().strip()
        if not text:
            continue
        # 遇到粗体行,就把当前块存入列表,开始新块
        if is_bold(text_obj):
            if current_block:
                data_blocks.append("\n".join(current_block))
                current_block = []
            current_block.append(text)
        else:
            current_block.append(text)
    # 加入最后一个数据块
    if current_block:
        data_blocks.append("\n".join(current_block))
    return data_blocks

def parse_block_to_list(block):
    # 解析单个数据块,转换成你期望的列表格式
    field_patterns = [
        (r'^([A-Z0-9\s]+(LLC|INC|CO))$', None),  # 公司名称(粗体行)
        (r'Premises No.:\s*(.*)', 'Premises No.: {}'),
        (r'License Key:\s*(.*)', 'License Key: {}'),
        (r'Date Entered:\s*(\d{2}/\d{2}/\d{4})', 'Date Entered:{}'),
        (r'Tradename:\s*(.*?)(?=\s*Date Received|$)', 'Tradename: {}'),
        (r'Email Address:\s*(.*?)(?=\s*License Type|$)', 'Email Address: {}'),
        (r'License Type/Action:\s*(.*)', 'License Type/Action: {}')
    ]
    result = []
    # 先加入公司名称
    name_match = re.match(field_patterns[0][0], block.split('\n')[0], re.IGNORECASE)
    if name_match:
        result.append(name_match.group(1).strip())
    # 匹配其他字段
    for pattern, fmt in field_patterns[1:]:
        match = re.search(pattern, block)
        if match:
            result.append(fmt.format(match.group(1).strip()))
    return result

# 主逻辑:写入CSV并输出目标列表
with open('data.csv', 'w', newline='', encoding='utf-8') as csv_file:
    csv_writer = csv.writer(csv_file)
    csv_headers = [
        'Name', 'Date Entered', 'Tradename', 
        'Address', 'Email Address', 'License Type/Action'
    ]
    csv_writer.writerow(csv_headers)

    with pdfplumber.open('pdf_file.pdf') as pdf:
        # 处理第一页(对应你之前的pdf[0])
        page = pdf.pages[0]
        # 从第314个字符开始提取(对应原代码的pdf[0][313:])
        target_objs = extract_target_text_objects(page, 313)
        # 拆分数据块
        data_blocks = split_into_data_blocks(target_objs)
        # 处理每个数据块
        for block in data_blocks:
            # 转换成你想要的列表格式
            output_list = parse_block_to_list(block)
            print(output_list)
            # 提取CSV需要的字段(从output_list中解析)
            csv_fields = {}
            for item in output_list:
                if item.startswith('Date Entered:'):
                    csv_fields['Date Entered'] = item.split(':')[1].strip()
                elif item.startswith('Tradename:'):
                    csv_fields['Tradename'] = item.split(':')[1].strip()
                elif item.startswith('Email Address:'):
                    csv_fields['Email Address'] = item.split(':')[1].strip()
                elif item.startswith('License Type/Action:'):
                    csv_fields['License Type/Action'] = item.split(':')[1].strip()
                elif 'LLC' in item or 'INC' in item:
                    csv_fields['Name'] = item.strip()
                # 地址需要单独匹配
                address_match = re.search(r'Address:\s*(.*?)(?=\s*Email Address|$)', block)
                if address_match:
                    csv_fields['Address'] = address_match.group(1).strip()
            # 按表头顺序写入CSV
            csv_row = [csv_fields.get(header, '') for header in csv_headers]
            csv_writer.writerow(csv_row)

代码解释

  1. is_bold:检查文本对象的字体是否包含"Bold",判断是否为粗体行。
  2. extract_target_text_objects:从PDF页面的指定位置(第314个字符)开始提取文本对象,过滤掉前面不需要的内容。
  3. split_into_data_blocks:遍历文本对象,遇到粗体行就拆分出一个新的数据块,确保每个数据块都是独立的记录。
  4. parse_block_to_list:用正则从数据块中提取各个字段,转换成你期望的列表格式。
  5. 主逻辑:把处理好的数据写入CSV,同时打印出你想要的列表。

这样处理后,就能精准地按PDF里的粗体换行符拆分数据,再也不用纠结字段混排的问题啦!

内容的提问来源于stack exchange,提问作者codewithawais

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 11:57:41