You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDF模板文档爬取问题求助:字段匹配与整合异常

处理PDF文本提取与Excel整合的问题

遇到的问题

  • CASE1:需将Definition and Class Characteristic的各类变体内容(如Definition and Class Characteristics、Definition and Distinguishing Characteristics等)合并至Excel的Job Description列同一单元格
  • CASE2:Python大小写敏感,需处理Characteristic的单复数拼写差异(如Distinguishing CharacteristicS或无复数后缀的Characteristic)
  • CASE3:需将EXAMPLE OF DUTIES内容合并至Excel同一单元格
  • CASE4:Education and Experience列需整合Knowledge of、Ability to及Education and Training的内容

修改后的完整代码

import fitz  # PyMuPDF
import pandas as pd
import re
import os


# 从PDF提取文本
def extract_text_from_pdf(pdf_path):
    document = fitz.open(pdf_path)
    pdf_text = ""

    for page_num in range(len(document)):
        page = document.load_page(page_num)
        pdf_text += page.get_text()

    return pdf_text


# 定义正则表达式,统一处理大小写和单复数差异
# 匹配Job Title
job_title_re = re.compile(r'UNIT:\s*(.*?)\s*Class Specification', re.DOTALL | re.IGNORECASE)
# 匹配Definition + 各类Characteristic变体(处理单复数和大小写)
definition_characteristics_re = re.compile(
    r'DEFINITION\s*(.*?)\s*(CLASS|DISTINGUISHING)\s*CHARACTERISTIC[S]?\s*(.*?)\s*EXAMPLE OF DUTIES',
    re.DOTALL | re.IGNORECASE
)
# 匹配Example of Duties
example_of_duties_re = re.compile(r'EXAMPLE OF DUTIES\s*(.*?)\s*(MINIMUM QUALIFICATIONS|WORKING CONDITIONS)', re.DOTALL | re.IGNORECASE)
# 匹配Knowledge of + Ability to + Experience/Training内容
qualifications_re = re.compile(
    r'Knowledge of:\s*(.*?)\s*Ability to:\s*(.*?)\s*(Experience and Training|Education and Training)\s*(.*?)\s*WORKING CONDITIONS',
    re.DOTALL | re.IGNORECASE
)


def extract_section(text, pattern, group_idx=1):
    match = pattern.search(text)
    if match:
        return match.group(group_idx).strip() if group_idx <= len(match.groups()) else ""
    return ""


def process_pdfs_in_folder(pdf_folder, excel_path):
    all_data = []

    for pdf_file in os.listdir(pdf_folder):
        if pdf_file.endswith('.pdf'):
            pdf_path = os.path.join(pdf_folder, pdf_file)
            pdf_text = extract_text_from_pdf(pdf_path)

            # 提取各部分内容
            job_title = extract_section(pdf_text, job_title_re)
            # 提取Definition + Characteristics合并内容
            def_char_match = definition_characteristics_re.search(pdf_text)
            job_description = ""
            if def_char_match:
                job_description = def_char_match.group(1).strip() + "\n" + def_char_match.group(3).strip()
            # 提取Example of Duties
            example_of_duties = extract_section(pdf_text, example_of_duties_re)
            # 提取Knowledge/Ability/Experience整合内容
            qual_match = qualifications_re.search(pdf_text)
            education_experience = ""
            if qual_match:
                education_experience = qual_match.group(1).strip() + "\n" + qual_match.group(2).strip() + "\n" + qual_match.group(4).strip()

            # 整理DataFrame数据
            data = {
                "Job Title": job_title,
                "Job Description": job_description,
                "Example of Duties": example_of_duties,
                "Education and Experience": education_experience
            }

            all_data.append(data)

    # 创建DataFrame并写入Excel
    new_df = pd.DataFrame(all_data)
    if os.path.exists(excel_path):
        existing_df = pd.read_excel(excel_path)
        final_df = pd.concat([existing_df, new_df], ignore_index=True)
    else:
        final_df = new_df

    final_df.to_excel(excel_path, index=False)
    print(f"数据已写入 {excel_path}")


# 路径配置
pdf_folder = r'C:\Users\donna\PycharmProjects\scrapePdf\pythonProject\.venv\OceansideCA\JobTitle'
excel_path = r'C:\Users\donna\PycharmProjects\scrapePdf\pythonProject\.venv\OceansideCA\JD_Excel.xlsx'

# 执行脚本
process_pdfs_in_folder(pdf_folder, excel_path)

关键修改说明

针对CASE1:合并Definition与Characteristic变体

  • 使用单个正则表达式definition_characteristics_re,通过(CLASS|DISTINGUISHING)匹配不同特征类型,同时捕获Definition和对应特征的内容,直接合并为Job Description的内容,避免多个正则重复提取后拼接。

针对CASE2:处理大小写与单复数差异

  • 正则中使用re.IGNORECASE忽略大小写;
  • 用CHARACTERISTIC[S]?匹配单复数([S]?表示S可选,既匹配Characteristic也匹配Characteristics)。

针对CASE3:合并Example of Duties

  • 优化example_of_duties_re的终止条件,兼容不同PDF的后续章节(比如有些是MINIMUM QUALIFICATIONS,有些可能直接到WORKING CONDITIONS);
  • 提取后直接存入同一单元格,确保内容完整。

针对CASE4:整合Education and Experience列

  • 用单个正则qualifications_re一次性捕获Knowledge of、Ability to、Experience/Education and Training的内容,统一拼接后存入Education and Experience列;
  • 兼容Experience and Training和Education and Training两种表述。

内容的提问来源于stack exchange,提问作者Donna Esperas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 11:24:55