You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python与NLP从多PDF提取关键词及关联人名的技术求助

需求与现有实现

需要从多个PDF文件中提取信息,输出包含文件名、PDF文本内容、指定关键词检测结果、关联人名的DataFrame。目前已完成前三个字段的代码开发,需借助NLP工具实现:当检测到关键词“Chair Man”时,提取距离该关键词最近的人名并填入DataFrame。

现有代码

import pandas as pd
import numpy as np
from tkinter import filedialog
from pathlib import Path
import os

import PyPDF2


file_path = filedialog.askdirectory()
file_list = Path(file_path + '/').rglob('*.pdf')

file_names = os.listdir(file_path)
file_names = [file for file in file_names if file.endswith('.pdf')]
print(len(file_names))

file_df = []
context = []

for file in file_list:
    print(file.stem)
    file_name = file.stem

    pdf = PyPDF2.PdfFileReader(file)
    page_num = pdf.getNumPages()
    text = ''
    for i in range(0, page_num):
        PgOb = pdf.getPage(i)
        text += PgOb.extractText()
    # print(text)
    file_df.append(file_name)
    context.append(text)

pdf_df = pd.DataFrame({'File': file_df, 'Context': context})


pdf_df['Chair_Man_Check'] = np.nan
pdf_df.loc[pdf_df['Context'].str.contains('Chair Man'), 'Chair_Man_Check'] = 'Chair Man  - Yes'
print(pdf_df)

期望输出示例

File_Name                    Context  Chair_Man_Check          Name
  xxx1.pdf  file context blah blah...   Chair Man -Yes   James Volks
  xxx2.pdf                  blah.....              NaN           NaN
  xxx3.pdf                   blahhhhh  Chair Man - Yes  Tim Aloepman

解决方案

使用Spacy的命名实体识别(NER)功能提取人名,核心逻辑是定位文本中所有“Chair Man”的位置,提取所有人名实体后,计算两者的距离并选取最近的人名。

步骤1:安装依赖

先安装Spacy和英文预训练模型:

pip install spacy
python -m spacy download en_core_web_sm

步骤2:修改后的完整代码

import pandas as pd
import numpy as np
from tkinter import filedialog
from pathlib import Path
import os
import PyPDF2
import spacy

# 加载Spacy英文模型
nlp = spacy.load("en_core_web_sm")

def get_closest_name(text, keyword="Chair Man"):
    if keyword not in text:
        return np.nan
    
    # 获取关键词的所有起始位置
    keyword_positions = []
    start_idx = 0
    while True:
        idx = text.find(keyword, start_idx)
        if idx == -1:
            break
        keyword_positions.append(idx)
        start_idx = idx + len(keyword)
    
    # 提取文本中的所有人名实体
    doc = nlp(text)
    name_entities = []
    for ent in doc.ents:
        if ent.label_ == "PERSON":
            # 存储实体的起始位置、结束位置和文本
            name_entities.append({
                "text": ent.text,
                "start": ent.start_char,
                "end": ent.end_char
            })
    
    if not name_entities:
        return np.nan
    
    # 计算每个人名与所有关键词位置的最小距离,再取整体最小的人名
    min_distance = float('inf')
    closest_name = np.nan
    for name in name_entities:
        # 人名的中间位置
        name_mid = (name["start"] + name["end"]) // 2
        # 找到离这个人名最近的关键词位置
        for kw_pos in keyword_positions:
            # 关键词的中间位置
            kw_mid = kw_pos + len(keyword) // 2
            distance = abs(name_mid - kw_mid)
            if distance < min_distance:
                min_distance = distance
                closest_name = name["text"]
    
    return closest_name

file_path = filedialog.askdirectory()
file_list = Path(file_path + '/').rglob('*.pdf')

file_df = []
context = []

for file in file_list:
    print(file.stem)
    file_name = file.stem

    pdf = PyPDF2.PdfFileReader(file)
    page_num = pdf.getNumPages()
    text = ''
    for i in range(page_num):
        PgOb = pdf.getPage(i)
        text += PgOb.extractText()
    file_df.append(file_name)
    context.append(text)

pdf_df = pd.DataFrame({'File_Name': file_df, 'Context': context})

# 标记关键词检测结果
pdf_df['Chair_Man_Check'] = np.nan
pdf_df.loc[pdf_df['Context'].str.contains('Chair Man'), 'Chair_Man_Check'] = 'Chair Man - Yes'

# 提取关联人名
pdf_df['Name'] = pdf_df['Context'].apply(get_closest_name)

print(pdf_df)

关键逻辑说明

  1. get_closest_name函数:
    • 先判断文本中是否存在目标关键词,不存在则返回NaN
    • 定位所有关键词的出现位置,计算每个关键词的中间点
    • 用Spacy提取所有PERSON类型的实体(人名),计算每个人名的中间点
    • 遍历所有人名和关键词位置,计算两者中间点的距离,选取距离最小的人名
  2. 通过apply方法将函数应用到DataFrame的Context列,生成Name字段

内容的提问来源于stack exchange,提问作者XaviorL

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 16:05:35