You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从非结构化PDF爬取Plum Book数据并实现结构化处理

处理《Plum Book》半结构化数据:提取部门/办公室与职位关联的方案

针对2016年前《Plum Book》的PDF/TXT(实际是HTML页面)半结构化数据,要实现职位与所属部门、办公室的关联,这里提供几个实用的Python方案,帮你避免手动处理大量条目:

优先从HTML版入手(你的“TXT”链接实际是HTML页面)

govinfo提供的HTML版自带基础结构,比PDF和纯文本更容易定位层级,是最高效的处理途径。

Python实现步骤(基于BeautifulSoup)

  1. 解析HTML页面,抓取核心内容区域;
  2. 遍历页面元素,实时追踪当前所属的部门和办公室;
  3. 将职位条目与当前的部门、办公室绑定,导出为CSV(方便Excel后续填充)。
import requests
from bs4 import BeautifulSoup
import csv

# 抓取2012版页面内容
url = "https://www.govinfo.gov/content/pkg/GPO-PLUMBOOK-2012/html/GPO-PLUMBOOK-2012.htm"
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# 定位主内容容器(可根据页面源码调整标签/class)
content_block = soup.find('div', class_='entry-content')

output_data = []
current_agency = None
current_office = None

# 遍历内容块内的所有元素
for elem in content_block.children:
    # 识别部门标题(通常是<h2>标签,可根据实际页面调整)
    if elem.name == 'h2':
        current_agency = elem.get_text(strip=True)
        current_office = None  # 切换部门后重置办公室
    # 识别办公室标题(可能是<h4>或带加粗的<p>标签,需根据页面调整)
    elif elem.name == 'h4' or (elem.name == 'p' and elem.find('strong')):
        current_office = elem.get_text(strip=True)
    # 识别职位条目(通常是<li>标签)
    elif elem.name == 'li':
        job_detail = elem.get_text(strip=True)
        # 仅当部门和办公室信息完整时存入数据
        if current_agency and current_office:
            output_data.append({
                'agency': current_agency,
                'office': current_office,
                'job': job_detail
            })

# 导出为CSV文件
with open('plum_book_2012.csv', 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.DictWriter(csv_file, fieldnames=['agency', 'office', 'job'])
    writer.writeheader()
    writer.writerows(output_data)

注意:如果页面里的办公室标题标签不同(比如是<h3>而非<h4>),需要手动查看页面源码调整判断条件。导出的CSV中,每个办公室的所有职位都会关联对应办公室名称;如果有遗漏,你可以用Excel的填充功能快速补全空白行。

PDF版备选方案

如果HTML版无法使用,可借助pdfplumber提取文本并通过字体大小区分层级:

  • 部门标题通常字体最大,办公室标题次之,职位条目字体最小;
  • 同样通过维护current_agency和current_office变量,绑定职位与所属层级。
import pdfplumber
import csv

output_data = []
current_agency = None
current_office = None

# 读取PDF文件(需先下载到本地)
with pdfplumber.open("GPO-PLUMBOOK-2012.pdf") as pdf:
    for page in pdf.pages:
        # 提取页面所有文本行
        lines = page.extract_text().split('\n')
        # 获取每个文字的字体信息(用于判断层级)
        word_details = page.extract_words()
        
        for line in lines:
            line_clean = line.strip()
            if not line_clean:
                continue
            
            # 匹配当前行的字体大小
            line_font_size = None
            for word in word_details:
                if word['text'] in line_clean:
                    line_font_size = word['size']
                    break
            if not line_font_size:
                continue
            
            # 根据字体大小判断层级(需根据实际PDF调整阈值)
            if line_font_size >= 14:  # 部门标题字体阈值
                current_agency = line_clean
                current_office = None
            elif line_font_size >= 12:  # 办公室标题字体阈值
                current_office = line_clean
            else:  # 职位条目
                if current_agency and current_office:
                    output_data.append({
                        'agency': current_agency,
                        'office': current_office,
                        'job': line_clean
                    })

# 导出CSV
with open('plum_book_2012_pdf.csv', 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.DictWriter(csv_file, fieldnames=['agency', 'office', 'job'])
    writer.writeheader()
    writer.writerows(output_data)

注意:字体大小阈值需要你用pdfplumber查看实际PDF的字体参数后调整,确保能准确区分标题和职位行。

关于MODS/PREMIS格式的说明

这两种是书籍元数据格式,仅包含出版日期、作者等基础信息,不涉及职位条目内容,对你的项目没有帮助,无需投入时间研究。

内容的提问来源于stack exchange,提问作者Drew Godsell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 22:15:07