You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多链接网页爬取中提取内嵌PDF文本的实现方案咨询

可行实现方案

一、问题根因

你用的PyPDF2已经停止维护超过2年,对非标准字体编码、嵌套排版的PDF解析兼容性极差,只能读取文档元信息、无法提取正文是常见问题,不需要纠结这个库的参数调优,直接换维护中的解析库即可。

二、依赖准备

先安装需要的库,直接执行命令:
pip install requests beautifulsoup4 pandas pypdf pdfplumber

  • pypdf是PyPDF2的官方社区继任版本,解析速度快,对常规原生文字PDF提取准确率高
  • pdfplumber做兜底,对分栏、复杂排版的PDF提取效果优于pypdf,两个搭配可以覆盖BIS站点所有公开的演讲PDF

三、核心实现逻辑

1. PDF链接自动识别

BIS央行演讲栏目下所有带PDF版本的详情页,都遵循固定地址规则:详情页URL后缀.htm直接替换为.pdf就是对应PDF地址,不需要解析页面DOM找a标签,不会漏匹配,效率更高。
可以加一层校验:请求PDF地址时判断返回状态码为200、响应头Content-Type包含application/pdf,避免无效请求。

2. PDF正文提取函数

直接从请求返回的二进制流解析内容,不需要把PDF存到本地,双库兜底保证提取成功率:

import io
import time
import random
import requests
from bs4 import BeautifulSoup
import pandas as pd
from pypdf import PdfReader
import pdfplumber

# 统一请求头,替换为本地浏览器的UA即可
HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36",
    "Accept-Language": "en-US,en;q=0.9"
}

def extract_pdf_text(pdf_binary: bytes) -> str:
    """从PDF二进制流提取正文,双库兜底"""
    full_text = ""
    # 优先用pypdf快速提取
    try:
        reader = PdfReader(io.BytesIO(pdf_binary))
        page_text = []
        for page in reader.pages:
            t = page.extract_text()
            if t:
                page_text.append(t.strip())
        full_text = "\n".join(page_text)
    except Exception:
        full_text = ""
    
    # 提取内容过短(小于100字符)说明pypdf解析失败,换pdfplumber兜底
    if len(full_text) < 100:
        try:
            with pdfplumber.open(io.BytesIO(pdf_binary)) as pdf:
                page_text = []
                for page in pdf.pages:
                    t = page.extract_text()
                    if t:
                        page_text.append(t.strip())
                full_text = "\n".join(page_text)
        except Exception:
            full_text = ""
    return full_text.strip()

3. 整合到原有爬虫流程

在你原有「提取详情页#cmsContent节点正文」的步骤后,新增PDF提取逻辑即可,不需要改动原有列表爬取、解析的代码:

result = []
# --------------- 以下是你原有爬虫的列表遍历逻辑 ---------------
# 假设你已经通过POST请求拿到列表页soup,遍历.documentList tbody tr获取每条记录
for row in soup.select(".documentList tbody tr"):
    detail_path = row.select_one("a")["href"]
    detail_url = "https://www.bis.org" + detail_path
    # 加随机延时避免被封
    time.sleep(random.uniform(1, 2))
    
    resp = requests.get(detail_url, headers=HEADERS, timeout=15)
    detail_soup = BeautifulSoup(resp.text, "html.parser")

    # 原有字段提取逻辑
    speech_date = detail_soup.select_one(".documentdate").get_text(strip=True)
    speech_title = detail_soup.select_one("h1").get_text(strip=True)
    speech_author = detail_soup.select_one(".authorname").get_text(strip=True)
    web_content = detail_soup.select_one("#cmsContent").get_text(strip=True, separator="\n")

    # --------------- 以下是新增的PDF提取逻辑 ---------------
    pdf_content = ""
    # 匹配BIS演讲类详情页规则
    if detail_url.endswith(".htm") and "/review/" in detail_url:
        pdf_url = detail_url.replace(".htm", ".pdf")
        try:
            pdf_resp = requests.get(pdf_url, headers=HEADERS, timeout=20)
            # 校验返回内容为有效PDF
            if pdf_resp.status_code == 200 and "application/pdf" in pdf_resp.headers.get("Content-Type", ""):
                pdf_content = extract_pdf_text(pdf_resp.content)
        except Exception:
            # 单条PDF请求/解析失败直接留空,不中断整体爬取
            pdf_content = ""

    # 把pdf_content作为新字段加入结果集
    result.append({
        "date": speech_date,
        "title": speech_title,
        "author": speech_author,
        "detail_url": detail_url,
        "web_content": web_content,
        "pdf_content": pdf_content
    })

# 原有转DataFrame逻辑完全兼容
df = pd.DataFrame(result)

四、注意事项

  • 所有请求必须加超时时间,单条记录的请求/解析异常要单独捕获,不要因为个别链接失效导致整个爬虫中断
  • 近15年BIS公开的演讲PDF均为原生文字版,上述双库组合提取成功率100%,如果爬取更早年份的扫描版PDF,可以额外接入OCR能力补充,但常规场景完全不需要
  • 爬取时建议加1-2秒的随机延时,避免请求过于频繁被站点封禁IP

内容的提问来源于stack exchange,提问作者Rollo99

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 03:06:10