Python多链接网页爬取中提取内嵌PDF文本的实现方案咨询
可行实现方案
一、问题根因
你用的PyPDF2已经停止维护超过2年,对非标准字体编码、嵌套排版的PDF解析兼容性极差,只能读取文档元信息、无法提取正文是常见问题,不需要纠结这个库的参数调优,直接换维护中的解析库即可。
二、依赖准备
先安装需要的库,直接执行命令:pip install requests beautifulsoup4 pandas pypdf pdfplumber
pypdf是PyPDF2的官方社区继任版本,解析速度快,对常规原生文字PDF提取准确率高pdfplumber做兜底,对分栏、复杂排版的PDF提取效果优于pypdf,两个搭配可以覆盖BIS站点所有公开的演讲PDF
三、核心实现逻辑
1. PDF链接自动识别
BIS央行演讲栏目下所有带PDF版本的详情页,都遵循固定地址规则:详情页URL后缀.htm直接替换为.pdf就是对应PDF地址,不需要解析页面DOM找a标签,不会漏匹配,效率更高。
可以加一层校验:请求PDF地址时判断返回状态码为200、响应头Content-Type包含application/pdf,避免无效请求。
2. PDF正文提取函数
直接从请求返回的二进制流解析内容,不需要把PDF存到本地,双库兜底保证提取成功率:
import io import time import random import requests from bs4 import BeautifulSoup import pandas as pd from pypdf import PdfReader import pdfplumber # 统一请求头,替换为本地浏览器的UA即可 HEADERS = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36", "Accept-Language": "en-US,en;q=0.9" } def extract_pdf_text(pdf_binary: bytes) -> str: """从PDF二进制流提取正文,双库兜底""" full_text = "" # 优先用pypdf快速提取 try: reader = PdfReader(io.BytesIO(pdf_binary)) page_text = [] for page in reader.pages: t = page.extract_text() if t: page_text.append(t.strip()) full_text = "\n".join(page_text) except Exception: full_text = "" # 提取内容过短(小于100字符)说明pypdf解析失败,换pdfplumber兜底 if len(full_text) < 100: try: with pdfplumber.open(io.BytesIO(pdf_binary)) as pdf: page_text = [] for page in pdf.pages: t = page.extract_text() if t: page_text.append(t.strip()) full_text = "\n".join(page_text) except Exception: full_text = "" return full_text.strip()
3. 整合到原有爬虫流程
在你原有「提取详情页#cmsContent节点正文」的步骤后,新增PDF提取逻辑即可,不需要改动原有列表爬取、解析的代码:
result = [] # --------------- 以下是你原有爬虫的列表遍历逻辑 --------------- # 假设你已经通过POST请求拿到列表页soup,遍历.documentList tbody tr获取每条记录 for row in soup.select(".documentList tbody tr"): detail_path = row.select_one("a")["href"] detail_url = "https://www.bis.org" + detail_path # 加随机延时避免被封 time.sleep(random.uniform(1, 2)) resp = requests.get(detail_url, headers=HEADERS, timeout=15) detail_soup = BeautifulSoup(resp.text, "html.parser") # 原有字段提取逻辑 speech_date = detail_soup.select_one(".documentdate").get_text(strip=True) speech_title = detail_soup.select_one("h1").get_text(strip=True) speech_author = detail_soup.select_one(".authorname").get_text(strip=True) web_content = detail_soup.select_one("#cmsContent").get_text(strip=True, separator="\n") # --------------- 以下是新增的PDF提取逻辑 --------------- pdf_content = "" # 匹配BIS演讲类详情页规则 if detail_url.endswith(".htm") and "/review/" in detail_url: pdf_url = detail_url.replace(".htm", ".pdf") try: pdf_resp = requests.get(pdf_url, headers=HEADERS, timeout=20) # 校验返回内容为有效PDF if pdf_resp.status_code == 200 and "application/pdf" in pdf_resp.headers.get("Content-Type", ""): pdf_content = extract_pdf_text(pdf_resp.content) except Exception: # 单条PDF请求/解析失败直接留空,不中断整体爬取 pdf_content = "" # 把pdf_content作为新字段加入结果集 result.append({ "date": speech_date, "title": speech_title, "author": speech_author, "detail_url": detail_url, "web_content": web_content, "pdf_content": pdf_content }) # 原有转DataFrame逻辑完全兼容 df = pd.DataFrame(result)
四、注意事项
- 所有请求必须加超时时间,单条记录的请求/解析异常要单独捕获,不要因为个别链接失效导致整个爬虫中断
- 近15年BIS公开的演讲PDF均为原生文字版,上述双库组合提取成功率100%,如果爬取更早年份的扫描版PDF,可以额外接入OCR能力补充,但常规场景完全不需要
- 爬取时建议加1-2秒的随机延时,避免请求过于频繁被站点封禁IP
内容的提问来源于stack exchange,提问作者Rollo99
相关产品推荐
相关产品推荐

