使用Beautiful Soup批量爬取PDF异常:仅下载少量旧文件
问题分析
- 动态内容加载:目标网站的PDF列表是动态渲染的,初始GET请求仅返回页面框架和少量历史PDF(2007年),2022年及以后的内容需要通过选择年份筛选器触发AJAX请求才能加载,静态解析HTML无法抓取这部分内容。
- 请求头不完整:直接使用
requests.get()发送请求时,缺少浏览器标识(User-Agent)等必要头信息,可能被网站服务器识别为爬虫,返回不完整内容。
解决方案
1. 模拟浏览器交互获取动态内容
使用selenium模拟浏览器操作,选择年份筛选器并加载对应内容,再提取PDF链接。如果不想依赖浏览器驱动,也可以通过抓包分析AJAX接口,直接请求接口获取数据。
方案一:Selenium模拟交互
import os import requests from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import Select import time def extract_url_pdf(input_url, folder_path='D:/Datos/Ordenanzas municipales/Municipalidad'): # 创建保存目录 if not os.path.exists(folder_path): os.mkdir(folder_path) # 初始化无头浏览器(需提前安装对应浏览器驱动,如ChromeDriver) options = webdriver.ChromeOptions() options.add_argument('--headless=new') driver = webdriver.Chrome(options=options) driver.get(input_url) time.sleep(2) # 等待页面基础框架加载 try: # 定位年份筛选下拉框(需根据页面实际HTML结构调整选择器) year_select = Select(driver.find_element(By.CSS_SELECTOR, 'select[name="anio"]')) # 遍历2022年及以后的年份 for option in year_select.options: year_text = option.text if not year_text.isdigit() or int(year_text) < 2022: continue year_select.select_by_visible_text(year_text) time.sleep(3) # 等待筛选后的内容加载完成 # 提取当前页面所有PDF链接 pdf_links = driver.find_elements(By.CSS_SELECTOR, 'a[href$=".pdf"]') counter = 0 for link in pdf_links: pdf_href = link.get_attribute('href') filename = os.path.join(folder_path, pdf_href.split('/')[-1]) # 带请求头下载PDF,避免被反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } pdf_response = requests.get(pdf_href, headers=headers) with open(filename, 'wb') as f: f.write(pdf_response.content) counter += 1 print(f"{year_text}年 - {counter} 已下载:{filename}") except Exception as e: print(f"执行出错:{str(e)}") finally: driver.quit() extract_url_pdf(input_url="https://munihuamanga.gob.pe/normas-legales/ordenanzas-municipales/")
方案二:抓包AJAX接口直接请求
- 打开浏览器开发者工具(F12),切换到Network标签;
- 选择年份筛选器,观察新出现的XHR请求,找到返回PDF列表的接口(例如带年份参数的请求);
- 直接请求该接口,解析返回的HTML或JSON数据提取PDF链接。
示例代码:
import os import requests from bs4 import BeautifulSoup from urllib.parse import urljoin def extract_url_pdf(folder_path='D:/Datos/Ordenanzas municipales/Municipalidad'): if not os.path.exists(folder_path): os.mkdir(folder_path) # 模拟浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8' } # 遍历2022到2024年(可根据实际调整年份范围) for year in range(2022, 2025): # 替换为抓包得到的实际接口地址 api_url = f"https://munihuamanga.gob.pe/normas-legales/ordenanzas-municipales/?anio={year}" response = requests.get(api_url, headers=headers) response.encoding = 'utf-8' # 解析返回的HTML片段 soup = BeautifulSoup(response.text, 'html.parser') pdf_links = soup.select('a[href$=".pdf"]') counter = 0 for link in pdf_links: pdf_href = urljoin(api_url, link['href']) filename = os.path.join(folder_path, pdf_href.split('/')[-1]) pdf_response = requests.get(pdf_href, headers=headers) with open(filename, 'wb') as f: f.write(pdf_response.content) counter += 1 print(f"{year}年 - {counter} 已下载:{filename}") extract_url_pdf()
2. 优化静态爬取的请求头(快速验证)
如果网站仅因请求头缺失返回不完整内容,可先尝试给原有代码的requests.get()添加完整请求头:
# 修改原有代码中的response请求部分 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8' } response = requests.get(url, headers=headers)
关键注意事项
- 动态内容是核心问题:多数政府网站的列表内容采用AJAX异步加载,静态解析初始HTML只能拿到预渲染的旧数据;
- 反爬规避:必须添加
User-Agent等请求头,同时控制请求频率,避免频繁请求导致IP被封禁; - 元素定位调整:Selenium代码中的元素选择器需根据页面实际HTML结构修改,可通过浏览器开发者工具查看元素属性。
内容的提问来源于stack exchange,提问作者Ivan A. Ramírez Zapata
相关产品推荐
相关产品推荐

