You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup批量爬取PDF异常:仅下载少量旧文件

问题分析
  • 动态内容加载:目标网站的PDF列表是动态渲染的,初始GET请求仅返回页面框架和少量历史PDF(2007年),2022年及以后的内容需要通过选择年份筛选器触发AJAX请求才能加载,静态解析HTML无法抓取这部分内容。
  • 请求头不完整:直接使用requests.get()发送请求时,缺少浏览器标识(User-Agent)等必要头信息,可能被网站服务器识别为爬虫,返回不完整内容。
解决方案

1. 模拟浏览器交互获取动态内容

使用selenium模拟浏览器操作,选择年份筛选器并加载对应内容,再提取PDF链接。如果不想依赖浏览器驱动,也可以通过抓包分析AJAX接口,直接请求接口获取数据。

方案一:Selenium模拟交互

import os
import requests
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import Select
import time

def extract_url_pdf(input_url, folder_path='D:/Datos/Ordenanzas municipales/Municipalidad'):
    # 创建保存目录
    if not os.path.exists(folder_path):
        os.mkdir(folder_path)
    
    # 初始化无头浏览器(需提前安装对应浏览器驱动,如ChromeDriver)
    options = webdriver.ChromeOptions()
    options.add_argument('--headless=new')
    driver = webdriver.Chrome(options=options)
    driver.get(input_url)
    time.sleep(2)  # 等待页面基础框架加载
    
    try:
        # 定位年份筛选下拉框(需根据页面实际HTML结构调整选择器)
        year_select = Select(driver.find_element(By.CSS_SELECTOR, 'select[name="anio"]'))
        # 遍历2022年及以后的年份
        for option in year_select.options:
            year_text = option.text
            if not year_text.isdigit() or int(year_text) < 2022:
                continue
            year_select.select_by_visible_text(year_text)
            time.sleep(3)  # 等待筛选后的内容加载完成
            
            # 提取当前页面所有PDF链接
            pdf_links = driver.find_elements(By.CSS_SELECTOR, 'a[href$=".pdf"]')
            counter = 0
            for link in pdf_links:
                pdf_href = link.get_attribute('href')
                filename = os.path.join(folder_path, pdf_href.split('/')[-1])
                # 带请求头下载PDF,避免被反爬拦截
                headers = {
                    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
                }
                pdf_response = requests.get(pdf_href, headers=headers)
                with open(filename, 'wb') as f:
                    f.write(pdf_response.content)
                counter += 1
                print(f"{year_text}年 - {counter} 已下载:{filename}")
    except Exception as e:
        print(f"执行出错:{str(e)}")
    finally:
        driver.quit()

extract_url_pdf(input_url="https://munihuamanga.gob.pe/normas-legales/ordenanzas-municipales/")

方案二:抓包AJAX接口直接请求

  1. 打开浏览器开发者工具(F12),切换到Network标签;
  2. 选择年份筛选器,观察新出现的XHR请求,找到返回PDF列表的接口(例如带年份参数的请求);
  3. 直接请求该接口,解析返回的HTML或JSON数据提取PDF链接。

示例代码:

import os
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

def extract_url_pdf(folder_path='D:/Datos/Ordenanzas municipales/Municipalidad'):
    if not os.path.exists(folder_path):
        os.mkdir(folder_path)
    
    # 模拟浏览器请求头
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8'
    }
    
    # 遍历2022到2024年(可根据实际调整年份范围)
    for year in range(2022, 2025):
        # 替换为抓包得到的实际接口地址
        api_url = f"https://munihuamanga.gob.pe/normas-legales/ordenanzas-municipales/?anio={year}"
        response = requests.get(api_url, headers=headers)
        response.encoding = 'utf-8'
        
        # 解析返回的HTML片段
        soup = BeautifulSoup(response.text, 'html.parser')
        pdf_links = soup.select('a[href$=".pdf"]')
        
        counter = 0
        for link in pdf_links:
            pdf_href = urljoin(api_url, link['href'])
            filename = os.path.join(folder_path, pdf_href.split('/')[-1])
            pdf_response = requests.get(pdf_href, headers=headers)
            with open(filename, 'wb') as f:
                f.write(pdf_response.content)
            counter += 1
            print(f"{year}年 - {counter} 已下载:{filename}")

extract_url_pdf()

2. 优化静态爬取的请求头(快速验证)

如果网站仅因请求头缺失返回不完整内容,可先尝试给原有代码的requests.get()添加完整请求头:

# 修改原有代码中的response请求部分
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8'
}
response = requests.get(url, headers=headers)
关键注意事项
  • 动态内容是核心问题:多数政府网站的列表内容采用AJAX异步加载,静态解析初始HTML只能拿到预渲染的旧数据;
  • 反爬规避:必须添加User-Agent等请求头,同时控制请求频率,避免频繁请求导致IP被封禁;
  • 元素定位调整:Selenium代码中的元素选择器需根据页面实际HTML结构修改,可通过浏览器开发者工具查看元素属性。

内容的提问来源于stack exchange,提问作者Ivan A. Ramírez Zapata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 01:36:24