You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何抓取HTML下拉按钮下方的PDF文件?

如何抓取HTML下拉按钮下方的PDF文件?

问题出在你用requests获取的是页面初始静态HTML,而页面里的下拉折叠面板(就是你提到的collapsible类按钮对应的内容)是通过JavaScript动态渲染的——这类内容不会在初始请求里返回,所以BeautifulSoup根本看不到折叠面板里的PDF链接。

要解决这个问题,你需要用能模拟浏览器交互的工具,比如Selenium,它可以像真实用户一样点击按钮、等待页面加载动态内容,之后再抓取链接。下面是完整的解决方案:

问题根源分析

你原代码的逻辑对静态页面完全有效,但这个页面用了Angular框架(从_ngcontent-pir-c40这类类名能看出来),折叠面板里的内容是用户点击按钮后才通过JS动态渲染到页面上的。requests只能拿到页面首次加载的静态代码,自然抓不到这些隐藏的PDF链接。

完整解决方案代码

import requests
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

# 初始化Chrome浏览器(确保ChromeDriver版本和你的浏览器匹配,可配置到环境变量或指定路径)
driver = webdriver.Chrome()
url = "https://cjj.gob.mx/fraction;article=2;fraction=XIII;subsection=109"
driver.get(url)

try:
    # 等待折叠按钮加载完成,最多等待10秒
    WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.CLASS_NAME, "collapsible"))
    )
    
    # 逐个点击折叠按钮,展开所有隐藏内容
    collapsible_buttons = driver.find_elements(By.CLASS_NAME, "collapsible")
    for button in collapsible_buttons:
        # 用JS点击避免按钮被页面元素遮挡的问题
        driver.execute_script("arguments[0].click();", button)
        time.sleep(0.5)  # 给页面一点时间加载展开的内容

    # 此时页面已加载所有动态内容,提取完整HTML
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, 'html.parser')

    # 筛选PDF链接,处理相对路径问题
    base_domain = "https://cjj.gob.mx"
    valid_pdf_links = []
    for link in soup.find_all('a'):
        href = link.get('href', '')
        if '.pdf' in href:
            # 补全相对路径为完整URL
            if not href.startswith('http'):
                href = f"{base_domain}{href}"
            valid_pdf_links.append(href)

    # 批量下载PDF文件
    for index, pdf_url in enumerate(valid_pdf_links, 1):
        print(f"开始下载文件: {index}")
        try:
            pdf_response = requests.get(pdf_url)
            pdf_response.raise_for_status()  # 捕获请求失败的情况
            with open(f"downloaded_pdf_{index}.pdf", "wb") as pdf_file:
                pdf_file.write(pdf_response.content)
            print(f"文件 downloaded_pdf_{index}.pdf 下载完成")
        except Exception as e:
            print(f"下载 {pdf_url} 失败: {str(e)}")

finally:
    # 不管成功失败,最后关闭浏览器
    driver.quit()

额外提示

  1. 你原代码里有个拼写错误:responde.content应该是response.content,这个小问题会导致下载时报错
  2. Selenium需要对应浏览器的驱动(比如ChromeDriver),一定要确保驱动版本和你的浏览器版本完全一致
  3. 如果页面加载较慢,可以把time.sleep(0.5)的时间适当延长,或者用更智能的显式等待替代固定休眠
  4. 加入异常捕获是为了避免单个链接失败导致整个程序中断,提升代码健壮性

备注:内容来源于stack exchange,提问作者aimee prieto

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 11:27:59