You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+Selenium抓取网站动态生成的隐藏PDF链接?

修复Selenium抓取动态PDF链接的方案

原脚本核心问题

  • 绝对XPATH定位脆弱:页面结构稍有变动就会失效
  • 属性获取逻辑错误:get_attribute("dvg-link-doc dvg-certidao")是错误用法,这两个是class属性值,并非单个属性名
  • WebDriverWait使用不当:until方法需接收符合expected_conditions规范的条件,不能直接传自定义函数
  • 变量名不统一:定义了documentss却调用documents.append,存在语法错误
  • 未等待动态加载:直接查找元素会因页面未加载完成导致元素未找到

修复后的脚本实现

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

site = "https://divulgacandcontas.tse.jus.br/divulga/#/candidato/2022/2040602022/AP/30001653385"

# 初始化浏览器(注意:ChromeDriver版本需匹配Chrome版本,建议用Service类替代直接传路径)
driver = webdriver.Chrome()
driver.get(site)
documents = []

try:
    # 等待"Documentos"区域的Certidão链接列表加载完成
    # 用相对XPATH+class定位,稳定性更强
    certidao_links = WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located(
            (By.XPATH, '//a[contains(@class, "dvg-link-doc") and contains(@class, "dvg-certidao")]')
        )
    )

    # 遍历链接提取PDF真实地址
    for link in certidao_links:
        # 动态生成的PDF链接通常存在data-url属性中
        pdf_url = link.get_attribute("data-url")
        if pdf_url and pdf_url.endswith(".pdf"):
            print(f"找到PDF链接: {pdf_url}")
            documents.append({"link": pdf_url})

finally:
    driver.quit()

# 输出所有抓取结果
print("\n所有抓取到的PDF链接:")
for doc in documents:
    print(doc["link"])

关键修复说明

  • 定位优化:改用相对XPATH结合class属性定位,避免绝对路径的脆弱性
  • 等待机制修正:用WebDriverWait配合presence_of_all_elements_located等待动态元素加载,超时时间设为10秒更合理
  • 正确提取链接:动态PDF链接一般存储在元素的data-url属性中,通过get_attribute("data-url")获取真实地址
  • 变量统一:修正documentss为documents,解决语法错误
  • 资源清理:用try...finally确保浏览器进程正常关闭

替代方案:监听网络请求

如果元素属性无法获取链接,可通过监听浏览器网络请求捕获PDF地址:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.options import Options

# 配置Chrome监听网络请求
chrome_options = Options()
chrome_options.add_argument("--log-level=3")
driver = webdriver.Chrome(options=chrome_options)
driver.get(site)

try:
    # 等待Certidão链接加载完成并点击第一个示例链接
    certidao_links = WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located(
            (By.XPATH, '//a[contains(@class, "dvg-link-doc") and contains(@class, "dvg-certidao")]')
        )
    )
    certidao_links[0].click()

    # 筛选网络请求中的PDF链接
    logs = driver.get_log("performance")
    pdf_links = []
    for log in logs:
        log_entry = eval(log["message"])["message"]
        if "Network.responseReceived" in log_entry["method"]:
            url = log_entry["params"]["response"]["url"]
            if url.endswith(".pdf") and url not in pdf_links:
                pdf_links.append(url)
    
    print("捕获到的PDF链接:")
    for link in pdf_links:
        print(link)

finally:
    driver.quit()

内容的提问来源于stack exchange,提问作者Reinaldo Chaves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 04:48:29