You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Selenium爬取联邦议院网站动态PDF链接遇问题

问题:德国联邦议院官网PDF链接爬取重复,翻页后无法获取新链接

我尝试从德国联邦议院官网爬取PDF文件链接,但用Python Selenium写的代码只能获取首个PDF链接,翻页后还是重复添加这个链接,没法抓取每页的所有PDF链接。

我的代码:

import time

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome("C:/Program Files/Chrome Driver/chromedriver.exe")

driver.get('https://www.bundestag.de/protokolle')
button_count = 0
pdf_links = []  # Initialize empty list to store links
mywaits = 100
for a in range(10):
     WebDriverWait(driver, mywaits).until(EC.presence_of_element_located((By.CSS_SELECTOR,
    "div.slick-initialized")))
     elements = driver.find_elements(By.CSS_SELECTOR, "div.slick-initialized")
     for element in elements:
         pdf_link = element.find_element(By.CSS_SELECTOR, "a")
         pdf_links.append(pdf_link.get_attribute('href'))  # Append the href attribute to list
     WebDriverWait(driver, mywaits).until(EC.element_to_be_clickable((By.CSS_SELECTOR,
    "div.slick-initialized")))
     button = driver.find_element(By.CSS_SELECTOR, "div.slick-initialized")
     button.click()
     print(pdf_links)
     button_count += 1
     time.sleep(1)  # Add a short delay to allow the new content to load
driver.close()

print(pdf_links)

输出结果:

['https://dserver.bundestag.de/btp/20/20090.pdf']
['https://dserver.bundestag.de/btp/20/20090.pdf', 'https://dserver.bundestag.de/btp/20/20090.pdf']
['https://dserver.bundestag.de/btp/20/20090.pdf', 'https://dserver.bundestag.de/btp/20/20090.pdf', 'https://dserver.bundestag.de/btp/20/20090.pdf']...

网站界面中,红色标记为要抓取的文档,绿色标记为翻页按钮。


解决方案

问题核心是CSS选择器定位错误:

  1. 你用div.slick-initialized定位的是整个轮播容器,不是单个PDF条目,所以每次只能拿到容器里的第一个链接;
  2. 翻页按钮的选择器也错了,点击轮播容器根本不会触发翻页,页面没变化自然一直拿到同一个链接。

修改后的代码:

import time
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome("C:/Program Files/Chrome Driver/chromedriver.exe")
driver.get('https://www.bundestag.de/protokolle')

pdf_links = []
mywaits = 20

for a in range(10):
    # 等待当前页面所有PDF条目加载完成
    WebDriverWait(driver, mywaits).until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.slick-slide > div.document-teaser"))
    )
    # 精准定位所有PDF下载链接
    pdf_elements = driver.find_elements(By.CSS_SELECTOR, "div.document-teaser a.document-teaser__download")
    for elem in pdf_elements:
        link = elem.get_attribute('href')
        if link not in pdf_links:  # 去重避免重复添加
            pdf_links.append(link)
    
    # 定位并点击真正的下一页按钮
    try:
        next_btn = WebDriverWait(driver, mywaits).until(
            EC.element_to_be_clickable((By.CSS_SELECTOR, "button.slick-next"))
        )
        next_btn.click()
        time.sleep(1)  # 等待页面切换完成
    except:
        print("已到达最后一页,停止抓取")
        break

# 输出所有抓取到的链接
print("\n抓取到的PDF链接:")
for idx, link in enumerate(pdf_links, 1):
    print(f"{idx}. {link}")

driver.close()

关键修改说明:

  • 用div.slick-slide > div.document-teaser定位单个PDF条目容器,确保能获取当前页面所有PDF元素;
  • 用div.document-teaser a.document-teaser__download精准定位PDF下载链接,避免误抓其他无关链接;
  • 翻页按钮改为button.slick-next,这才是页面上绿色标记的翻页按钮对应的选择器;
  • 添加去重逻辑,防止同一链接被多次添加;
  • 缩短等待时间到20秒,避免不必要的长时间等待,同时添加异常处理,遇到最后一页时自动停止。

内容的提问来源于stack exchange,提问作者corvusMidnight

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 12:45:11