You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Selenium从网页表格中按序提取PDF链接

提取PDF链接的最简实现

基于你已有的代码,只需在获取表格元素后,筛选并收集所有指向PDF的链接即可,以下是最简写法:

driver.get(company_link)
announcement_link = driver.find_element(By.XPATH, '//*[@id="heading1"]/h1/a').get_attribute('href')
driver.get(announcement_link)
table = driver.find_element(By.XPATH, '//*[@id="lblann"]/table/tbody/tr[4]/td')

# 用列表推导式直接生成PDF链接列表(保持页面顺序)
pdf_links = [
    link.get_attribute('href') 
    for link in table.find_elements(By.TAG_NAME, 'a') 
    if link.get_attribute('href') and link.get_attribute('href').endswith('.pdf')
]

# 输出结果
print(pdf_links)

关键说明

  • 使用find_elements(By.TAG_NAME, 'a')获取表格内所有链接元素(复数形式确保拿到全部条目)。
  • 通过endswith('.pdf')筛选出PDF文件链接,同时判断href不为空避免空值报错。
  • 列表推导式让代码更简洁,且生成的列表顺序与页面表格中的链接展示顺序完全一致。

内容的提问来源于stack exchange,提问作者KawaiKx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 00:20:33