You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用Selenium抓取带分页动态datatable并保留链接的问题

问题核心原因

  • 现有代码仅完成了首次页面加载后的单次解析,没有增加分页点击、翻页后的内容等待逻辑,因此只能拿到第一页数据
  • 打印整页soup对象自然会输出全站HTML,仅解析目标表格节点即可过滤出需要的内容
  • 缺少表格内链接的提取逻辑,原有代码仅提取了节点文本内容

可直接运行的修改后代码

from bs4 import BeautifulSoup
import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as ec
from selenium.webdriver.chrome.options import Options
from selenium.common.exceptions import TimeoutException

chrome_options = Options()
chrome_options.add_argument('--headless')
# Selenium 4+ 无需手动指定chromedriver路径会自动匹配,旧版本可以保留你原来的路径写法
driver = webdriver.Chrome(options=chrome_options)
wait = WebDriverWait(driver, 20)

url = "https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/compliance-actions-and-activities/warning-letters"
driver.get(url)

# 等待表格初始加载完成
wait.until(ec.presence_of_element_located((By.ID, "datatable")))

all_data = []
page_num = 1

while True:
    # 解析当前页表格
    soup = BeautifulSoup(driver.page_source, 'lxml')
    table = soup.find('table', id="datatable")
    table_body = table.find('tbody')
    rows = table_body.find_all('tr')
    
    for row in rows:
        cols = row.find_all('td')
        row_data = []
        for ele in cols:
            text = ele.text.strip()
            # 提取单元格内所有链接
            links = [a['href'] for a in ele.find_all('a', href=True)]
            # 可根据需求调整存储格式,此处同时保留文本和对应链接
            row_data.append({
                "text": text,
                "links": links
            })
        all_data.append(row_data)
    
    print(f"第{page_num}页抓取完成,累计已获取{len(all_data)}条数据")
    
    # 翻页逻辑:点击下一页直到按钮不可用
    try:
        next_btn = wait.until(ec.element_to_be_clickable((By.CSS_SELECTOR, 'li.paginate_button.next:not(.disabled)')))
        next_btn.click()
        # 等待表格刷新,避免读取到上一页缓存内容
        wait.until(ec.staleness_of(table_body))
        page_num += 1
    except TimeoutException:
        # 下一页按钮不可点击,已到最后一页,终止循环
        break

driver.quit()

# 后续可自行将all_data转换为CSV、JSON等格式存储

补充说明

  • 采用下一页按钮定位的方案比直接按data-dt-idx定位页码兼容性更好,不会因为页码过多导致分页栏折叠失效
  • 所有单元格的链接已经统一提取存储在links字段中,可根据实际需求调整存储结构
  • 如果你坚持要按data-dt-idx从索引为2的页码开始遍历,只需要把翻页逻辑替换为按属性遍历即可,注意额外处理分页栏页码折叠的场景

内容的提问来源于stack exchange,提问作者alminmd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 23:45:05