You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python Selenium实现分页网站的数据抓取?

解决SEC Edgar搜索结果的分页抓取问题

要实现所有页面数据的抓取,核心是循环识别并点击分页栏的「Next」按钮,直到没有下一页为止。下面是修改后的完整代码,包含分页逻辑:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException, TimeoutException

# 初始化浏览器和等待对象
driver = webdriver.Chrome()
wait = WebDriverWait(driver, 10)

# 目标URL
target_url = "https://www.sec.gov/edgar/search/#/dateRange=custom&category=custom&startdt=2023-08-01&enddt=2023-08-01&forms=10-K%252C10-Q%252C8-K"
driver.get(target_url)

def scrape_current_page():
    """抓取当前页面的所有实体名称和文件链接"""
    # 等待表格行加载完成
    wait.until(EC.presence_of_all_elements_located((By.XPATH, '//*[@id="hits"]/table/tbody/tr')))
    rows = driver.find_elements(By.XPATH, '//*[@id="hits"]/table/tbody/tr')
    
    for row in rows:
        try:
            # 提取实体名称
            entity_name = row.find_element(By.CLASS_NAME, 'entity-name').text
            
            # 点击行内链接打开模态框
            element_to_click = row.find_element(By.XPATH, './td[1]/a')
            element_to_click.click()
            
            # 等待文件链接加载并提取
            link_element = wait.until(EC.element_to_be_clickable((By.ID, "open-file")))
            link = link_element.get_attribute("href")
            
            # 关闭模态框
            close_button = wait.until(EC.element_to_be_clickable((By.ID, "close-modal")))
            close_button.click()
            
            # 输出结果
            print("Entity Name:", entity_name)
            print("Link:", link)
        except (NoSuchElementException, TimeoutException) as e:
            print(f"处理行时出错: {str(e)}")
            # 关闭模态框(如果意外打开)
            try:
                close_button = driver.find_element(By.ID, "close-modal")
                close_button.click()
            except:
                pass
            continue

# 循环处理所有页面
while True:
    # 抓取当前页面
    scrape_current_page()
    
    try:
        # 查找「Next」按钮,判断是否可点击(未禁用)
        next_button = wait.until(EC.element_to_be_clickable((By.XPATH, '//button[@aria-label="Next page"]')))
        # 检查按钮是否处于禁用状态(通过class判断,SEC的禁用按钮会有「disabled」类)
        if "disabled" in next_button.get_attribute("class"):
            break
        # 点击下一页
        next_button.click()
        # 等待页面切换完成(等待表格内容更新)
        wait.until(EC.staleness_of(driver.find_element(By.XPATH, '//*[@id="hits"]/table/tbody/tr[1]')))
    except (NoSuchElementException, TimeoutException):
        # 找不到Next按钮或超时,说明已到最后一页
        break

# 关闭浏览器
driver.quit()

关键说明:

  • 分页逻辑:通过定位「Next」按钮(aria-label="Next page"),检查其是否处于禁用状态,来判断是否还有下一页。
  • 页面等待:使用staleness_of等待上一页的表格行失效,确保新页面内容已经加载完成。
  • 异常处理:增加了异常捕获,避免单个行的处理错误导致整个脚本中断,同时处理模态框未正常关闭的情况。
  • 代码封装:把单页抓取逻辑封装成函数,让代码结构更清晰。

内容的提问来源于stack exchange,提问作者Ashish Talgotra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 04:12:47