如何使用Python Selenium实现分页网站的数据抓取?
解决SEC Edgar搜索结果的分页抓取问题
要实现所有页面数据的抓取,核心是循环识别并点击分页栏的「Next」按钮,直到没有下一页为止。下面是修改后的完整代码,包含分页逻辑:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException, TimeoutException # 初始化浏览器和等待对象 driver = webdriver.Chrome() wait = WebDriverWait(driver, 10) # 目标URL target_url = "https://www.sec.gov/edgar/search/#/dateRange=custom&category=custom&startdt=2023-08-01&enddt=2023-08-01&forms=10-K%252C10-Q%252C8-K" driver.get(target_url) def scrape_current_page(): """抓取当前页面的所有实体名称和文件链接""" # 等待表格行加载完成 wait.until(EC.presence_of_all_elements_located((By.XPATH, '//*[@id="hits"]/table/tbody/tr'))) rows = driver.find_elements(By.XPATH, '//*[@id="hits"]/table/tbody/tr') for row in rows: try: # 提取实体名称 entity_name = row.find_element(By.CLASS_NAME, 'entity-name').text # 点击行内链接打开模态框 element_to_click = row.find_element(By.XPATH, './td[1]/a') element_to_click.click() # 等待文件链接加载并提取 link_element = wait.until(EC.element_to_be_clickable((By.ID, "open-file"))) link = link_element.get_attribute("href") # 关闭模态框 close_button = wait.until(EC.element_to_be_clickable((By.ID, "close-modal"))) close_button.click() # 输出结果 print("Entity Name:", entity_name) print("Link:", link) except (NoSuchElementException, TimeoutException) as e: print(f"处理行时出错: {str(e)}") # 关闭模态框(如果意外打开) try: close_button = driver.find_element(By.ID, "close-modal") close_button.click() except: pass continue # 循环处理所有页面 while True: # 抓取当前页面 scrape_current_page() try: # 查找「Next」按钮,判断是否可点击(未禁用) next_button = wait.until(EC.element_to_be_clickable((By.XPATH, '//button[@aria-label="Next page"]'))) # 检查按钮是否处于禁用状态(通过class判断,SEC的禁用按钮会有「disabled」类) if "disabled" in next_button.get_attribute("class"): break # 点击下一页 next_button.click() # 等待页面切换完成(等待表格内容更新) wait.until(EC.staleness_of(driver.find_element(By.XPATH, '//*[@id="hits"]/table/tbody/tr[1]'))) except (NoSuchElementException, TimeoutException): # 找不到Next按钮或超时,说明已到最后一页 break # 关闭浏览器 driver.quit()
关键说明:
- 分页逻辑:通过定位「Next」按钮(
aria-label="Next page"),检查其是否处于禁用状态,来判断是否还有下一页。 - 页面等待:使用
staleness_of等待上一页的表格行失效,确保新页面内容已经加载完成。 - 异常处理:增加了异常捕获,避免单个行的处理错误导致整个脚本中断,同时处理模态框未正常关闭的情况。
- 代码封装:把单页抓取逻辑封装成函数,让代码结构更清晰。
内容的提问来源于stack exchange,提问作者Ashish Talgotra
相关产品推荐
相关产品推荐

