You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的Selenium爬取GeeksforGeeks多页搜索结果

Selenium爬取GeeksforGeeks多页搜索结果的解决方案

获取总页数

GeeksforGeeks的分页栏提供两种获取总页数的途径,根据页面实际结构选择:

  • 直接提取总页数文本
    找到分页栏里显示“共X页”的元素,提取数字即可:
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

wait = WebDriverWait(driver, 10)
total_text = wait.until(EC.visibility_of_element_located((By.CLASS_NAME, "ant-pagination-total-text"))).text
total_pages = int(total_text.split()[1]) # 适配"共 N 页"格式的文本
  • 通过页码按钮获取总页数
    如果页面未显示总页数文本,直接取最后一个页码按钮的文本:
page_items = wait.until(EC.visibility_of_all_elements_located((By.CLASS_NAME, "ant-pagination-item")))
total_pages = int(page_items[-1].text)

实现Next按钮翻页导航

你当前的定位方式可能不够精准,且未处理按钮禁用场景,优化方案如下:

  1. 精准定位Next按钮
    GeeksforGeeks的Next按钮嵌套在ant-pagination-next类的li标签内,用XPATH定位更可靠:
next_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//li[@class='ant-pagination-next']/button[@class='ant-pagination-item-link']")))
  1. 循环翻页逻辑
    结合总页数实现循环爬取,同时判断按钮是否处于禁用状态:
current_page = 1
while current_page < total_pages:
    # 此处编写当前页数据爬取逻辑
    # ...
    
    # 点击下一页并等待页面切换完成
    next_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//li[@class='ant-pagination-next']/button[@class='ant-pagination-item-link']")))
    parent_li = next_btn.find_element(By.XPATH, "./..")
    if "ant-pagination-disabled" not in parent_li.get_attribute("class"):
        next_btn.click()
        wait.until(EC.text_to_be_present_in_element((By.CLASS_NAME, "ant-pagination-item-active"), str(current_page + 1)))
        current_page += 1
    else:
        break

额外提醒

  • 爬取时添加适当延迟(比如time.sleep(1))或随机间隔,避免触发反爬机制。
  • 页面结构可能更新,定期检查元素定位器的有效性。

内容的提问来源于stack exchange,提问作者Marwa Mriwa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 18:55:24