You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取分页网站URL时翻页失效的问题求助

问题分析与解决方案

你的代码存在几个关键问题导致分页失效,同时还有冗余的写入逻辑,以下是具体修复方案:

核心问题点

  1. 页面未等待加载完成:调用driver.get()后立即获取元素,动态网站可能还未完成新页面渲染,导致读取的仍是上一页内容。
  2. 循环变量混淆+重复写入:外层分页循环用i,内层写入循环也用i(虽不直接影响分页,但逻辑冗余),且嵌套循环会让每个URL重复写入num_page_items次。
  3. 潜在的缓存问题:部分网站会缓存页面,URL变更后仍加载旧内容。

修复步骤与完整代码

1. 添加显式等待确保页面加载

使用Selenium的显式等待,等待目标元素出现后再执行后续操作,保证页面渲染完成。

2. 修正循环逻辑与写入方式

替换重复的嵌套循环,用csv.writer规范写入,避免重复和编码问题。

3. 禁用浏览器缓存(可选)

防止网站缓存导致页面不更新。

import csv
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.webdriver.firefox.options import Options

MAX_PAGE_NUM = 3
MAX_PAGE_DIG = 1

# 禁用浏览器缓存,避免旧页面残留
options = Options()
options.set_preference("browser.cache.disk.enable", False)
options.set_preference("browser.cache.memory.enable", False)
options.set_preference("browser.cache.offline.enable", False)
options.set_preference("network.http.use-cache", False)

driver = webdriver.Firefox(options=options)

for page_idx in range(1, MAX_PAGE_NUM + 1):
    # 生成分页编号
    page_num = (MAX_PAGE_DIG - len(str(page_idx))) * '0' + str(page_idx)
    target_url = f"https://www.example.com/user/learn/freehelp/dynTest/1/Landing/1/page{page_num}"
    driver.get(target_url)
    
    # 等待目标元素加载,超时时间10秒
    try:
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.XPATH, '//div[@class="col-md-12"]/a'))
        )
    except Exception as e:
        print(f"第 {page_idx} 页加载超时,跳过: {e}")
        continue
    
    # 获取页面所有目标链接
    find_href = driver.find_elements(By.XPATH, '//div[@class="col-md-12"]/a')
    
    # 写入CSV文件,避免重复和编码问题
    with open('links1.csv', 'a', newline='', encoding='utf-8') as f:
        writer = csv.writer(f)
        for my_href in find_href:
            href = my_href.get_attribute("href")
            if href:  # 过滤空链接
                writer.writerow([href])

driver.close()

额外排查建议

手动在浏览器中访问page2、page3的URL,确认是否真的能跳转到对应分页。如果手动访问也显示第一页,说明你拼接的URL规则错误,需要重新分析网站的分页参数(比如可能是?page=2而非page2格式)。

内容的提问来源于stack exchange,提问作者GennieM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 02:05:23