You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium while循环爬取列表仅保留首次迭代结果

问题根因

你遇到的问题核心是三个代码逻辑缺陷,刚好和调试/直接运行的场景差异匹配:

  • 定位语法写错:By.CLASS_NAME仅支持传入单个类名,你传入了带.分隔的CSS选择器格式的复合类字符串,翻页后DOM结构轻微变化时就会定位不到元素。
  • 异常捕获范围太窄:你只捕获了TimeoutException,但第一次点击下一页时,按钮在页面底部,会被亚马逊固定在顶部的导航栏遮挡,直接触发ElementClickInterceptedException(点击拦截异常);另外翻页时DOM刷新还可能触发StaleElementReferenceException(陈旧元素引用异常),这两类异常都没被捕获,会直接终止程序——而第一次循环的逻辑是先存第一页数据,再点下一页,所以程序崩溃时列表里刚好只存了第一页的结果。
  • 翻页无稳定等待:点击下一页后没有等新页面加载完成就进入下一轮查找,直接运行时速度快,会拿到未刷新的旧DOM或者空DOM;调试时单步执行速度慢,页面有足够时间加载,所以能看到列表长度正常增长。
修复方案
  1. 把商品名的复合类定位从By.CLASS_NAME改成By.CSS_SELECTOR,匹配Selenium的定位语法规则
  2. 点击下一页按钮前,先用JS把按钮滚动到避开顶部导航栏的位置,避免点击被遮挡
  3. 扩大异常捕获范围,把翻页过程中常见的异常都纳入捕获,触发时直接终止循环即可
  4. 翻页后增加显式等待,等旧页面的商品元素完全失效、新页面商品加载完成后再执行抓取,保证拿到的是新页数据
  5. 补全缺失的依赖导入,给cookie按钮也加上显式等待提升稳定性

修复后的可运行代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.common.exceptions import TimeoutException, ElementClickInterceptedException, StaleElementReferenceException
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from time import sleep

# 目标URL
url = "https://www.amazon.se/-/en/s?k=mirror+sticker&language=en_GB&crid=3LCT7C6GU8FUS&qid=1656847509&sprefix=mirror+sticker%2Caps%2C91&ref=sr_pg_1"
driver = webdriver.Chrome("chromedriver.exe")
driver.get(url)
driver.maximize_window()
wait = WebDriverWait(driver, 10)

# 等待cookie弹窗出现并点击接受
wait.until(EC.element_to_be_clickable((By.ID, "sp-cc-accept"))).click()

def get_text_store(web_elements_lst, storage_lst):
    for element in web_elements_lst:
        content = element.get_attribute("textContent").strip()
        storage_lst.append(content if content else "No data")

names_txt = []
prices_txt = []
while True:
    try:
        # 等待当前页商品列表加载完成
        wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".a-size-base-plus.a-color-base.a-text-normal")))
        web_elements_names = driver.find_elements(By.CSS_SELECTOR, ".a-size-base-plus.a-color-base.a-text-normal")
        web_elements_prices = driver.find_elements(By.CLASS_NAME, "a-price-whole")
        
        # 暂存当前页第一个商品元素,用于后续判断翻页是否完成
        first_item = web_elements_names[0]
        
        get_text_store(web_elements_names, names_txt)
        get_text_store(web_elements_prices, prices_txt)
        
        # 定位下一页按钮,滚动到避开顶部导航的位置再点击
        next_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//a[text()='Next']")))
        driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", next_btn)
        sleep(0.5) # 留短时间等滚动动画完成
        next_btn.click()
        
        # 等待旧页元素失效,确认翻页完成
        wait.until(EC.staleness_of(first_item))
    except (TimeoutException, ElementClickInterceptedException, StaleElementReferenceException):
        print("爬取终止,已到最后一页或触发异常")
        break

print(names_txt)
print(prices_txt)
driver.quit()
额外说明
  • 爬取时建议给driver加maximize_window()最大化窗口,减少元素被遮挡的概率
  • 价格目前只抓取了整数部分,如果需要完整价格可以把小数部分和货币符号的逻辑补上
  • 如果爬取过程中触发亚马逊人机验证,可以在每次翻页后加1-2秒的随机延时,降低被反爬拦截的概率

内容的提问来源于stack exchange,提问作者Stash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 05:42:31