You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python循环迭代时Selenium提取的数据被覆盖,如何解决?

解决Selenium爬取亚马逊商品数据时数据被覆盖的问题

问题原因

你的代码里item字典是在循环外部初始化的,每次循环只是修改这个字典的内容,而product.append(item)添加的是字典的引用,不是独立副本。最终列表里所有元素指向的都是同一个字典,所以最后所有数据都会被最后一次循环的内容覆盖。

修复方案

把item字典的初始化移到for循环内部,每次迭代都创建一个新的字典,这样每个商品的数据都会存在独立的字典里,不会互相覆盖。

修改后的完整代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.wait import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

PATH="C:\\Program Files (x86)\\chromedriver.exe"
driver = webdriver.Chrome(PATH)
df_urls = pd.read_csv('D:/selenium/inputs/amazone-asin.csv',encoding='utf-8')
list_dicts_urls = df_urls.to_dict('records')

product = []
for url in list_dicts_urls:
    # 每次循环创建新的item字典,避免引用覆盖
    item = dict()
    product_url = 'https://' + url['MARKETPLACE'] + '/dp/' + url['ASIN']
    driver.get(product_url)

    try:
        item['title'] = driver.find_element(By.CSS_SELECTOR,'span#productTitle').text
    except:
        item['title'] = ''
        
    try:
        item['brand'] = driver.find_element(By.CSS_SELECTOR,'a#bylineInfo').text.replace('Visit the','').replace('Store','').strip()
    except:
        item['brand'] = ''
    try:
        rating = driver.find_element(By.CSS_SELECTOR,'span#acrCustomerReviewText').text.replace('ratings','').strip()
        rating = int(rating.replace(',', ''))
        item['rating'] = rating
    except:
        item['rating'] = ''
        
    # 用WebDriverWait替代time.sleep,更高效可靠
    WebDriverWait(driver, 5).until(
        EC.presence_of_element_located((By.XPATH, '//span[@class="a-price-whole"]'))
    )
    try:
        p1 = driver.find_element(By.XPATH, '//span[@class="a-price-whole"]').text
        p2 = driver.find_element(By.XPATH, '//span[@class="a-price-fraction"]').text
        item['price'] = p1 + p2
    except:
        item['price'] = ''
        
    product.append(item)

# 关闭浏览器释放资源
driver.quit()

df = pd.DataFrame(product)
df.to_csv("ama.csv", index=False)

额外优化建议

  • 替换time.sleep为WebDriverWait:等待目标元素加载完成再执行操作,避免固定等待时间导致的效率低下或元素未加载完成的问题。
  • 避免裸except:捕获具体异常类型(如NoSuchElementException),防止隐藏其他未知错误。
  • 导出CSV时添加index=False,避免生成不必要的索引列。

内容的提问来源于stack exchange,提问作者developer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 11:10:22