You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium批量导出爬取内容到CSV时替代件抓取异常问题求助

Selenium爬取联想兼容配件页替代件匹配问题

我正在使用Selenium实现爬取内容批量导出到CSV的功能,目标爬取页面为联想数据中心支持的兼容配件页:https://datacentersupport.lenovo.com/gb/en/products/storage/lenovo-storage/s3200/70l8/parts/display/compatible。现有代码可以正常爬取大部分字段,但无法正确抓取关联的替代件信息,最终导出的CSV中替代件单独成行,没有和对应主配件信息匹配。我需要实现每个主配件对应的替代件编号和主配件信息关联,输出到CSV的最后一列,以下是我的完整代码,请问逻辑哪里有问题?

from selenium import webdriver
from time import sleep
import pandas as pd
from selenium.common.exceptions import NoSuchElementException, TimeoutException
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC

# initializing webdriver 
driver = webdriver.Chrome(executable_path="~~chromedriver.exe")
url = " https://datacentersupport.lenovo.com/gb/en/products/storage/lenovo-storage/s3200/70l8/parts/display/compatible."
driver.get(url)
sleep(5)

results = []

#getting breadcrumbs
bread1 = driver.find_element_by_xpath("//span[@class='prod-catagory-name']")
bread2 = driver.find_element_by_xpath("//span[@class='prod-catagory-name']/a")

#grabbing table data and navigating 
pages = int(driver.find_element_by_xpath("//div[@class='page-container']/span[@class='icon-s-right active']/preceding-sibling::span[1]").text)
num = pages -1 

for _ in range(pages):
        rows = driver.find_elements_by_xpath("//table/tbody/tr/td[2]/div")
        for row in rows:
            parts = row.text
            results.append([url,bread1.text,parts])
        try: 
            for element in WebDriverWait(driver, 10).until(EC.visibility_of_all_elements_located((By.CSS_SELECTOR, "span[class='icon-s-down']"))):
                driver.execute_script("arguments[0].click();", element)
                sleep(5)
                substitute = WebDriverWait(driver, 10).until(EC.visibility_of_all_elements_located((By.XPATH, "//span[@class='icon-s-up']//following::tr[3]/td[contains(@class,'enabled-border')]//div[text()]")))
                for sub in substitute:
                    subs = sub.text
                    results.append(subs)
        except TimeoutException:
            pass
        except NoSuchElementException:
            break
        finally:
            try:
                pagination = driver.find_element_by_xpath("//div[@class='page-container']/span[@class='icon-s-right active']").click()
                sleep(3)
            except NoSuchElementException:
                break
df = pd.DataFrame(results)
df.to_csv('datacenter2.csv', index=False)
driver.quit()

当前导出的CSV效果可参考:
替代件单独出现在第12至16行


代码逻辑问题说明

  • 主配件与替代件无关联绑定:代码先批量把所有主配件信息存入结果列表,之后单独抓取的替代件也作为独立元素追加到结果列表,两类数据没有对应关系,自然出现替代件单独成行的问题。
  • 选择器全局匹配导致数据错乱:抓取替代件用的是全局Xpath,每次点击下拉按钮后会匹配页面所有符合条件的替代件,而非当前点击的主配件对应的替代件,会出现匹配错误、数据重复的问题。
  • 基础参数错误:你定义的URL开头有空格、末尾多了英文句号,可能导致页面加载异常。
  • 数据结构不匹配:主配件存入时是包含3个元素的列表,替代件存入时是单个字符串,转换为DataFrame时结构不统一,也会导致输出错位。

内容的提问来源于stack exchange,提问作者Reggie18

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 23:24:05