Selenium批量导出爬取内容到CSV时替代件抓取异常问题求助
Selenium爬取联想兼容配件页替代件匹配问题
我正在使用Selenium实现爬取内容批量导出到CSV的功能,目标爬取页面为联想数据中心支持的兼容配件页:https://datacentersupport.lenovo.com/gb/en/products/storage/lenovo-storage/s3200/70l8/parts/display/compatible。现有代码可以正常爬取大部分字段,但无法正确抓取关联的替代件信息,最终导出的CSV中替代件单独成行,没有和对应主配件信息匹配。我需要实现每个主配件对应的替代件编号和主配件信息关联,输出到CSV的最后一列,以下是我的完整代码,请问逻辑哪里有问题?
from selenium import webdriver from time import sleep import pandas as pd from selenium.common.exceptions import NoSuchElementException, TimeoutException from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC # initializing webdriver driver = webdriver.Chrome(executable_path="~~chromedriver.exe") url = " https://datacentersupport.lenovo.com/gb/en/products/storage/lenovo-storage/s3200/70l8/parts/display/compatible." driver.get(url) sleep(5) results = [] #getting breadcrumbs bread1 = driver.find_element_by_xpath("//span[@class='prod-catagory-name']") bread2 = driver.find_element_by_xpath("//span[@class='prod-catagory-name']/a") #grabbing table data and navigating pages = int(driver.find_element_by_xpath("//div[@class='page-container']/span[@class='icon-s-right active']/preceding-sibling::span[1]").text) num = pages -1 for _ in range(pages): rows = driver.find_elements_by_xpath("//table/tbody/tr/td[2]/div") for row in rows: parts = row.text results.append([url,bread1.text,parts]) try: for element in WebDriverWait(driver, 10).until(EC.visibility_of_all_elements_located((By.CSS_SELECTOR, "span[class='icon-s-down']"))): driver.execute_script("arguments[0].click();", element) sleep(5) substitute = WebDriverWait(driver, 10).until(EC.visibility_of_all_elements_located((By.XPATH, "//span[@class='icon-s-up']//following::tr[3]/td[contains(@class,'enabled-border')]//div[text()]"))) for sub in substitute: subs = sub.text results.append(subs) except TimeoutException: pass except NoSuchElementException: break finally: try: pagination = driver.find_element_by_xpath("//div[@class='page-container']/span[@class='icon-s-right active']").click() sleep(3) except NoSuchElementException: break df = pd.DataFrame(results) df.to_csv('datacenter2.csv', index=False) driver.quit()
当前导出的CSV效果可参考:
代码逻辑问题说明
- 主配件与替代件无关联绑定:代码先批量把所有主配件信息存入结果列表,之后单独抓取的替代件也作为独立元素追加到结果列表,两类数据没有对应关系,自然出现替代件单独成行的问题。
- 选择器全局匹配导致数据错乱:抓取替代件用的是全局Xpath,每次点击下拉按钮后会匹配页面所有符合条件的替代件,而非当前点击的主配件对应的替代件,会出现匹配错误、数据重复的问题。
- 基础参数错误:你定义的URL开头有空格、末尾多了英文句号,可能导致页面加载异常。
- 数据结构不匹配:主配件存入时是包含3个元素的列表,替代件存入时是单个字符串,转换为DataFrame时结构不统一,也会导致输出错位。
内容的提问来源于stack exchange,提问作者Reggie18
相关产品推荐
相关产品推荐

