为何我的Zara商品爬取代码输出为空DataFrame?
爬取Zara女装夹克数据返回空DataFrame的问题排查与解决
问题描述
尝试爬取Zara网站女装夹克数据用于趋势分析,但运行代码后得到空DataFrame,输出如下:
Empty DataFrame Columns: [Title, Price, Discount] Index: []爬取代码:
import pandas as pd import time from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support.ui import Select from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By # Specify the full path to the ChromeDriver executable driver=webdriver.Chrome() driver.get('https://www.zara.com/us/en/search?searchTerm=women%20jackets§ion=WOMAN') driver.maximize_window() container=driver.find_element(by='xpath', value='.//ul[contains(@class, "product-grid__product-list")]') products=container.find_elements(by='xpath', value='.//li[contains(@class, "product-grid-product")]') product_title=[] product_price=[] discount=[] for product in products: product_title.append(product.find_element(by='xpath',value='.//a[@class="product-link _item product-grid-product-info__name link"]').text) product_price.append(product.find_element(by='xpath', value='.//span[@class="money-amount__main"]').text) discount.append(product.find_element(by='xpath',value='.//span[@class="price-current__discount-percentage"]').text) driver.quit() df=pd.DataFrame({'Title':product_title,'Price':product_price, 'Discount':discount}) df.to_csv('Zara_Jackets.csv',index=False) print(df)
问题原因与解决步骤
1. 页面未完全加载就查找元素
Zara的产品列表是动态渲染的,直接执行find_element时页面DOM可能还未加载完成,导致找不到容器和产品元素,最终products为空列表。
解决:使用WebDriverWait显式等待元素加载完成,确保元素存在后再进行操作。
2. 元素定位器过于严格或已失效
原代码中标题的XPath使用精确匹配@class="product-link _item product-grid-product-info__name link",但页面元素的class可能随网站更新变化,或存在动态添加的class,导致定位失败。另外,部分产品没有折扣,直接用find_element查找折扣元素会抛出异常,中断循环。
解决:
- 对class使用
contains模糊匹配,提高定位稳定性; - 用
find_elements(复数形式)判断折扣元素是否存在,不存在则填充默认值。
3. ChromeDriver路径未正确配置
原代码注释提示需要指定ChromeDriver路径,但实际代码未配置,可能导致驱动加载异常,影响页面渲染。
解决:通过Service类指定ChromeDriver的实际路径,确保驱动与Chrome浏览器版本匹配。
修改后的完整代码
import pandas as pd from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 替换为你的ChromeDriver实际路径 service = Service(executable_path='./chromedriver') driver = webdriver.Chrome(service=service) driver.get('https://www.zara.com/us/en/search?searchTerm=women%20jackets§ion=WOMAN') driver.maximize_window() # 显式等待产品列表容器加载,最长等待10秒 wait = WebDriverWait(driver, 10) container = wait.until(EC.presence_of_element_located( (By.XPATH, './/ul[contains(@class, "product-grid__product-list")]') )) products = container.find_elements(By.XPATH, './/li[contains(@class, "product-grid-product")]') product_title = [] product_price = [] discount = [] for product in products: # 获取产品标题(模糊匹配class) title_elem = product.find_element(By.XPATH, './/a[contains(@class, "product-grid-product-info__name")]') product_title.append(title_elem.text.strip()) # 获取产品价格 price_elem = product.find_element(By.XPATH, './/span[@class="money-amount__main"]') product_price.append(price_elem.text.strip()) # 处理折扣:存在则取文本,不存在填"0%" discount_elems = product.find_elements(By.XPATH, './/span[@class="price-current__discount-percentage"]') discount.append(discount_elems[0].text.strip() if discount_elems else "0%") driver.quit() df = pd.DataFrame({'Title': product_title, 'Price': product_price, 'Discount': discount}) df.to_csv('Zara_Jackets.csv', index=False) print(df)
关键修改说明
- 添加显式等待,确保产品容器加载完成后再获取产品列表;
- 标题定位改用
contains匹配class,避免因class动态变化导致定位失败; - 折扣元素使用
find_elements判断存在性,避免无折扣产品引发的异常; - 正确配置ChromeDriver路径,保证驱动正常加载;
- 对文本结果使用
strip()去除多余空格,优化数据整洁度。
内容的提问来源于stack exchange,提问作者Win123
相关产品推荐
相关产品推荐

