You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Shein页面报错:AttributeError: 'NoneType'无text属性

解决Shein墨西哥站点女装爬虫的AttributeError问题

问题概述

爬取Shein墨西哥站点女装页面时触发AttributeError: 'NoneType' object has no attribute 'text',错误发生在提取商品价格的代码段。

错误原因

  1. 类名匹配错误:HTML中价格元素的类是normal-price-ctn__sale-price,但代码中误写为normal-price-ctn__sale-prices(多了末尾的s),导致find方法返回None。
  2. 全局查找而非局部查找:代码中使用soup.find查找价格元素,会在整个页面中找第一个匹配项,而非当前商品节点下的元素,逻辑错误。
  3. 未处理元素缺失场景:没有判断元素是否存在就直接调用.text,一旦元素找不到就会触发报错。

修复方案

  • 修正价格元素的类名,去掉多余的s;
  • 改用item.find在当前商品节点范围内查找价格元素;
  • 添加判断逻辑,当价格元素不存在时,赋值默认值(如"N/A"),避免程序崩溃。
  • 补充请求头模拟浏览器访问,降低被Shein反爬机制拦截的概率。

修复后的完整代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 目标URL
url = "https://www.shein.com.mx/style/Women-Clothing-sc-001121425.html?ici=mx_tab01navbar04&src_module=topcat&src_tab_page_id=page_select_class1686667964514&src_identifier=fc%3DWomen%60sc%3DROPA%60tc%3D0%60oc%3D0%60ps%3Dtab01navbar04%60jc%3DitemPicking_001121425&srctype=category&userpath=category-ROPA"

# 添加请求头,模拟浏览器访问
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 发送请求
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.content, "html.parser")

# 初始化存储列表
products = []
prices = []
urls = []

# 查找所有商品项
product_items = soup.find_all("div", class_="product-list-v2")

for item in product_items:
    # 提取商品名称
    product_name_elem = item.find("div", class_="S-product-item__name")
    if product_name_elem:
        product = product_name_elem.text.strip()
        products.append(product)
        print(f"商品名: {product}")
    else:
        products.append("N/A")
        print("商品名未找到")

    # 提取商品价格(修复核心部分)
    price_element = item.find("span", class_="normal-price-ctn__sale-price")
    if price_element:
        price = price_element.text.strip()
        prices.append(price)
        print(f"价格: {price}")
    else:
        prices.append("N/A")
        print("价格未找到")

    # 提取商品URL
    product_link_elem = item.find("a", class_="S-product-item__link")
    if product_link_elem and "href" in product_link_elem.attrs:
        product_url = "https://www.shein.com.mx" + product_link_elem["href"]
        urls.append(product_url)
        print(f"URL: {product_url}")
    else:
        urls.append("N/A")
        print("URL未找到")

# 生成DataFrame并保存
data = {
    "Product": products,
    "Price": prices,
    "URL": urls
}
df = pd.DataFrame(data)
df.to_excel("shein_data.xlsx", index=False)
print("数据已保存到shein_data.xlsx")

额外说明

  • Shein的页面结构可能随时变更,若后续再次出现类似问题,需重新检查HTML元素的类名或结构;
  • 如果遇到动态加载的商品(页面滚动后才加载),单纯使用requests无法获取全部数据,此时需要结合Selenium或Playwright等工具模拟浏览器行为。

内容的提问来源于stack exchange,提问作者Alexis Rodas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 21:47:24