Python爬取Shein页面报错:AttributeError: 'NoneType'无text属性
解决Shein墨西哥站点女装爬虫的AttributeError问题
问题概述
爬取Shein墨西哥站点女装页面时触发AttributeError: 'NoneType' object has no attribute 'text',错误发生在提取商品价格的代码段。
错误原因
- 类名匹配错误:HTML中价格元素的类是
normal-price-ctn__sale-price,但代码中误写为normal-price-ctn__sale-prices(多了末尾的s),导致find方法返回None。 - 全局查找而非局部查找:代码中使用
soup.find查找价格元素,会在整个页面中找第一个匹配项,而非当前商品节点下的元素,逻辑错误。 - 未处理元素缺失场景:没有判断元素是否存在就直接调用
.text,一旦元素找不到就会触发报错。
修复方案
- 修正价格元素的类名,去掉多余的
s; - 改用
item.find在当前商品节点范围内查找价格元素; - 添加判断逻辑,当价格元素不存在时,赋值默认值(如
"N/A"),避免程序崩溃。 - 补充请求头模拟浏览器访问,降低被Shein反爬机制拦截的概率。
修复后的完整代码
import requests from bs4 import BeautifulSoup import pandas as pd # 目标URL url = "https://www.shein.com.mx/style/Women-Clothing-sc-001121425.html?ici=mx_tab01navbar04&src_module=topcat&src_tab_page_id=page_select_class1686667964514&src_identifier=fc%3DWomen%60sc%3DROPA%60tc%3D0%60oc%3D0%60ps%3Dtab01navbar04%60jc%3DitemPicking_001121425&srctype=category&userpath=category-ROPA" # 添加请求头,模拟浏览器访问 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # 发送请求 response = requests.get(url, headers=headers) soup = BeautifulSoup(response.content, "html.parser") # 初始化存储列表 products = [] prices = [] urls = [] # 查找所有商品项 product_items = soup.find_all("div", class_="product-list-v2") for item in product_items: # 提取商品名称 product_name_elem = item.find("div", class_="S-product-item__name") if product_name_elem: product = product_name_elem.text.strip() products.append(product) print(f"商品名: {product}") else: products.append("N/A") print("商品名未找到") # 提取商品价格(修复核心部分) price_element = item.find("span", class_="normal-price-ctn__sale-price") if price_element: price = price_element.text.strip() prices.append(price) print(f"价格: {price}") else: prices.append("N/A") print("价格未找到") # 提取商品URL product_link_elem = item.find("a", class_="S-product-item__link") if product_link_elem and "href" in product_link_elem.attrs: product_url = "https://www.shein.com.mx" + product_link_elem["href"] urls.append(product_url) print(f"URL: {product_url}") else: urls.append("N/A") print("URL未找到") # 生成DataFrame并保存 data = { "Product": products, "Price": prices, "URL": urls } df = pd.DataFrame(data) df.to_excel("shein_data.xlsx", index=False) print("数据已保存到shein_data.xlsx")
额外说明
- Shein的页面结构可能随时变更,若后续再次出现类似问题,需重新检查HTML元素的类名或结构;
- 如果遇到动态加载的商品(页面滚动后才加载),单纯使用
requests无法获取全部数据,此时需要结合Selenium或Playwright等工具模拟浏览器行为。
内容的提问来源于stack exchange,提问作者Alexis Rodas
相关产品推荐
相关产品推荐

