使用BeautifulSoup爬取Zalora商品信息时遇Output None问题求助
解决Zalora商品信息爬取返回None的问题
我来帮你排查下爬取Zalora时返回None的问题,大概率是这几个原因导致的,咱们一步步解决:
1. 选择器不完整(最直接的原因)
你代码里的class选择器写了一半:"b-catalogList__itmBrand fsm ...",带省略号的写法会让BeautifulSoup找不到对应元素,直接返回None。
你可以这么调整:
- 用完整的class值:先通过浏览器F12查看元素的完整class(比如实际可能是
b-catalogList__itmBrand fsm f--bold),然后精确匹配; - 或者用前缀匹配:只取class里唯一的部分,比如:
itemBrand = soup.find("span", class_="b-catalogList__itmBrand") # 或者用CSS选择器更灵活 itemBrand = soup.select_one("span.b-catalogList__itmBrand")
2. 页面动态加载,静态HTML无数据
Zalora的商品列表很多是通过JavaScript动态渲染的,requests.get()只能拿到初始的空框架HTML,里面根本没有商品数据,自然查不到元素。这时候有两个靠谱方案:
方案A:用Selenium模拟浏览器渲染
安装Selenium和对应浏览器驱动(比如ChromeDriver),让浏览器帮你加载完所有动态内容后再爬取:
from selenium import webdriver from bs4 import BeautifulSoup # 初始化浏览器驱动 driver = webdriver.Chrome() url = 'https://www.zalora.com.hk/men/clothing/shirt/?gender=men&dir=desc&sort=popularity&category_id=31&enable_visual_sort=1' driver.get(url) # 获取渲染后的页面源码 soup = BeautifulSoup(driver.page_source, 'html.parser') # 开始提取数据 brands = soup.select("span.b-catalogList__itmBrand") names = soup.select("div.b-catalogList__itmName") original_prices = soup.select("span.b-catalogList__itmOriginalPrice") # 记得关闭浏览器 driver.quit()
方案B:抓包获取API接口
打开浏览器F12的Network标签,筛选XHR请求,找到Zalora用来拉取商品数据的API接口,直接请求这个接口拿JSON数据——这比爬HTML高效多了,还不容易被反爬。
3. 反爬机制拦截了你的请求
Zalora会检测请求是否来自真实浏览器,你需要给requests.get()加个请求头,模拟浏览器访问:
def make_soup(url): headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers) response.encoding = response.apparent_encoding # 避免乱码 return BeautifulSoup(response.text, 'html.parser')
完整测试代码
结合上面的优化,给你一个可以直接测试的代码(如果页面静态能拿到数据的话):
from bs4 import BeautifulSoup import requests def make_soup(url): headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers) response.encoding = response.apparent_encoding return BeautifulSoup(response.text, 'html.parser') url = 'https://www.zalora.com.hk/men/clothing/shirt/?gender=men&dir=desc&sort=popularity&category_id=31&enable_visual_sort=1' soup = make_soup(url) # 批量提取所有商品信息 items = soup.select("div.b-catalogList__itm") for idx, item in enumerate(items, 1): brand = item.select_one("span.b-catalogList__itmBrand").text.strip() if item.select_one("span.b-catalogList__itmBrand") else "无品牌" name = item.select_one("div.b-catalogList__itmName").text.strip() if item.select_one("div.b-catalogList__itmName") else "无名称" original_price = item.select_one("span.b-catalogList__itmOriginalPrice").text.strip() if item.select_one("span.b-catalogList__itmOriginalPrice") else "无原价" print(f"商品{idx}:\n品牌: {brand}\n名称: {name}\n原价: {original_price}\n---")
内容的提问来源于stack exchange,提问作者TedLLH
相关产品推荐
相关产品推荐

