You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Zalora商品信息时遇Output None问题求助

解决Zalora商品信息爬取返回None的问题

我来帮你排查下爬取Zalora时返回None的问题,大概率是这几个原因导致的,咱们一步步解决:

1. 选择器不完整(最直接的原因)

你代码里的class选择器写了一半:"b-catalogList__itmBrand fsm ...",带省略号的写法会让BeautifulSoup找不到对应元素,直接返回None。

你可以这么调整:

  • 用完整的class值:先通过浏览器F12查看元素的完整class(比如实际可能是b-catalogList__itmBrand fsm f--bold),然后精确匹配;
  • 或者用前缀匹配:只取class里唯一的部分,比如:
itemBrand = soup.find("span", class_="b-catalogList__itmBrand")
# 或者用CSS选择器更灵活
itemBrand = soup.select_one("span.b-catalogList__itmBrand")

2. 页面动态加载,静态HTML无数据

Zalora的商品列表很多是通过JavaScript动态渲染的,requests.get()只能拿到初始的空框架HTML,里面根本没有商品数据,自然查不到元素。这时候有两个靠谱方案:

方案A:用Selenium模拟浏览器渲染

安装Selenium和对应浏览器驱动(比如ChromeDriver),让浏览器帮你加载完所有动态内容后再爬取:

from selenium import webdriver
from bs4 import BeautifulSoup

# 初始化浏览器驱动
driver = webdriver.Chrome()
url = 'https://www.zalora.com.hk/men/clothing/shirt/?gender=men&dir=desc&sort=popularity&category_id=31&enable_visual_sort=1'
driver.get(url)

# 获取渲染后的页面源码
soup = BeautifulSoup(driver.page_source, 'html.parser')

# 开始提取数据
brands = soup.select("span.b-catalogList__itmBrand")
names = soup.select("div.b-catalogList__itmName")
original_prices = soup.select("span.b-catalogList__itmOriginalPrice")

# 记得关闭浏览器
driver.quit()

方案B:抓包获取API接口

打开浏览器F12的Network标签,筛选XHR请求,找到Zalora用来拉取商品数据的API接口,直接请求这个接口拿JSON数据——这比爬HTML高效多了,还不容易被反爬。

3. 反爬机制拦截了你的请求

Zalora会检测请求是否来自真实浏览器,你需要给requests.get()加个请求头,模拟浏览器访问:

def make_soup(url):
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    response = requests.get(url, headers=headers)
    response.encoding = response.apparent_encoding  # 避免乱码
    return BeautifulSoup(response.text, 'html.parser')

完整测试代码

结合上面的优化,给你一个可以直接测试的代码(如果页面静态能拿到数据的话):

from bs4 import BeautifulSoup
import requests

def make_soup(url):
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    response = requests.get(url, headers=headers)
    response.encoding = response.apparent_encoding
    return BeautifulSoup(response.text, 'html.parser')

url = 'https://www.zalora.com.hk/men/clothing/shirt/?gender=men&dir=desc&sort=popularity&category_id=31&enable_visual_sort=1'
soup = make_soup(url)

# 批量提取所有商品信息
items = soup.select("div.b-catalogList__itm")
for idx, item in enumerate(items, 1):
    brand = item.select_one("span.b-catalogList__itmBrand").text.strip() if item.select_one("span.b-catalogList__itmBrand") else "无品牌"
    name = item.select_one("div.b-catalogList__itmName").text.strip() if item.select_one("div.b-catalogList__itmName") else "无名称"
    original_price = item.select_one("span.b-catalogList__itmOriginalPrice").text.strip() if item.select_one("span.b-catalogList__itmOriginalPrice") else "无原价"
    print(f"商品{idx}:\n品牌: {brand}\n名称: {name}\n原价: {original_price}\n---")

内容的提问来源于stack exchange,提问作者TedLLH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:44:49