You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫抓取新加坡屈臣氏网站数据为空,求代码排查

新加坡屈臣氏爬虫结果为空的问题排查与解决

问题描述

我用Python编写爬虫,尝试抓取新加坡屈臣氏网站的商品数据,但运行后结果为空,不清楚代码哪里出现问题。以下是我的代码:

import requests
from bs4 import BeautifulSoup
import pandas as pd


# 设置pandas显示选项
pd.options.display.width = 1000
pd.options.display.max_rows = 1000

HEADERS = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/105.0.0.0 Safari/537.36'
}

# 创建存储数据的列表
items = []
prices = []


# 定义页面抓取函数
def scrape_page(page):
    try:
        url = f'https://www.watsons.com.sg/health/c/2100000?currentPage={page}'
        response = requests.get(url, headers=HEADERS)
        response.raise_for_status()  # 检查请求是否成功

        soup = BeautifulSoup(response.text, 'html.parser')

        # 查找商品元素
        product_items = soup.find_all('e2-product-tile', class_='ng-star-inserted hasPromotion-2')
        print(soup)

        for item in product_items:
            # 提取商品名称
            name = item.find('h2', class_='productName').get_text(strip=True)
            items.append(name)

            # 提取商品价格
            price = item.find('div', class_='formatted-value ng-star-inserted').get_text(strip=True)
            prices.append(price)

    except requests.RequestException as e:
        print(f"发生错误: {e}")

# 示例:抓取第一页
for page in range(1, 2):
    scrape_page(page)

# 创建DataFrame
df = pd.DataFrame({'Item': items, 'Price': prices})

# 打印结果
print(df)

问题原因

  1. 页面动态渲染:该网站基于Angular框架开发,商品数据是通过JavaScript动态加载渲染的。requests只能获取到初始的静态HTML,其中并没有实际的商品内容,BeautifulSoup自然找不到目标元素。
  2. 元素定位不准确:代码中使用的e2-product-tile自定义标签和hasPromotion-2类名,在静态HTML中仅作为占位符存在,且实际渲染后的元素类名可能存在变动,导致无法匹配到目标商品节点。

解决方法

方法一:使用Selenium模拟浏览器渲染

通过模拟真实浏览器的行为,等待页面JS渲染完成后再抓取数据,代码示例:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import pandas as pd

pd.options.display.width = 1000
pd.options.display.max_rows = 1000

items = []
prices = []

def scrape_page(page):
    url = f'https://www.watsons.com.sg/health/c/2100000?currentPage={page}'
    driver.get(url)
    # 等待商品元素加载完成,超时10秒
    try:
        WebDriverWait(driver, 10).until(
            EC.presence_of_all_elements_located((By.TAG_NAME, 'e2-product-tile'))
        )
        # 获取渲染后的页面源码
        page_source = driver.page_source
        soup = BeautifulSoup(page_source, 'html.parser')
        # 匹配所有商品节点
        product_items = soup.find_all('e2-product-tile', class_='ng-star-inserted')
        for item in product_items:
            name = item.find('h2', class_='productName').get_text(strip=True)
            items.append(name)
            price = item.find('div', class_='formatted-value').get_text(strip=True)
            prices.append(price)
    except Exception as e:
        print(f"第{page}页抓取失败: {str(e)}")

# 初始化Chrome浏览器(需提前下载对应版本的ChromeDriver)
driver = webdriver.Chrome()

# 抓取第一页数据
scrape_page(1)

# 关闭浏览器
driver.quit()

# 生成DataFrame并打印
df = pd.DataFrame({'商品名称': items, '价格': prices})
print(df)

方法二:直接调用API接口(高效推荐)

通过浏览器开发者工具(F12)抓包,找到加载商品数据的API接口,直接请求获取JSON格式数据,效率更高。示例代码:

import requests
import pandas as pd

pd.options.display.width = 1000
pd.options.display.max_rows = 1000

HEADERS = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/105.0.0.0 Safari/537.36',
    'Accept': 'application/json'
}

items = []
prices = []

def scrape_page(page):
    # 抓包得到的商品列表API接口
    api_url = f'https://www.watsons.com.sg/api/catalog/products?categoryId=2100000&currentPage={page}&pageSize=24'
    try:
        response = requests.get(api_url, headers=HEADERS)
        response.raise_for_status()
        data = response.json()
        # 提取商品名称和价格
        for product in data['products']:
            items.append(product['name'])
            prices.append(product['price']['formattedValue'])
    except requests.RequestException as e:
        print(f"API请求失败: {str(e)}")

# 抓取第一页数据
scrape_page(1)

# 生成DataFrame并打印
df = pd.DataFrame({'商品名称': items, '价格': prices})
print(df)

注意事项

  • 使用Selenium时,需确保浏览器驱动(如ChromeDriver)版本与浏览器版本一致,可从官方渠道下载对应驱动。
  • 调用API时,若请求被拦截,可补充Referer、Origin等请求头参数,模拟真实浏览器请求。
  • 爬取数据时需遵守网站的robots.txt规则,控制请求频率,避免IP被封禁。

内容的提问来源于stack exchange,提问作者jjbkd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 10:54:52