You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BS4爬取Mecca护肤品成分时,相同p类标签导致无法精准提取求助

精准提取Mecca澳洲官网护肤品成分的解决方案

问题场景

需要爬取Mecca澳洲官网护肤品板块的产品成分列表,以Drunk Elephant Lala Retro Whipped Cream为例,成分内容位于如下HTML结构中:

<div role="region" title="Ingredients" id="sect5e670152-13f8-46ca-9f0b-4c60b533d079" aria-labelledby="accordion-5e670152-13f8-46ca-9f0b-4c60b533d079" aria-controls="accordion-5e670152-13f8-46ca-9f0b-4c60b533d079-container" class="css-1g2bsc4">
  <div>
    <p class="css-pdiqc3 e151354a4">Water/aqua/eau, glycerin, caprylic/capric triglyceride, isopropyl isostearate, pseudozyma epicola/camellia sinensis seed oil/glucose/glycine soja (soybean) meal/malt extract/yeast extract ferment filtrate, glyceryl stearate se, cetearyl alcohol, palmitic acid, stearic acid, pentylene glycol, plantago lanceolata leaf extract, adansonia digitata seed oil, citrullus lanatus (watermelon) seed oil, passiflora edulis seed oil, schinziophyton rautanenii kernel oil, sclerocarya birrea seed oil, polyglyceryl6 ximenia americana seedate, cholesterol, ceramide ap, ceramide eop, sodium hyaluronate crosspolymer, ceramide np, phytosphingosine, ceteareth20, trisodium ethylenediamine disuccinate, tocopherol, sodium lauroyl lactylate, sodium hydroxide, citric acid, carbomer, xanthan gum, caprylyl glycol, chlorphenesin, phenoxyethanol, ethylhexylglycerin.</p>
  </div>
</div>

当前爬取代码因页面存在多个相同class="css-pdiqc3 e151354a4"的p标签,无法精准提取目标成分:

import requests
from bs4 import BeautifulSoup

url = "https://www.mecca.com"

productlinks = []

for x in range(1,5):
    soup = BeautifulSoup(requests.get(f'https://www.mecca.com/en-au/skincare/?page={x}').content, "html.parser")
    products = soup.find_all('div', class_="css-1r7iqog")

    for item in products:
        for link in item.find_all('a', href=True):
            productlinks.append(url + link['href'])
    
for link in productlinks:
    response = requests.get(link)

    soup = BeautifulSoup(response.content, "html.parser")
    brand = soup.find("span", class_="product-brand css-738lkl e11u5ot719").string
    name = soup.find('span', class_='product-name css-x4jxi0 e11u5ot718').string
    Ingred = soup.find('p', class_='css-pdiqc3 e151354a4') 

    print(brand)
    print(name)
    print(Ingred)

尝试过以下无效方案:

# 方案1:通过索引定位,但索引可能随页面结构变化失效
Ingred = soup.find_all('p', class_='css-pdiqc3 e151354a4')[8].get_text()

# 方案2:错误的标签定位方式
Details = soup.find_all('Ingredients', {'class':'css-1g2bsc4'})
Ingred = BeautifulSoup(str(Details).strip()).get_text()

有效解决方案

利用成分所在父容器的title="Ingredients"属性精准定位,再提取内部的p标签内容:

修改代码中提取成分的部分为:

# 先定位包含成分的父容器
ingredients_container = soup.find('div', title="Ingredients")
# 从容器内提取目标p标签的文本
if ingredients_container:
    Ingred = ingredients_container.find('p', class_='css-pdiqc3 e151354a4').get_text(strip=True)
else:
    Ingred = "成分未找到"

完整修改后的代码:

import requests
from bs4 import BeautifulSoup

url = "https://www.mecca.com"

productlinks = []

for x in range(1,5):
    soup = BeautifulSoup(requests.get(f'https://www.mecca.com/en-au/skincare/?page={x}').content, "html.parser")
    products = soup.find_all('div', class_="css-1r7iqog")

    for item in products:
        for link in item.find_all('a', href=True):
            productlinks.append(url + link['href'])
    
for link in productlinks:
    response = requests.get(link)

    soup = BeautifulSoup(response.content, "html.parser")
    brand = soup.find("span", class_="product-brand css-738lkl e11u5ot719").string
    name = soup.find('span', class_='product-name css-x4jxi0 e11u5ot718').string
    
    # 精准提取成分
    ingredients_container = soup.find('div', title="Ingredients")
    Ingred = ingredients_container.find('p', class_='css-pdiqc3 e151354a4').get_text(strip=True) if ingredients_container else "成分未找到"

    print(brand)
    print(name)
    print(Ingred)
    print("---")

方案说明

  • 父容器的title="Ingredients"是唯一标识成分区域的属性,比单纯依赖class定位更可靠,避免其他相同class元素的干扰
  • 增加了空值判断,避免因页面结构异常导致代码报错
  • 使用get_text(strip=True)可以去除文本前后的空白字符,让成分列表更整洁

内容的提问来源于stack exchange,提问作者user22276310

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 13:12:04