You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Beautiful Soup爬虫异常:无法提取指定网页的成分列表

解决Beautiful Soup提取成分列表的问题

我来帮你搞定这个爬虫问题!从你描述的标签结构来看,问题大概率出在没有准确定位目标<td>元素或者没处理掉<sup>标签里的上标数字上,下面是一步步的解决方案:

1. 先确保能正确获取网页内容

很多网站会拦截无浏览器标识的请求,所以先给请求加个User-Agent头模拟浏览器访问:

import requests
from bs4 import BeautifulSoup

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}
url = "https://skinsalvationsf.com/2012/08/updated-comedogenic-ingredients-list/"
response = requests.get(url, headers=headers)
# 确认请求成功(状态码200)
if response.status_code != 200:
    print("请求网页失败,请检查网络或URL")
    exit()
soup = BeautifulSoup(response.text, 'html.parser')

2. 精准定位目标<td>元素

你提到的成分都在valign="top"且width="33%"的<td>标签里,用find_all直接筛选这些标签:

# 筛选符合属性的所有td标签
ingredient_tds = soup.find_all('td', attrs={'valign': 'top', 'width': '33%'})

3. 清理文本并提取成分

每个<td>里的成分文本后面跟着<sup>标签的数字,我们需要把这些数字去掉,保留纯净的成分名称:

import re

ingredients = []
for td in ingredient_tds:
    # 提取td内的所有文本并去除首尾空格
    raw_text = td.get_text(strip=True)
    # 用正则去掉末尾的数字(对应sup标签里的上标)
    cleaned_ingredient = re.sub(r'\d+$', '', raw_text).strip()
    # 过滤掉空字符串(避免网页里的空td)
    if cleaned_ingredient:
        ingredients.append(cleaned_ingredient)

4. 截取目标范围的成分列表

现在我们已经有了所有成分,只需要截取从Acetylated Lanolin到Octyl Palmitate的部分:

try:
    start_idx = ingredients.index('Acetylated Lanolin')
    end_idx = ingredients.index('Octyl Palmitate') + 1  # 包含最后一个成分
    target_list = ingredients[start_idx:end_idx]
    
    # 打印结果验证
    print("提取的目标成分列表:")
    for idx, ing in enumerate(target_list, 1):
        print(f"{idx}. {ing}")
except ValueError as e:
    print(f"未找到目标成分:{e},请检查文本清理是否正确")

可能遇到的坑

  • 如果index报错找不到目标成分,可能是文本清理不彻底(比如有些成分有特殊空格或大小写问题),可以改用模糊匹配,比如:
    # 找起始索引的模糊方式
    start_idx = next(i for i, ing in enumerate(ingredients) if ing.startswith('Acetylated Lanolin'))
    
  • 若网页结构有变动,可打开开发者工具重新确认<td>的属性是否还是valign="top"和width="33%",随时调整筛选条件。

内容的提问来源于stack exchange,提问作者pynewbee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:32:23