You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取li类商品时报ResultSet无find属性错误如何解决

问题修复与全量爬取实现方案

错误根源

你遇到的AttributeError核心原因是:find_all()返回的是ResultSet类型(元素列表集合),列表没有find()方法,你把列表当成单个元素操作了。
你提到的ul和li之间的未知内容不需要单独处理,直接定位商品节点即可规避层级干扰。

修复后的可运行代码

from bs4 import BeautifulSoup
import requests

# 增加请求头模拟浏览器,避免被网站反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

# 全量爬取需要处理分页,先定义单页爬取函数
def crawl_page(page_num):
    url = f"https://www.uplannerperu.com/shop/page/{page_num}/" if page_num > 1 else "https://www.uplannerperu.com/shop/"
    resp = requests.get(url, headers=headers)
    resp.encoding = "utf-8"
    soup = BeautifulSoup(resp.text, 'lxml')
    # 直接定位所有商品li节点,跳过中间ul层级,避免未知内容干扰
    product_items = soup.find_all("li", class_="product")
    page_products = []
    for item in product_items:
        # 可根据需求调整提取的字段
        product_a = item.find("a", class_="woocommerce-LoopProduct-link")
        product_name = product_a.get("aria-label") if product_a else ""
        product_url = product_a.get("href") if product_a else ""
        product_price = item.find("span", class_="price").text.strip() if item.find("span", class_="price") else ""
        page_products.append({
            "name": product_name,
            "url": product_url,
            "price": product_price
        })
    return page_products

# 先获取总页数
first_resp = requests.get("https://www.uplannerperu.com/shop/", headers=headers)
first_soup = BeautifulSoup(first_resp.text, 'lxml')
page_nums = int(first_soup.find_all("a", class_="page-numbers")[-2].text)

# 遍历所有页爬取
all_products = []
for page in range(1, page_nums + 1):
    print(f"正在爬取第{page}页")
    page_data = crawl_page(page)
    all_products.extend(page_data)

# 输出爬取结果
print(f"共爬取到{len(all_products)}个商品")
for p in all_products:
    print(p)

关键修复点说明

  • 增加User-Agent请求头:避免网站识别为爬虫返回异常页面,导致元素定位失败
  • 直接定位商品节点:不逐层查找div→ul→li,跳过中间的未知标签干扰,直接匹配商品的公共class属性product
  • 分页逻辑处理:先拿总页数再循环爬取所有分页,实现全量商品采集
  • 增加空值判断:避免部分商品缺少字段时抛出异常

内容的提问来源于stack exchange,提问作者carpeta94

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 01:54:05