You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python爬取WooCommerce网站遇AttributeError问题求解

解决爬取WooCommerce站点时的AttributeError错误

你遇到的AttributeError: 'NoneType' object has no attribute 'full_text'错误,本质是代码中调用html.find()后直接访问.full_text,但该方法没找到匹配元素返回了None,导致属性调用失败。结合你的代码,以下是具体修复方案:

核心问题点

  1. 未对元素查找结果做存在性校验,直接调用属性
  2. 部分CSS选择器写法错误(如nav.ruk_rating_snippet data-sku)
  3. 未导入time和csv模块(代码中用到但没声明)
  4. 价格选择器用索引[1]存在越界风险

修复后的完整代码

from requests_html import HTMLSession
import http.client
import time
import csv  # 补全缺失的模块

http.client._MAXHEADERS = 1000
url='https://agedcarestore.com.au/product-category/physio-products/arthritis/'

s=HTMLSession()

def get_links(url):
    r=s.get(url)
    items=r.html.find('div.product-small.box')
    links=[]
    for item in items:
        # 先检查a标签存在再取href
        a_tag = item.find('a', first=True)
        if a_tag and 'href' in a_tag.attrs:
            links.append(a_tag.attrs['href'])
    return links

def get_product(link):
    r=s.get(link)
    product = {}
    
    # 处理标题:先判断元素是否存在
    title_elem = r.html.find('h1', first=True)
    product['title'] = title_elem.full_text.strip() if title_elem else '无标题'
    
    # 处理价格:避免索引越界,兼容不同价格展示场景
    price_elems = r.html.find('span.woocommerce-Price-amount.amount bdi')
    if len(price_elems) >= 2:
        product['price'] = price_elems[1].full_text.strip()
    elif price_elems:
        product['price'] = price_elems[0].full_text.strip()
    else:
        product['price'] = '无价格'
    
    # 处理SKU:优先取span.sku,再尝试从rating元素取属性
    sku_elem = r.html.find('span.sku', first=True)
    if sku_elem:
        product['sku'] = sku_elem.full_text.strip()
    else:
        rating_elem = r.html.find('nav.ruk_rating_snippet', first=True)
        product['sku'] = rating_elem.attrs.get('data-sku', '无SKU') if rating_elem else '无SKU'
    
    # 处理标签:判断元素是否存在
    tag_elem = r.html.find('a[rel=tag]', first=True)
    product['tag'] = tag_elem.full_text.strip() if tag_elem else '无标签'
    
    print(product)
    return product 

results = []
links = get_links(url)

for link in links:
    results.append(get_product(link))
    time.sleep(1)

# 写入CSV时先判断数据是否为空,避免索引错误
if results:
    with open('version1.csv', 'w', encoding='utf8', newline='') as f:
        fc = csv.DictWriter(f, fieldnames=results[0].keys())
        fc.writeheader()
        fc.writerows(results)  # 补全写入数据的逻辑
else:
    print("未获取到任何产品数据")

关键修复说明

  • 元素存在性校验:所有find()结果都先判断是否为None,再调用属性,从根源避免报错
  • 修正选择器逻辑:将错误的nav.ruk_rating_snippet data-sku改为读取元素的data-sku属性
  • 价格鲁棒性处理:不再硬编码索引[1],而是根据匹配到的元素长度动态取值
  • 模块补全:导入代码中实际用到的time和csv模块
  • CSV写入完善:补充writerows()写入数据的逻辑,同时判断结果列表是否为空,防止results[0]索引错误

内容的提问来源于stack exchange,提问作者pengoul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 08:32:17