You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Playwright与BeautifulSoup爬取Vans新加坡站点时触发AttributeError: 'NoneType' object has no attribute 'text'的原因及代码问题排查

问题分析与解决方案

首先,咱们来拆解这个AttributeError:它的意思是你尝试调用一个None对象的.text属性——说白了就是soup.find('span', {'class': 'price'})没找到对应的元素,返回了None,你直接去取.text就触发错误了。结合你的代码和页面结构,问题主要出在这几个地方:

1. 变量名冲突导致逻辑混乱

你把BeautifulSoup的实例命名为soup,然后循环的时候又写了for soup in clothes:——这会把原来的soup对象给覆盖掉!虽然当前循环里的soup是单个<li>元素,但这个命名很容易混淆,而且会让后续的查找逻辑变得不可靠。

2. 盲目遍历所有<li>元素

你用soup.find_all('li')获取了页面里所有的<li>,但页面里的<li>不一定都是商品项——比如可能存在分页控件、筛选栏的<li>,这些元素里根本没有price和product-item-link这类商品相关的标签,所以find()会返回None,调用.text自然就报错了。

3. 页面加载判断不够可靠

你用page.is_visible('div.layer-product-list')来判断页面加载状态,但这个方法只是瞬间检查元素是否可见,如果页面还在异步加载商品数据,可能元素已经显示但内容还没渲染完成,导致你拿到的HTML里缺少部分商品的信息。


修复后的代码示例

from playwright.sync_api import sync_playwright
from bs4 import BeautifulSoup

with sync_playwright() as p:
    browser = p.chromium.launch(headless=False, slow_mo=50)
    page = browser.new_page()
    page.goto('https://www.vans.com.sg/customer/account/login/')
    page.fill('input#email', 'abc')
    page.fill('input#pass', 'abc')
    page.click('button#send2')
    page.goto('https://www.vans.com.sg/men/clothing.html')
    
    # 替换is_visible为wait_for_selector,确保元素完全加载完成
    page.wait_for_selector('#layer-product-list', state='visible')
    html = page.inner_html('#layer-product-list')
    
    soup = BeautifulSoup(html, 'html.parser')
    # 只定位带商品类别的li,避免无关元素
    clothes_items = soup.find_all('li', class_='item product product-item')
    
    for item in clothes_items:
        # 先获取元素,判断存在后再取text
        price_elem = item.find('span', class_='price')
        title_elem = item.find('a', class_='product-item-link')
        
        if price_elem and title_elem:
            price = price_elem.text.strip()
            title = title_elem.text.strip()
            print(f'Title = {title}, Price = {price}')
        else:
            # 可以跳过或打印日志记录异常项
            print("Skipping item with missing price or title")
    
    browser.close()

关键修复点说明

  • 变量名重命名:把循环变量从soup改成item,避免覆盖BeautifulSoup实例。
  • 精准选择商品元素:用find_all('li', class_='item product product-item')只获取商品对应的<li>,过滤掉无关元素。
  • 增加存在性判断:在调用.text前先检查元素是否存在,避免None对象报错。
  • 更可靠的加载等待:用page.wait_for_selector替代is_visible,确保元素及其内容完全渲染后再获取HTML。

内容的提问来源于stack exchange,提问作者Ryan Loh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 19:07:45