You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多URL网页爬取遇AttributeError:'NoneType'对象无'text'属性求助

问题分析与解决方案

嘿,我来帮你搞定这个爬虫问题!先拆解一下你遇到的几个核心问题:

1. 变量名冲突(最关键的错误)

你这段代码里犯了一个很容易忽略的错误:

individual_items = all_products_on_page.find_all(class_='product-item-info')
# 下面的循环变量和上面的列表同名了!
aero_product_name = [individual_items.find(class_='product-item-link').text for individual_items in individual_items]

当你用individual_items作为循环变量时,会在迭代过程中覆盖掉原来存储所有商品的列表,导致后续的查找逻辑混乱,这是引发错误的重要原因之一。

2. 未处理None类型的情况

当某个商品页面上找不到price类的元素时(比如缺货商品、预购商品没有显示价格),individual_items.find(class_='price')会返回None,这时候再调用.text就会触发AttributeError。

3. 数据被重复覆盖

你每次循环都新建一个DataFrame并写入同一个CSV文件,这会导致前一个页面的数据被后一个页面完全覆盖,最终CSV里只会保留最后一个页面的内容。


修正后的完整代码

#imports
import pandas as pd
import requests
from bs4 import BeautifulSoup

#Product Websites For Consolidation
urls = ['https://www.aeroprecisionusa.com/ar15/lower-receivers/stripped-lowers?product_list_limit=all', 
        'https://www.aeroprecisionusa.com/ar15/lower-receivers/complete-lowers?product_list_limit=all']

# 初始化空DataFrame,用来存放所有页面的商品数据
master_df = pd.DataFrame(columns=['Product Name', 'Prices', 'Source URL'])

headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:71.0) Gecko/20100101 Firefox/71.0"}

for website in urls:
    page = requests.get(website, headers=headers)
    soup = BeautifulSoup(page.content, 'html.parser')
    
    # 先检查是否找到商品列表容器,避免后续报错
    all_products_on_page = soup.find(class_='products wrapper container grid products-grid')
    if not all_products_on_page:
        print(f"警告:在 {website} 未找到商品列表,跳过该页面")
        continue
    
    # 这里把列表命名为items,避免和循环变量冲突
    items = all_products_on_page.find_all(class_='product-item-info')
    
    # 提取商品名称和价格,处理None的情况
    product_names = []
    product_prices = []
    for item in items:
        # 提取名称
        name_tag = item.find(class_='product-item-link')
        product_name = name_tag.text.strip() if name_tag else "未知名称"
        product_names.append(product_name)
        
        # 提取价格,处理找不到价格的情况
        price_tag = item.find(class_='price')
        product_price = price_tag.text.strip() if price_tag else "无价格"
        product_prices.append(product_price)
    
    # 把当前页面的数据转换成DataFrame,追加到总表中
    current_df = pd.DataFrame({
        'Product Name': product_names,
        'Prices': product_prices,
        'Source URL': [website]*len(product_names)  # 记录商品来自哪个URL
    })
    master_df = pd.concat([master_df, current_df], ignore_index=True)

# 最后统一保存所有数据到CSV
master_df.to_csv('MasterPriceTracker.csv', index=False)

关键修改点说明

  • 修复变量名冲突:把存储商品列表的变量改成items,循环变量用item,避免覆盖原列表。
  • 增加错误处理:
    • 检查all_products_on_page是否存在,避免后续调用find_all时出错。
    • 对每个商品的名称和价格标签做判断,找不到时用默认值替代,防止触发AttributeError。
  • 避免数据覆盖:初始化一个总DataFrame,每次循环把当前页面的数据追加进去,最后一次性保存到CSV。
  • 新增来源URL字段:方便你区分商品来自哪个页面,后续分析更清晰。

内容的提问来源于stack exchange,提问作者PythonLearner23

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 09:07:35