You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取商品名称与价格返回NONE问题排查求助

问题解决:爬取商品信息返回None的修复

问题描述

爬取lordgun网站的商品页面时,尝试提取商品名称、价格和分类,但输出结果均为None,无法获取预期信息。目标商品页为该网站下的specialized-turbo-kenevo-comp-bike-emtb-1商品页(带颜色参数)。

错误原因分析

  1. BeautifulSoup属性参数错误:你在使用soup.find()时,错误地给非class属性添加了下划线(比如name_、itemprop_)。BeautifulSoup仅对class属性需要用class_作为参数名,其他HTML属性直接使用原属性名即可(如name、itemprop)。
  2. 缺少请求头(可选但重要):部分网站会识别非浏览器请求,返回不完整内容,导致无法定位元素,建议添加User-Agent模拟浏览器请求。

修正后的代码

#Packages needed
from bs4 import BeautifulSoup
import requests
import pandas as pd

#Base website URL
baseurl = "https://www.lordgun.com/"

#New list to store all the links of the different bikes that will later be scraped
productlinks =[]

for x in range(1,2):
    # 添加headers模拟浏览器请求
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    r = requests.get(f'https://www.lordgun.com/_road-mtb-bikes-frames?p={x}', headers=headers)
    soup = BeautifulSoup(r.content, 'lxml')

    # 定位页面中的商品列表
    productlist = soup.find_all('article', class_="article product")

    for item in productlist:
        for link in item.find_all('a', href=True):
            # 处理链接,避免重复拼接baseurl
            full_link = link['href'] if link['href'].startswith('http') else baseurl + link['href']
            productlinks.append(full_link)

# 打印获取到的商品链接数量
print(len(productlinks))

# 测试单个商品链接
testlink = 'https://www.lordgun.com/specialized-turbo-kenevo-comp-bike-emtb-1?color=Gloss%20Dark%20Moss%20Green%20Oak%20Green'
r = requests.get(testlink, headers=headers)

soup = BeautifulSoup(r.content, 'lxml')

# 修正属性参数写法,提取商品信息
name = soup.find('h1', name='product title')
price = soup.find('div', class_= 'prd-price')
category = soup.find('span', itemprop='name')

# 打印提取到的文本内容,空值时显示提示
print(price.get_text(strip=True) if price else '价格未找到')
print(name.get_text(strip=True) if name else '名称未找到')
print(category.get_text(strip=True) if category else '分类未找到')

额外说明

  • 核心修正:将soup.find()中的错误参数(如name_、itemprop_)改为正确的属性名(name、itemprop);
  • 添加User-Agent请求头,降低被网站反爬拦截的概率;
  • 优化链接拼接逻辑,避免重复添加baseurl导致无效链接;
  • 增加空值判断,打印更友好的提示信息,而非直接输出None。

内容的提问来源于stack exchange,提问作者tpteun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 05:43:21