You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取Jumia电商网站时遇AttributeError: NoneType对象无get_text属性

解决Jumia爬虫AttributeError: 'NoneType' object has no attribute 'get_text'问题

问题场景

爬取Jumia尼日利亚站(jumia.com.ng)的iOS手机商品列表,遍历商品链接提取详情时触发以下错误:

Traceback (most recent call last):
  File "c:\Users\LP\Documents\jumia\jumia.py", line 32, in <module>       
    name = soup.find('h1', class_='-fs20 -pts -pbxs').get_text(strip=True)
AttributeError: 'NoneType' object has no attribute 'get_text'

错误原因

  1. 反爬拦截:Jumia的反爬机制识别到爬虫请求,返回的页面不是正常商品详情页(可能是验证码页、跳转页或空白内容),导致soup.find找不到目标元素,返回None。
  2. 页面结构变更:目标元素的class名称或HTML结构已更新,原选择器失效。
  3. 商品页缺失元素:部分商品可能没有评论、评分或某些特征项,直接调用get_text会报错。
  4. 请求头不完整:仅携带User-Agent不足以模拟正常浏览器请求,被服务器识别为异常请求。

解决方案

1. 完善请求头

补充浏览器常用请求头,模拟真实访问行为:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://www.jumia.com.ng/ios-phones/',
    'DNT': '1',
    'Connection': 'keep-alive',
    'Upgrade-Insecure-Requests': '1'
}

2. 增加元素存在性判断

在调用get_text前先检查元素是否存在,避免None对象调用属性:

# 提取商品名称
name_tag = soup.find('h1', class_='-fs20 -pts -pbxs')
name = name_tag.get_text(strip=True) if name_tag else '无名称'

# 提取价格
amount_tag = soup.find('span', class_='-b -ltr -tal -fs24')
amount = amount_tag.get_text(strip=True) if amount_tag else '无价格'

# 提取评分
review_tag = soup.find('div', class_='stars _s _al')
review = review_tag.get_text(strip=True) if review_tag else '无评分'

# 提取评论数
rating_tag = soup.find('a', class_='-plxs _more')
rating = rating_tag.get_text(strip=True) if rating_tag else '无评论'

3. 处理特征列表索引越界

特征列表可能不足6项,改用循环遍历而非固定索引取值:

features = soup.find_all('li', attrs={'style': 'box-sizing: border-box; padding: 0px; margin: 0px;'})
feature_texts = [feat.get_text(strip=True) for feat in features]

# 打印所有提取到的特征
print('Key Features')
for idx, feat in enumerate(feature_texts, 1):
    print(f"Feature {idx}: {feat}")

4. 添加异常捕获与请求延迟

  • 捕获请求和解析时的异常,避免程序崩溃
  • 添加请求间隔,降低被反爬拦截的概率:
import time

for link in productlinks:
    try:
        r = requests.get(link, headers=headers)
        # 检查响应状态码,判断是否被拦截
        if r.status_code != 200:
            print(f"请求失败,状态码:{r.status_code},链接:{link}")
            continue
        soup = BeautifulSoup(r.content, 'lxml')
        
        # 元素提取代码...
        
        time.sleep(1)
    except Exception as e:
        print(f"处理链接{link}时出错:{str(e)}")
        continue

5. 验证页面结构

如果修改后仍找不到元素,手动访问商品链接查看页面源码,确认目标元素的class名称是否变更,更新选择器。

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import time

baseurl = 'https://www.jumia.com'

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://www.jumia.com.ng/ios-phones/',
    'DNT': '1',
    'Connection': 'keep-alive',
    'Upgrade-Insecure-Requests': '1'
}

productlinks = []

# 爬取商品列表链接
for x in range(1, 51):
    try:
        r = requests.get(f'https://www.jumia.com.ng/ios-phones/?page={x}#catalog-listing/', headers=headers)
        if r.status_code != 200:
            print(f"列表页{x}请求失败,状态码:{r.status_code}")
            continue
        soup = BeautifulSoup(r.content, 'lxml')
        productlist = soup.find_all('article', class_='prd _fb col c-prd')
        
        for product in productlist:
            link = product.find('a', href=True)
            if link:
                productlinks.append(baseurl + link['href'])
        time.sleep(0.5)
    except Exception as e:
        print(f"爬取列表页{x}出错:{str(e)}")
        continue

print(f"共收集到{len(productlinks)}个商品链接")

# 遍历商品链接提取详情
for link in productlinks:
    try:
        r = requests.get(link, headers=headers)
        if r.status_code != 200:
            print(f"商品页请求失败,状态码:{r.status_code},链接:{link}")
            time.sleep(1)
            continue
        soup = BeautifulSoup(r.content, 'lxml')
        
        # 提取商品信息,增加存在性判断
        name_tag = soup.find('h1', class_='-fs20 -pts -pbxs')
        name = name_tag.get_text(strip=True) if name_tag else '无名称'
        
        amount_tag = soup.find('span', class_='-b -ltr -tal -fs24')
        amount = amount_tag.get_text(strip=True) if amount_tag else '无价格'
        
        review_tag = soup.find('div', class_='stars _s _al')
        review = review_tag.get_text(strip=True) if review_tag else '无评分'
        
        rating_tag = soup.find('a', class_='-plxs _more')
        rating = rating_tag.get_text(strip=True) if rating_tag else '无评论'
        
        features = soup.find_all('li', attrs={'style': 'box-sizing: border-box; padding: 0px; margin: 0px;'})
        feature_texts = [feat.get_text(strip=True) for feat in features]
        
        # 打印信息
        print(f"Name: {name}")
        print(f"Amount: {amount}")
        print(f"Review: {review}")
        print(f"Rating: {rating}")
        print('Key Features')
        for idx, feat in enumerate(feature_texts, 1):
            print(f"Feature {idx}: {feat}")
        print('')
        
        time.sleep(1)
    except Exception as e:
        print(f"处理链接{link}时出错:{str(e)}")
        time.sleep(1)
        continue

内容的提问来源于stack exchange,提问作者Miracle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 04:05:26